Super-Resolution Method and System Based on Local Autoregressive Model and Discrete Dictionary

Through the combination of local autoregression model and high-frequency discrete learning dictionary, the problem of image blur generated by super-resolution methods in the existing technology is solved, and the generation and calculation efficiency of high-definition pictures are improved.

CN114240748BActive Publication Date: 2025-07-25中央广播电视总台 +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111475883.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-06
Publication Date
2025-07-25
Estimated Expiration
2041-12-06

AI Technical Summary

Technical Problem

The super-resolution method based on deep neural networks in the prior art is poor in practical applications, resulting in blurred high-resolution pictures generated, and the traditional method is not clear enough when face reconstruction.

Method used

Local autoregression model and high-frequency discrete learning dictionary are used to achieve low-frequency recovery through the rough super-segment model, and high-definition pictures are generated using high-frequency discrete encoding and local autoregression methods.

Benefits of technology

The generated high-definition pictures are clearer, the model training is more stable, the calculation time consumption is reduced, and the effect is more robust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114240748B_ABST
    Figure CN114240748B_ABST
Patent Text Reader

Abstract

The present invention provides a super-resolution method based on a local autoregressive model and a discrete dictionary, comprising: obtaining a data pair of a high-definition image and a low-definition image, processing the data pair based on a coarse super-resolution model and a high-frequency discrete learnable dictionary to obtain a local autoregressive model capable of obtaining discrete coding; and the low-definition image to be restored is passed through the coarse super-resolution model and the local autoregressive model to obtain a super-resolution image. The present invention realizes preliminary super-resolution through a coarse super-resolution module, thereby realizing low-frequency restoration. At the same time, discrete coding is performed on the high frequency and generated by a local autoregressive method to realize the enhancement of the low-definition image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method in the fields of computer vision and image processing. Specifically, it relates to a super-resolution method and system based on a local autoregressive model and a discrete dictionary. Background Art

[0002] Super-resolution is one of the most fundamental tasks in computer vision. In the past few years, many methods based on deep neural networks have achieved great success. However, most of these methods simulate data to pursue pixel-level restoration and use regression methods, and their effects in practical applications are not good. The main reason for this is that the pixel-level loss function will lead to blurred results, which are not suitable for human eye perception.

[0003] After retrieval, the Chinese invention patent publication number is CN104036482A, and the application number is 201410323594.X. This invention discloses a face image super-resolution method based on dictionary asymptotic update: in the training stage, the leave-one-out method is used to perform super-resolution reconstruction on each low-resolution face image in the low-resolution face image training set to obtain a layer of low-resolution intermediate dictionary; using this low-resolution intermediate dictionary as the input of the new low-resolution face image training set, a new layer of low-resolution intermediate dictionary is reconstructed; repeating the above process, finally obtaining multiple layers of low-resolution intermediate dictionaries. In the testing stage, according to the input low-resolution face image, the upper-layer low-resolution intermediate dictionary and the high-resolution face image training set, super-resolution reconstruction is performed on the input low-resolution face image to obtain a predicted high-resolution face image; repeating the above process, finally reconstructing a high-resolution face image.

[0004] This patent uses traditional methods to construct a dictionary and generate faces, which has the problem of inaccurate modeling of face reconstruction problems, and may also lead to unclear generated high-resolution pictures. Summary of the Invention

[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide a super-resolution method and system based on a local autoregressive model and a discrete dictionary.

[0006] According to one aspect of the present invention, there is provided a super-resolution method based on a local autoregressive model and a discrete dictionary, including:

[0007] Obtaining a data pair of a high-definition picture and a low-definition picture,

[0008] Processing the data pair based on a coarse super-resolution model and a high-frequency discrete learnable dictionary to obtain a local autoregressive model capable of obtaining discrete coding;

[0009] The low-resolution image to be processed is passed through the coarse super-resolution model and the local autoregressive model to obtain a super-resolution image.

[0010] Preferably, the acquisition of the data pair of the high-resolution image and the low-resolution image includes: high-resolution training images, and high-resolution and low-resolution training data pairs are obtained by downsampling.

[0011] Preferably, the processing of the data pair based on the coarse super-resolution model and the high-frequency discrete learnable dictionary to obtain an autoregressive model capable of obtaining discrete coding includes:

[0012] Construct a coarse super-resolution model and a high-frequency discrete learnable dictionary according to the data pair;

[0013] The low-resolution image is input into the coarse super-resolution model to obtain a coarse super-resolution result, and the high-resolution image is input into the high-frequency discrete learnable dictionary for discrete coding;

[0014] The coarse super-resolution result and the discrete coding are used to train and obtain a local autoregressive model.

[0015] Preferably, the coarse super-resolution model includes several sub-modules, where the encoding convolutional network consists of several convolutional and max-pooling operations to extract the visual features of the image; the high-frequency discrete learnable dictionary consists of several learnable vectors; the decoding convolutional network consists of several convolutional layers and upsampling operations.

[0016] Preferably, the construction of the coarse super-resolution model and the high-frequency discrete learnable dictionary according to the data pair includes:

[0017] The high-resolution images in the dataset are represented as X hr , and the corresponding low-resolution image obtained by downsampling is X lr ;

[0018] The low-resolution image is X lr Feed it into the coarse super-resolution model to obtain the coarse super-resolution result X c ;

[0019] The coarse super-resolution result X c and the high-resolution image X hr are used as the input to the encoder of the high-frequency discrete learnable dictionary; the encoder outputs a feature map f hf , and for each pixel position's feature vector in it, find the entry in the high-frequency dictionary I hf of the high-frequency discrete learnable dictionary that has the closest Euclidean distance to it and replace it to obtain the high-frequency feature map f' hf ;

[0020] The high-frequency feature map f' hf and the coarse super-resolution result X cThrough the decoder of the high-frequency discrete learnable dictionary, the high-definition picture Y is restored. hr .

[0021] Preferably, for the coarse super-resolution model and the high-frequency discrete learnable dictionary, the optimization objectives include the optimization of the coarse super-resolution model and the optimization of the high-frequency discrete learnable dictionary and its codec, where:

[0022] For the optimization of the coarse super-resolution model, the mean absolute error L c-sr :

[0023] L c-sr = ||X c - X hr ||

[0024] X c = C(X lr )

[0025] For the optimization of the codec of the high-frequency discrete learnable dictionary, the reparameterization technique is used, and the optimization objective is the Euclidean distance L hr between X hr and Y recons ,

[0026] L recons = ||Y hr - X hr ‖

[0027] Y hr = D(f hf + [f′ hf - f hf , X c )

[0028] f hf = E(X hr , X c )

[0029] where D represents the decoder neural network, E represents the encoder neural network, and [*] represents the gradient truncation operation for reparameterization;

[0030] For the optimization of the high-frequency discrete learnable dictionary, the high-frequency discrete learnable dictionary needs to be updated according to the dataset, and the update of the dictionary entries is carried out in a clustering manner. In the forward propagation of the neural network of the high-frequency discrete learnable dictionary, for any entry there is

[0031]

[0032] where the summation symbol is the sum over all i, j that satisfy the condition ;

[0033] represents the updated entry, ε represents a constant used to increase the stability of convergence, and N represents the number of all (i, j) that satisfy . represents the feature at the (i, j) position in the feature map before replacement, represents the feature at the (i, j) position in the feature map after replacement.

[0034] Preferably, the low-resolution image is super-resolved based on a coarse super-resolution model to obtain a coarse super-resolution result, and the high-resolution image is discretely encoded based on the high-frequency discrete learnable dictionary; the coarse super-resolution result and the discrete encoding are used to train a local autoregressive model, including:

[0035] For the high-resolution image and the corresponding low-resolution image data pair in the dataset, the coarse super-resolution result X c is obtained through the coarse super-resolution model, and its high-frequency discrete encoding II hf is obtained through the high-frequency discrete learnable dictionary. For any value hf at the position in II it is calculated by the following formula:

[0036]

[0037] Divide into regular non-overlapping small blocks, with the length and width of each small block being s, and the size of II hf being H * W;

[0038] All regions can be divided into H / / s * W / / s non-overlapping small blocks; each small block contains s * s positions;

[0039] For the same position in each small block, the same label is assigned, and the label value ranges from 1 to s * s;

[0040] The positions with the same label will be generated simultaneously; the number of autoregressive times is s * s;

[0041] Denote all the pixel points with the same label t in II hf as

[0042] Fit the conditional posterior probability where is the k-th element in, is the set of all satisfying l < t,

[0043] The fitted posterior probability is used for sampling generation in the autoregressive process.

[0044] Preferably, for the local autoregressive model, its optimization objective includes the optimization of the local autoregressive neural network, where: for the prediction result obtained by the local autoregressive neural network Use cross-entropy to calculate its loss

[0045] 9. The super-resolution method based on the local autoregressive model and the discrete dictionary according to claim 7, wherein the local autoregressive model is constructed by various network structures, including LSTM, mask attention or mask CNN.

[0046] According to a second aspect of the present invention, there is provided a super-resolution system based on a local autoregressive model and a high-frequency discrete learnable dictionary, including:[[]]

[0047] Coarse super-resolution and high-frequency discrete learnable dictionary construction module: According to the high-definition picture and the corresponding low-definition picture, train the coarse super-resolution model and the high-frequency discrete learnable dictionary, and the dictionary entries of the high-frequency discrete learnable dictionary correspond to the high-frequency part in the high-definition picture;

[0048] Local autoregressive model construction module: Use deep learning to train the local autoregressive model according to the coarse super-resolution result and the high-frequency dictionary encoding obtained according to the high-frequency discrete learnable dictionary;

[0049] High-definition super-resolution picture generation construction module: Use the input low-definition picture to obtain the coarse super-resolution result, generate its corresponding high-frequency discrete encoding based on the coarse super-resolution result using the local autoregressive model, and decode the final super-resolution picture according to the coarse super-resolution result and the generated high-frequency discrete encoding.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] 1. The present invention provides a super-resolution method based on a local autoregressive model and a high-frequency discrete learnable dictionary, which realizes preliminary super-resolution through the coarse super-resolution module, thereby realizing low-frequency recovery. At the same time, discrete encoding is performed on the high-frequency part, and it is generated by the local autoregressive method to realize the enhancement of the low-definition picture.

[0052] 2. The present invention uses the local autoregressive method to generate high-definition pictures, which is more stable in training compared with other generative models, and saves a large amount of computational time consumption compared with traditional autoregressive models.

[0053] 3. The present invention uses high-low frequency separation and discrete encoding, achieving better results on low-definition pictures, and the model training is more robust. Description of the Drawings

[0054] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0055] Figure 1 Flowchart of a super-resolution method based on a local autoregressive model and a discrete dictionary according to an embodiment provided by the present invention;

[0056] Figure 2 Schematic structural diagram of a local autoregressive model according to a preferred embodiment provided by the present invention;

[0057] Figure 3 Schematic diagram of a super-resolution system based on a local autoregressive model and a high-frequency discrete learnable dictionary according to an embodiment provided by the present invention. Detailed implementation manners

[0058] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0059] The present invention provides an embodiment of a super-resolution method based on a local autoregressive model and a discrete dictionary, including:

[0060] Obtain a data pair of a high-definition image and a low-definition image,

[0061] Based on a coarse super-resolution model and a high-frequency discrete learnable dictionary, process the data pair to obtain a local autoregressive model capable of obtaining discrete coding;

[0062] The low-definition image to be restored passes through the coarse super-resolution model and the local autoregressive model to obtain a super-resolution image.

[0063] Based on the further optimization of the above embodiment, the present invention provides a preferred embodiment. As Figure 1 shown, it is a flowchart of a super-resolution method based on a local autoregressive model and a discrete dictionary according to this embodiment. It includes:

[0064] S11. For the input high-definition training image, first obtain a training data pair of high-definition and low-definition through downsampling. Input the low-definition image into the coarse super-resolution model to obtain a coarse super-resolution result; input the high-definition image and the coarse super-resolution result into the high-frequency discrete learnable dictionary at the same time to construct a high-frequency discrete learnable dictionary. The high-frequency discrete learnable dictionary and the coarse super-resolution result can be used to restore the high-definition image.

[0065] S12. For the obtained high-definition and low-definition training data pairs, use the coarse super-resolution model to obtain the coarse super-resolution result, use the high-frequency discrete learnable dictionary to obtain the high-frequency discrete coding, and use the coarse super-resolution result and the corresponding high-frequency discrete coding to train the local recurrent autoregressive model.

[0066] S13. For the low-definition pictures that need to be enhanced, first use the coarse super-resolution model to obtain the coarse super-resolution result. Then send the coarse super-resolution result into the local autoregressive model to generate its high-frequency discrete coding through autoregression. Finally, use the coarse super-resolution result and the high-frequency discrete coding to generate the corresponding high-definition picture.

[0067] In the above embodiments, the coarse super-resolution model and the high-frequency discrete learnable dictionary are used to realize the separation of high and low frequencies. For the low-frequency part, the coarse super-resolution model is used for restoration; for the high-frequency part in the picture, a high-frequency discrete learnable dictionary is constructed. Different processing methods are adopted for high and low frequencies, ensuring that the generated super-resolution pictures are clearer and more realistic while maintaining as much fidelity as possible and not generating errors. At the same time, the proposed local autoregressive method greatly reduces the time consumption of the autoregressive model.

[0068] To enhance the final fidelity effect and reduce the distortion of the original picture, the present invention provides a preferred embodiment. In the steps of constructing the coarse super-resolution module and the high-frequency discrete learnable dictionary in this embodiment, the low-definition part of the high-definition picture is restored using the super-coarse super-resolution module, and the high-frequency part that cannot be restored by the coarse super-resolution module is embedded into the discrete high-frequency discrete learnable dictionary and generated using the local autoregressive method. That is, different processing methods are adopted for high and low frequencies, so as to output pictures that are clearer and more realistic while maintaining as much fidelity as possible and not generating errors. The methods for constructing the coarse super-resolution model and the high-frequency discrete learnable dictionary are as follows:

[0069] S101. The high-definition facial feature pictures in the dataset are denoted as X hr , and the low-definition facial feature pictures obtained through downsampling are X lr ;

[0070] S102. The low-definition facial feature picture is X lr as the input of the coarse super-resolution convolutional network. For the coarse super-resolution result X c output by the convolutional network and the high-definition picture X hr , they are used as the input of the encoder of the high-frequency discrete learnable dictionary. For the output feature map f hf of the encoder, for the feature vector at each pixel position, find the entry in the high-frequency dictionary I hf with the closest Euclidean distance to it and replace it to obtain f' hf ;

[0071] S103. f' hf and X c pass through the decoder convolutional network to finally restore the high-definition picture Yhr 。

[0072] Through the construction of the coarse super-resolution model and the high-frequency discrete learnable dictionary in this preferred embodiment, the dictionary can be directly learned, enhancing the fidelity and clarity of the output.

[0073] In some preferred embodiments of the present invention, a coarse super-resolution module and a high-frequency discrete learnable dictionary are constructed, where: the coarse super-resolution module network consists of several convolutional layers. The encoding convolutional network consists of several convolutional layers and max-pooling operations to extract the visual features of the image; the high-frequency discrete learnable dictionary consists of several learnable vectors; the decoding convolutional network consists of several convolutional layers and upsampling operations.

[0074] To achieve precise optimization, the present invention uses a preferred embodiment. The coarse super-resolution model and the high-frequency discrete learnable dictionary are learned, and the optimization objectives include the optimization of the coarse super-resolution model and the optimization of the high-frequency discrete learnable dictionary and its encoders and decoders, where:

[0075] For the optimization of the coarse super-resolution model, the mean absolute error L c-sr :

[0076] L c-sr =||X c -X hr ||,

[0077] X c =C(X lr )

[0078] For the optimization of the high-frequency discrete learnable dictionary encoders and decoders, the reparameterization trick is used, and the optimization objective is the Euclidean distance L hr between X hr and Y recons

[0079] L recons =||Y hr -X hr ||,

[0080] Y hr =D(f hf +[f′ hf -f hf ,X c ),

[0081] f hf =E(X hr ,X c )

[0082] where D represents the decoder neural network, E represents the encoder neural network, and [*] represents the gradient truncation operation for reparameterization;

[0083] Meanwhile, for the optimization of the dictionary, the high-frequency discrete learnable dictionary needs to be updated according to the dataset. The update of the dictionary entries adopts a clustering method. The specific update method is that in the forward propagation of the neural network, for any entry there is

[0084]

[0085] The rightmost summation symbol in the above formula sums over all i, j that satisfy the condition ;

[0086] Among them, represents the updated entry, ε represents a relatively small constant used to increase the stability of convergence, N represents the number of all (i, j) that satisfy , represents the feature at the (i, j) position in the feature map before replacement, represents the feature at the (i, j) position in the feature map after replacement.

[0087] To reduce the time consumption of the autoregressive model, the present invention provides a preferred embodiment. In this embodiment, the coarse super-resolution model obtains the coarse super-resolution result X c , and obtains its high-frequency discrete encoding II hf through the high-frequency discrete learnable dictionary. For any value hf at any position in II it is calculated through Here . The steps for constructing the local autoregressive model include:

[0088] S201, the local autoregressive method divides into regular non-overlapping small blocks. Assuming that the length and width of the small blocks are both s, and the size of II hf is H*W, then all regions can be divided into H / / s * W / / s non-overlapping small blocks. Each small block contains s*s positions.

[0089] S202, for the same positions in each small block, the same label is given, and the label values range from 1 to s*s. In the process of local autoregression, the positions with the same label will be generated simultaneously. Therefore, the total number of autoregressions is s*s. We denote all the pixel points with the same label t in II hf as The local autoregressive neural network can be built with various network structures, such as LSTM, mask attention, mask CNN, etc. Its function is to fit the conditional posterior probability where is the k-th element in is all a set that satisfies l < t. The obtained posterior probability is used for sampling generation in the autoregressive process. As Figure 2 shown, it is a schematic diagram of the partitioning method with s = 4 in the local autoregressive model of this embodiment.

[0090] To achieve precise optimization, the present invention provides a preferred embodiment. The prediction result obtained by the local autoregressive neural network uses cross-entropy to calculate its loss

[0091]

[0092] Through the local autoregressive model of the above preferred embodiment, the time consumption problem of the traditional autoregressive model can be greatly reduced, so that the autoregressive model can be applied to the super-resolution problem.

[0093] Furthermore, in the above high-frequency discrete learnable dictionary encoding local autoregressive step, the autoregressive from the coarse super-resolution result to the high-frequency encoding is implemented by LSTM or mask attention or mask CNN. The internal structure makes the pixel of the current label unable to obtain the information of the pixel of this label and the pixels behind this label, so as to use the information before the pixel of this label to complete the fitting of the distribution of this pixel.

[0094] The above embodiments of the present invention utilize the coarse super-resolution model and the high-frequency discrete learnable dictionary to achieve high-low frequency separation and discrete encoding, and achieve better results on low-resolution images. Through high-low frequency separation and discrete encoding of the dictionary, a local autoregressive model is used to enhance low-resolution images.

[0095] In some embodiments of the present invention, for the high-definition image generation step, wherein: according to the input low-resolution image X lr input, use the coarse super-resolution network to obtain the coarse super-resolution result X c , generate high-frequency discrete encoding through the local autoregressive neural network Finally, use the decoder of the high-frequency discrete learnable dictionary to obtain the super-resolution result

[0096] Based on the same concept of the above embodiments, in another embodiment of the present invention, a super-resolution system of a local autoregressive model and a high-frequency discrete learnable dictionary is provided. As Figure 3 shown, it is a schematic diagram of the system of this embodiment. It includes:

[0097] A coarse super-resolution module and a high-frequency discrete learnable dictionary construction module: According to the high-definition image and the corresponding low-resolution image, train the coarse super-resolution model and the high-frequency discrete learnable dictionary. The dictionary entries of the high-frequency discrete learnable dictionary correspond to the high-frequency part in the high-definition image;

[0098] Local autoregressive model construction module: Use deep learning to train a local autoregressive model based on the result of coarse super-resolution and the high-frequency dictionary encoding obtained from the corresponding high-frequency discrete learnable dictionary.

[0099] High-definition super-resolution image generation construction module: Obtain the result of coarse super-resolution using the input low-resolution image, generate its corresponding high-frequency discrete encoding using the local autoregressive model based on the result of coarse super-resolution, and decode the final super-resolution image according to the result of coarse super-resolution and the generated high-frequency discrete encoding.

[0100] In summary, the above embodiments utilize the result of coarse super-resolution of a learnable coarse super-resolution model, discretely encode the high-frequency part of the image using a high-frequency discrete learnable dictionary, use the local autoregressive model to complete the generation from the result of coarse super-resolution to high-frequency dictionary encoding, and use the high-definition image generation module to generate the high-definition image corresponding to the final low-resolution image, thereby improving the fidelity and clarity of the images generated by the model, and at the same time greatly reducing the time consumption of the traditional autoregressive model.

[0101] The present invention can enhance low-resolution images using a publicly available dataset and achieve a good super-resolution effect.

[0102] It should be noted that the steps in the method provided by the present invention can be implemented by corresponding modules, devices, units, etc. in the system. Those skilled in the art can refer to the technical solution of the method to implement the composition of the system. That is, the embodiments in the method can be understood as preferred examples for constructing the system, which will not be elaborated here.

[0103] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention. The above preferred features can be combined arbitrarily without conflict.

Claims

1. A super-resolution method based on a local autoregressive model and a discrete dictionary, characterized in that, Including: Obtaining a data pair of a high-definition image and a low-definition image; Based on a coarse super-resolution model and a high-frequency discrete learnable dictionary, processing the data pair to obtain a local autoregressive model capable of obtaining discrete coding; The low-definition image to be processed passes through the coarse super-resolution model and the local autoregressive model to obtain a super-resolution image; Among them, the processing of the data pair based on the coarse super-resolution model and the high-frequency discrete learnable dictionary to obtain an autoregressive model capable of obtaining discrete coding includes: According to the data pair, constructing a coarse super-resolution model and a high-frequency discrete learnable dictionary; The low-definition image is input into the coarse super-resolution model to obtain a coarse super-resolution result, and the high-definition image is input into the high-frequency discrete learnable dictionary for discrete coding; the coarse super-resolution result and the discrete coding are trained to obtain a local autoregressive model. Specifically: A pair of high-definition images and corresponding low-definition image data, obtaining a rough super-resolution result X through the rough super-resolution model c , obtaining its high-frequency discrete encoding II through a high-frequency discrete learnable dictionary hf , for II hf The value at any position Calculate through the following formula: Divide into regular and non-overlapping small blocks, where the length and width of each small block are both s, II hf with a size of H*W; All regions are divided into H / / s * W / / s non-overlapping small blocks; each small block contains s * s positions; For the same positions in each small block, the same label is given, and the label values range from 1 to s * s; The positions with the same label will be generated simultaneously; the number of autoregressive times is s * s; Denote II hf all the pixel points with the same label t in Fitted conditional posterior probability where is the k-th element in and is the set of all The fitted posterior probability is used for sampling generation in the autoregressive process.

2. The super-resolution method based on the local autoregressive model and the discrete dictionary according to claim 1, characterized in that The obtaining of the data pair of the high-definition image and the low-definition image includes: the high-definition image is downsampled to obtain the low-definition image, and the high-definition image and the low-definition image form a data pair for training.

3. The super-resolution method based on a local autoregressive model and a discrete dictionary according to claim 1, wherein The coarse super-resolution model includes several sub-modules, and one sub-module is an encoding convolutional network composed of several convolutional layers and max-pooling operations for extracting visual features of the image; the high-frequency discrete learnable dictionary includes several learnable vectors, and one vector is a decoding convolutional network composed of several convolutional layers and upsampling operations.

4. The super-resolution method based on the local autoregressive model and the discrete dictionary according to claim 1, wherein The constructing of the coarse super-resolution model and the high-frequency discrete learnable dictionary according to the data pair includes: The high-definition images in the dataset are represented as X hr , and the corresponding low-definition images obtained through downsampling are X lr ; The low-resolution image is X lr Transmit the rough super-resolution model to obtain the rough super-resolution result X c ; The rough super-resolution result X c and the high-definition image X hr are used as the inputs of the encoder of the high-frequency discrete learnable dictionary; the encoder outputs a feature map f hf . For each pixel position of its feature vector, the entry in the high-frequency dictionary I hf of the high-frequency discrete learnable dictionary that has the closest Euclidean distance to it is found and replaced to obtain a high-frequency feature map f' hf ; The high-frequency feature map f' hf and the rough super-resolution result X c are passed through the decoder of the high-frequency discrete learnable dictionary to recover the high-definition image Y hr .

5. The super-resolution method based on the local autoregressive model and the discrete dictionary according to claim 4, characterized in that, For the coarse super-resolution model and the high-frequency discrete learnable dictionary, their optimization objectives include the optimization of the coarse super-resolution model and the optimization of the high-frequency discrete learnable dictionary and its encoders and decoders. Among them: For the optimization of the coarse super-resolution model, the mean absolute error L c-sr : L c-sr = ||X c -X hr || X c = C(X lr ) C is the coarse super-resolution model; The optimization of the codec for the high-frequency discrete learnable dictionary uses the reparameterization trick, and the optimization target is X hr and y hr The Euclidean distance L recons , L recons = ||Y hr -X hr || Y hr = D(f hf + [f′ hf - f hf , X c ) f hf = E(X hr , X c ) Among them, D represents the decoder neural network, E represents the encoder neural network, and [*] represents the gradient truncation operation for reparameterization operation; For the optimization of the high-frequency discrete learnable dictionary, the high-frequency discrete learnable dictionary is updated according to the data set. The update of the dictionary entries adopts a clustering method. In the forward propagation of the neural network of the high-frequency discrete learnable dictionary, for any entry there is wherein, the summation symbol sums over all i and j that satisfy the condition ; represents the updated entry, ε represents a constant that increases the stability of convergence, and N represents the number of all (i, j) that satisfy the quantity of (i, j), represents the feature at the (i, j) position in the feature map before replacement, represents the feature at the (i, j) position in the feature map after replacement.

6. The super-resolution method based on the local autoregressive model and the discrete dictionary according to claim 1, wherein The local autoregressive model, whose optimization objective includes the optimization of the local autoregressive neural network, where: for the prediction results obtained by the local autoregressive neural network Calculate its loss using cross-entropy 7. The super-resolution method based on a local autoregressive model and a discrete dictionary according to claim 1, characterized in that The local autoregressive model is built by various network structures, including LSTM, mask attention or mask CNN.

8. A super-resolution system based on a local autoregressive model and a high-frequency discrete learnable dictionary, characterized in that, Including: Coarse super-resolution and high-frequency discrete learnable dictionary construction module: According to the high-definition image and the corresponding low-definition image, training the coarse super-resolution model and the high-frequency discrete learnable dictionary, and the dictionary entries of the high-frequency discrete learnable dictionary correspond to the high-frequency part in the high-definition image; Local autoregressive model construction module: Using deep learning to train the local autoregressive model according to the coarse super-resolution result and the high-frequency dictionary coding obtained according to the high-frequency discrete learnable dictionary; High-definition super-resolution image generation construction module: Using the input low-definition image to obtain a coarse super-resolution result, generating its corresponding high-frequency discrete coding based on the coarse super-resolution result using the local autoregressive model, and decoding the final super-resolution image according to the coarse super-resolution result and the generated high-frequency discrete coding; Among them, processing the data pair based on the coarse super-resolution model and the high-frequency discrete learnable dictionary to obtain an autoregressive model capable of obtaining discrete coding includes: Constructing a coarse super-resolution model and a high-frequency discrete learnable dictionary according to the data pair; Inputting the low-resolution image into the coarse super-resolution model to obtain a coarse super-resolution result, and inputting the high-resolution image into the high-frequency discrete learnable dictionary for discrete coding; training the coarse super-resolution result and the discrete coding to obtain a local autoregressive model, specifically: A pair of high-definition images and corresponding low-definition image data, and a rough super-resolution result X is obtained through the rough super-resolution model c , and its high-frequency discrete encoding II is obtained through a high-frequency discrete learnable dictionary hf , for II hf The value at any position is calculated by the following formula: Divide into regular and non-overlapping small blocks, where the length and width of each small block are both s, II hf with a size of H*W; All regions are divided into H / / s*W / / s non-overlapping small blocks; each small block contains s*s positions; For the same positions in each small block, the same label is given, and the label values range from 1 to s*s; The positions with the same label will be generated simultaneously; the number of autoregressive times is s*s; Denote II hf all the pixel points with the same label t in Fitting conditional posterior probability where is the k-th element in and the set of all satisfies l < t The posterior probability obtained by fitting is used for sampling generation in the autoregressive process.

Citation Information

Patent Citations

  • Facial image super-resolution method based on dictionary asymptotic updating

    CN104036482A

  • A face image super-resolution method based on dictionary asymptotic updates

    CN104036482B

  • Face five-sense-organ super-resolution method and system based on learnable dictionary, and medium

    CN113628109A