A model training method and device for super-resolution of real complex scene images

By combining complex degradation models and key point pairing technology in the model training method in the field of computer image super-resolution, the problem of performance degradation in the real world degradation space is solved, and a highly generalized natural image super-resolution processing is achieved.

CN115205112BActive Publication Date: 2025-05-13HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210675580.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-05-13
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

Existing computer image super-resolution methods have significantly reduced performance when facing the vast degradation space in the real world, making it difficult to effectively process real and complex scene images.

Method used

A model training method for super-resolution of real and complex scene images is proposed. By acquiring sample images and initial super-resolution models, the key points with dimension invariance in the sample images are determined, and high-resolution images and low-resolution images are paired one by one to generate a registered image pair data set. Then, the relationship parameters between the low-resolution image and the degraded image are determined based on the complex degradation model, and the relationship parameters between the super-resolution image and the high-resolution image are determined using the initial super-resolution model to establish a target model.

Benefits of technology

By simulating the real degradation data set, the broad degradation space is covered, the generalization of the network model is improved, and the super-resolution processing of natural images in real and complex scenarios is realized, with the characteristics of strong generalization and wide application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205112B_ABST
    Figure CN115205112B_ABST
Patent Text Reader

Abstract

The present application provides a model training method and device for super-resolution of real complex scene images, by obtaining sample images and an initial super-resolution model; determining the key points of the sample images, and pairing the key points in the high-resolution image and the low-resolution image one by one to generate a registered image pair data set; determining the first relationship parameter between the low-resolution image and the degraded image in the registered image pair data set based on a complex degradation model; determining the second relationship parameter between the super-resolution image generated by the degraded image and the high-resolution image in the registered image pair data set based on the initial super-resolution model; and establishing a target model based on the second relationship parameter and the initial super-resolution model. By combining high- and low-resolution image pairs with complex degradation models, a wide degradation space is covered, and the generalization of the network model is improved. The super-resolution model has the characteristics of strong generalization and wide application, and can be used for super-resolution of natural images in a variety of real scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of computer image super-resolution, and in particular to a model training method and device for super-resolution of real complex scene images. Background Art

[0002] In the field of computer image super-resolution, traditional methods use interpolation algorithms and information in the original image to infer the value of the unknown pixel in the high-resolution image. The principle is simple, but it only relies on a defined linear interpolation kernel and input image for reconstruction, resulting in the loss of high-frequency information in the reconstructed image, and often causing edge blur and jaggedness in the reconstructed image. Thanks to the development of machine learning and the improvement of computing power, learning-based super-resolution algorithms have developed rapidly and have been widely used in recent years. The academic community has fully applied deep learning knowledge to the field of super-resolution, and the reconstruction effect has been significantly improved compared with traditional methods. However, since there are very few samples of low-resolution and high-resolution image pairs in the real world, and the low-resolution images used in experiments are often derived from simple linear degradation of high-resolution images, the final generated model performs well under simulated degradation conditions, but its performance drops sharply when facing the vast degradation space in the real world.

[0003] Therefore, there is a need for a method that can effectively perform super-resolution on real-world images to promote the practical application of super-resolution methods. Summary of the invention

[0004] In view of the above problems, the present application is proposed to provide a real complex scene image super-resolution model training method and device that overcomes the above problems or at least partially solves the above problems, including:

[0005] A super-resolution model training method for real complex scene images, which is used to perform super-resolution processing on real complex scene images.

[0006] The method comprises:

[0007] Acquire a sample image and an initial super-resolution model; wherein the sample image includes a high-resolution image and a low-resolution image;

[0008] Determining key points with size invariance in the sample image, and pairing the key points in the high-resolution image with the key points in the low-resolution image one by one to generate a registration image pair data set;

[0009] Determining a first relationship parameter between the low-resolution image and the degraded image in the registered image pair data set according to a complex degradation model;

[0010] Determining, according to the initial super-resolution model, a second relationship parameter between a super-resolution image generated by the degraded image and a high-resolution image in the registration image pair data set;

[0011] A target model is established according to the second relationship parameter and the initial super-resolution model.

[0012] Furthermore, the step of determining the key points with size invariance in the sample image, and pairing the key points in the high-resolution image and the low-resolution image one by one to generate a registered image pair data set includes:

[0013] Constructing a Gaussian difference pyramid based on the sample image, and determining a key point with scale invariance in the sample image based on the Gaussian difference pyramid; wherein the key point includes the position coordinates of the key point and the scale information of the key point;

[0014] Determine the main direction of the key point, and generate a key point description vector according to the key point and the main direction of the key point;

[0015] The key points of the high-resolution image and the low-resolution image are matched one by one according to the key point description vectors to generate a registered image pair data set.

[0016] Furthermore, the step of constructing a Gaussian difference pyramid based on the sample image and determining key points with scale invariance in the sample image based on the Gaussian difference pyramid includes:

[0017] Determining pixel values ​​of image detection points in the Gaussian difference pyramid;

[0018] Determine a first pixel value set of 8 pixel points adjacent to the image detection point;

[0019] Determine a second pixel value set of a pixel point at a corresponding position of an adjacent image in the same group as the image detection point in the Gaussian difference pyramid and eight adjacent pixel points;

[0020] When the pixel values ​​are all greater than or all less than the first pixel value set and the second pixel value set, the image detection point is determined to be the key point; wherein the key point is as shown in the following formula:

[0021] X=(x,y,σ)

[0022] Where X is the key point; (x, y) is the position coordinate of the key point; σ is the scale information of the key point.

[0023] Furthermore, the step of determining the main direction of the key point and generating a key point description vector according to the key point and the main direction of the key point includes:

[0024] Determine the direction information of each pixel of the key point within a preset radius; wherein the direction information includes the amplitude and argument of the pixel;

[0025] Establishing a histogram of a preset direction according to the direction information;

[0026] Gaussian weighting is performed on the direction information according to the distance from each pixel to the key point within a preset radius, and the highest column in the histogram is obtained as the main direction of the key point;

[0027] The key point description vector is generated according to the position coordinates, the scale information and the main direction.

[0028] Furthermore, the step of determining the first relationship parameter between the low-resolution image and the degraded image in the registered image pair data set according to the complex degradation model includes:

[0029] Degrading the low-resolution image in the registered image pair data set according to a blur kernel generated according to a first preset probability; wherein the blur kernel includes an isotropic Gaussian blur kernel and an anisotropic Gaussian blur kernel;

[0030] Degrading the low-resolution image in the registered image pair data set according to downsampling generated by the second preset probability; wherein the downsampling includes nearest neighbor downsampling, bilinear downsampling and bicubic downsampling;

[0031] The low-resolution image in the registration image pair data set is degraded according to the Gaussian noise generated by the third preset probability, the Poisson noise generated by the fourth preset probability, and the JPEG compression noise generated by the fifth preset probability.

[0032] Furthermore, the initial super-resolution model includes a first model and a second model, and the second relationship parameter includes a first relationship and a second relationship; the step of determining the second relationship parameter between the super-resolution image generated by the degraded image and the high-resolution image in the registration image pair data set according to the initial super-resolution model includes:

[0033] Determining a first relationship between the degraded image and the corresponding super-resolution image by using the first model;

[0034] A second relationship between the super-resolution image and the corresponding high-resolution image in the registered image pair dataset is determined by the second model.

[0035] Further, the first model includes a first sub-model and a second sub-model, and the first relationship includes a first sub-relationship and a second sub-relationship; and the step of determining the first relationship between the degraded image and the corresponding super-resolution image by using the first model includes:

[0036] Determining a first sub-relationship between the degraded image and the corresponding image feature through the first sub-model; wherein the first sub-model is composed of a first convolutional layer and a residual dense shrinkage block, and the residual dense shrinkage block includes a fusion network composed of a soft threshold function, a channel attention mechanism and a ResNet;

[0037] The second sub-relationship between the image feature and the corresponding super-resolution graph is determined by the second sub-model; wherein the second sub-model is composed of a sub-pixel convolution layer and a second convolution layer.

[0038] A super-resolution model training device for real complex scene images, the device is used to perform super-resolution processing on real complex scene images,

[0039] The device comprises:

[0040] An acquisition module, used to acquire a sample image and an initial super-resolution model; wherein the sample image includes a high-resolution image and a low-resolution image;

[0041] A registration module, used to determine key points with size invariance in the sample image, and to pair the key points in the high-resolution image with the key points in the low-resolution image one by one to generate a registered image pair data set;

[0042] A degradation module, used for determining a first relationship parameter between a low-resolution image and a degraded image in the registered image pair data set according to a complex degradation model;

[0043] A super-resolution module, used for determining, based on the initial super-resolution model, a second relationship parameter between a super-resolution image generated by the degraded image and a high-resolution image in the registration image pair data set;

[0044] A training module is used to establish a target model based on the second relationship parameter and the initial super-resolution model.

[0045] A computer device comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of a model training method for super-resolution of real complex scene images as described above are implemented.

[0046] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a model training method for super-resolution of real complex scene images as described above.

[0047] This application has the following advantages:

[0048] In an embodiment of the present application, a sample image and an initial super-resolution model are obtained; wherein the sample image includes a high-resolution image and a low-resolution image; key points with size invariance in the sample image are determined, and the key points in the high-resolution image and the low-resolution image are paired one by one to generate a registered image pair data set; a first relationship parameter between the low-resolution image and the degraded image in the registered image pair data set is determined based on a complex degradation model; a second relationship parameter between the super-resolution image generated by the degraded image and the high-resolution image in the registered image pair data set is determined based on the initial super-resolution model; a target model is established based on the second relationship parameter and the initial super-resolution model. This application combines the high- and low-resolution image pairs taken by a camera with a complex degradation model for the first time, simulates a real degraded data set to cover a wide degradation space, improves the generalization of the network model, and proposes a super-resolution model based on a residual shrinkage network and dense connections, which has the characteristics of strong generalization and wide application, and can be applied to natural image super-resolution in a variety of real scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the description of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0050] Figure 1 This is a flowchart of the steps of a model training method for real complex scene image super-resolution provided by an embodiment of the present application;

[0051] Figure 2 is a schematic diagram of a sample image provided by an embodiment of the present application;

[0052] Figure 3 is a schematic diagram of a sample image key point matching result provided by an embodiment of the present application;

[0053] Figure 4 is a schematic diagram of a pair of high and low resolution images after image registration provided by an embodiment of the present application;

[0054] Figure 5 is a structural schematic diagram of a complex degradation model provided in an embodiment of the present application;

[0055] Figure 6 is a structural schematic diagram of a super-resolution model provided in one embodiment of the present application;

[0056] Figure 7 It is a structural block diagram of a model training device for real complex scene image super-resolution provided by an embodiment of the present application;

[0057] Figure 8 It is a structural diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0058] In order to make the objects, features and advantages of the present application more obvious and understandable, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0059] Reference Figure 1 , shows a model training method for super-resolution of real complex scene images provided by an embodiment of the present application, wherein the model training method is applied to super-resolution processing of real complex scene images;

[0060] The method comprises:

[0061] S110, obtaining a sample image and an initial super-resolution model; wherein the sample image includes a high-resolution image and a low-resolution image;

[0062] S120, determining key points with size invariance in the sample image, and pairing the key points in the high-resolution image and the low-resolution image one by one to generate a registration image pair data set;

[0063] S130, determining a first relationship parameter between the low-resolution image and the degraded image in the registered image pair data set according to a complex degradation model;

[0064] S140, determining, according to the initial super-resolution model, a second relationship parameter between the super-resolution image generated by the degraded image and the high-resolution image in the registered image pair data set;

[0065] S150: Establish a target model according to the second relationship parameter and the initial super-resolution model.

[0066] In an embodiment of the present application, a sample image and an initial super-resolution model are obtained; wherein the sample image includes a high-resolution image and a low-resolution image; key points with size invariance in the sample image are determined, and the key points in the high-resolution image and the low-resolution image are paired one by one to generate a registered image pair data set; a first relationship parameter between the low-resolution image and the degraded image in the registered image pair data set is determined based on a complex degradation model; a second relationship parameter between the super-resolution image generated by the degraded image and the high-resolution image in the registered image pair data set is determined based on the initial super-resolution model; a target model is established based on the second relationship parameter and the initial super-resolution model. This application combines the high- and low-resolution image pairs taken by a camera with a complex degradation model for the first time, simulates a real degraded data set to cover a wide degradation space, improves the generalization of the network model, and proposes a super-resolution model based on a residual shrinkage network and dense connections, which has the characteristics of strong generalization and wide application, and can be applied to natural image super-resolution in a variety of real scenes.

[0067] Next, a model training method for super-resolution of real complex scene images in this exemplary embodiment will be further described.

[0068] As described in step S110, a sample image and an initial super-resolution model are obtained; wherein the sample image includes a high-resolution image and a low-resolution image.

[0069] In an embodiment of the present invention, the specific process of "obtaining a sample image and an initial super-resolution model; wherein the sample image includes a high-resolution image and a low-resolution image" in step S110 may be further explained in combination with the following description.

[0070] As an example, for the same scene, the camera is adjusted to aperture priority mode, and the camera aperture value is manually adjusted to control the depth of field so that the depth of field can cover the entire picture so that there is no blur in the entire picture; on the other hand, the camera is set to auto focus, white balance and exposure; in order to reduce camera noise, the ISO (sensitivity, a measure of the sensitivity of the film to light) of the camera is set as small as possible; in order to avoid lens shake caused by handheld camera, a camera tripod is used to fix the camera, and the focal length is adjusted under the condition of fixed position and lens angle to obtain two images of the same scene. Due to the different focal lengths, the size of this scene in the two images is different. The scene is larger at a large focal length. Therefore, the resolution of this scene in the two images with the same resolution is high and low, laying the foundation for the subsequent acquisition of super-resolution image pair data sets. Finally, the image data set obtained by shooting may partially have objects moving and blurring, and images that do not meet the requirements are discarded. Repeat the above method to shoot different focal length image pair data sets in real scenes as sample images.

[0071] In a specific implementation, referring to Figure 2 , are two sample images of the same object taken by a Canon camera at a focal length of 50mm and 25mm. The resolution of these two images is consistent. Due to the different focal lengths, the proportion of the object box in the two images is different, and its resolution in the two images is also different. The area contained in the red frame in the right image is consistent with the scene in the left image, but the resolution is lower. In order to obtain more valuable image pairs, the captured object needs to ensure that it will not move and the background will not change before and after adjusting the focal length, and the texture is rich. Repeat the above shooting method to obtain 1,000 pairs of different focal length image data sets of the same scene with rich textures and a wide variety.

[0072] As described in step S120, key points with size invariance in the sample image are determined, and key points in the high-resolution image and the low-resolution image are paired one by one to generate a registered image pair data set.

[0073] In one embodiment of the present invention, the specific process of "determining key points with size invariance in the sample image, and pairing the key points in the high-resolution image and the low-resolution image one by one to generate a registration image pair data set" in step S120 can be further explained in combination with the following description.

[0074] As described in the following steps, a Gaussian difference pyramid is constructed according to the sample image, and a key point with scale invariance in the sample image is determined according to the Gaussian difference pyramid; wherein the key point includes the position coordinates of the key point and the scale information of the key point;

[0075] As described in the following steps, determining the main direction of the key point, and generating a key point description vector according to the key point and the main direction of the key point;

[0076] As described in the following steps, the key points of the high-resolution image and the low-resolution image are paired one-to-one according to the key point description vectors to generate a registered image pair data set.

[0077] In an embodiment of the present invention, the specific process of "constructing a Gaussian difference pyramid according to the sample image, and determining key points with scale invariance in the sample image according to the Gaussian difference pyramid" can be further explained in combination with the following description.

[0078] As described in the following steps, the pixel value of the image detection point in the Gaussian difference pyramid is determined;

[0079] As described in the following steps, determining a first pixel value set of eight pixel points adjacent to the image detection point;

[0080] As described in the following steps, a second pixel value set of a pixel point at a corresponding position of an adjacent image in the same group as the image detection point in the Gaussian difference pyramid and eight neighboring pixel points is determined;

[0081] As described in the following steps, when the pixel values ​​are all greater than or all less than the first pixel value set and the second pixel value set, the image detection point is determined to be the key point; wherein the key point is shown in the following formula:

[0082] X=(x,y,σ)

[0083] Where X is the key point; (x, y) is the position coordinate of the key point; σ is the scale information of the key point.

[0084] It should be noted that the scale of an image can be understood in this way: for an image, the effects of observing the image from a close distance and from a distance are different. The former is clearer, while the latter is blurrier; the former is larger, while the latter is smaller. Through the former, you can see some detailed information of the image, while through the latter, you can see some outline information of the image. This is the scale of the image.

[0085] As an example, the image registration algorithm specifically adopts a scale-invariant feature transform (SIFT) algorithm, and the algorithm steps are divided into: constructing a Gaussian difference pyramid, locating key points, and calculating the main direction of the key points.

[0086] In a specific implementation, in order to align the two images of high resolution and low resolution in constructing Gaussian difference pyramid, it is necessary to ensure that the features in different scale spaces are accurately paired, so the extracted key points need to ensure scale invariance. Assuming that the scale of the original sample image is σ0, the image is first Gaussian pyramided, and the layers of the Gaussian pyramid are divided into groups. The first layer image of the bottom group is set to the original image, and the number of layers of each group is set to 6. The Gaussian pyramid is generated from bottom to top, and the bottom group is set to the first group. Starting from the first group, the i-th layer image in the group is the first layer image in the group. Use (2 1 / 3 ) i-1 The images of different scale spaces are obtained by Gaussian blurring with σ as the parameter (=1, 2, ..., 6); the first layer image in the jth group is the image obtained by sampling the second to last layer image in the j-1th group of the pyramid at alternate points (=2, 3, ...). The two-dimensional Gaussian function and the scale space calculation formula of the processed image are as follows:

[0087]

[0088] L(x,y,σ)=G(x,y,σ)*I(x,y)

[0089] Among them, σ is the variance of the Gaussian function, and I(x,y) is the original image.

[0090] It can be found that the scale of each layer in group 1 satisfies σ=(2 1 / 3 ) i-1 σ0,i=0,1,…6, so the scale of the second-to-last layer image in the first group is 2σ0, and therefore the scale of the first layer image in the second group of the pyramid is 2σ0, and so on, the scale of the i-th layer of the j-th group is σ=2 j-1+1 / 3*(i-1) σ0, so the scale of the image in the previous layer between two adjacent groups is twice that of the image in the next layer. For each layer of the pyramid, the images in the group are updated to 5 Gaussian difference images obtained by subtracting two adjacent images in the group. The Gaussian difference image reduces the low-frequency information of the image and further highlights the characteristics of the image, and finally obtains the Gaussian difference pyramid.

[0091] In a specific implementation, in the key point positioning, in order to make the detected key point stable and unchanged in the multi-scale space of the image, the pixel value of the image detection point in the Gaussian difference pyramid is compared not only with the values ​​of the adjacent 8 pixels, but also with the pixel value of the corresponding position in the adjacent image in the same group and the values ​​of the adjacent 8 pixels of the point in the adjacent image. A total of 26 pixel values ​​need to be compared. If all are greater than or less than, the point is an extreme point and is used as a key point. The above steps ensure that the key point has scale invariance, so that the corresponding key point can be detected in two images of the same scene but different resolutions. The key point information includes its position coordinates and scale information, which can be expressed as X = (x, y, σ).

[0092] In an embodiment of the present invention, the specific process of "determining the main direction of the key point, and generating a key point description vector according to the key point and the main direction of the key point" can be further explained in combination with the following description.

[0093] As described in the following steps, the direction information of each pixel of the key point within a preset radius is determined; wherein the direction information includes the amplitude and argument of the pixel;

[0094] As described in the following steps, a histogram of a preset direction is established according to the direction information;

[0095] As described in the following steps, Gaussian weighting is performed on the direction information according to the distance from each pixel to the key point within a preset radius, and the highest column in the histogram is obtained as the main direction of the key point;

[0096] As described in the following steps, the key point description vector is generated according to the position coordinates, the scale information and the main direction.

[0097] In a specific implementation, the main purpose of calculating the main direction of the key point is to ensure the rotation invariance of the key point so that the key point can remain consistent after the image is rotated. Therefore, the direction information of the key point in a certain neighborhood area is calculated here. For a key point, the amplitude m(x, y) and the angle θ(x, y) of all pixels in the area with a preset radius of 3*1.5σ with the key point as the center are calculated. The calculation formula is defined as follows:

[0098]

[0099] In order to count the direction information corresponding to each pixel in the neighborhood, a histogram is used for statistics. The columns in the histogram record the preset directions of 10 degrees, 20 degrees, ... 360 degrees. It is worth noting that since the distances of each pixel from the key point are different, the direction information is Gaussian weighted according to the distance from the key point, and finally the highest column is selected as the main direction of the key point. Finally, the above results are fused to obtain a key point description vector containing position, scale and direction information.

[0100] The above method is used to calculate all key points in two images of the same scene with different resolutions, and then the key points in the two images are paired according to the key point description, and the images are cropped according to the registration results of the feature points. Specifically, the image captured at a small focal length is cut out from the image that is consistent with the image captured at a large focal length. The above steps obtain multiple image pairs of different resolutions of the same object, that is, a super-resolution image pair dataset captured by a real camera.

[0101] In a specific implementation, Figure 2 The area in the right frame that is consistent with the scene in the left image is cut out to obtain a pair of high and low resolution images of the same scene. To this end, an image pairing algorithm is applied here. Figure 3 The following is a schematic diagram of the matching result of the sample image according to the key point description. According to this result, the left image taken at a small focal length can be intercepted to obtain the following Figure 4 The registered image pair of the same scene shown in Figure 4 From the zoomed-in results of the area in the lower right corner box of the two images, it can be seen that there is a significant difference in resolution between the two images.

[0102] As described in step S130, a first relationship parameter between the low-resolution image and the degraded image in the registered image pair data set is determined according to the complex degradation model.

[0103] It should be noted that the complex degradation model is composed of fuzzy, down-sampling, random noise generation of multiple random parameters and random order combination.

[0104] In an embodiment of the present invention, the specific process of "determining the first relationship parameter between the low-resolution image and the degraded image in the registered image pair data set according to the complex degradation model" in step S130 can be further explained in combination with the following description.

[0105] As described in the following steps, the low-resolution image in the registered image pair data set is degraded according to the blur kernel generated by the first preset probability; wherein the blur kernel includes an isotropic Gaussian blur kernel and an anisotropic Gaussian blur kernel;

[0106] As described in the following steps, the low-resolution image in the registered image pair data set is degraded according to the downsampling generated by the second preset probability; wherein the downsampling includes nearest neighbor downsampling, bilinear downsampling and bicubic downsampling;

[0107] As described in the following steps, the low-resolution image in the registered image pair data set is degraded according to the Gaussian noise generated by the third preset probability or the JPEG compression noise generated by the fourth probability.

[0108] It should be noted that an image may degrade over time in the real world, or during transmission and storage. Therefore, in order to cover as much degradation space as possible, the complex degradation model specifically includes three key factors: blur, downsampling, and noise level.

[0109] As an example, for a pair of registered high-resolution and low-resolution images, the low-resolution image is used as the low-resolution image in the dataset after simulating real degradation using a complex degradation model, thereby obtaining a dataset of high-resolution and low-resolution image pairs simulating real degradation. Due to the variability of parameters and various types of key factors, the dataset simulating real degradation covers a wide degradation space, improving the generalization of the network model from the perspective of the dataset.

[0110] In a specific implementation, blurring is one of the common degradation operations in the image field. This application uses the most commonly used isotropic Gaussian blur kernel and anisotropic Gaussian blur kernel in the super-resolution field. A coordinate system with a center point coordinate of (0,0) is established in the blur kernel with a kernel size of 2t+1. The weight k(I,J) of each point is calculated as shown in the following formula:

[0111]

[0112] Where N is the normalization constant, C is the spatial coordinate, and σ is the standard deviation.

[0113] When σ1=σ2, the blur kernel is an isotropic Gaussian blur kernel; otherwise, it is anisotropic. In order to cover as much degradation space as possible, random selection is made on various parameters: in the selection of kernel size, t of both blur kernels is uniformly sampled from integers between 2 and 10; in the selection of standard deviation, the variance σ of the isotropic Gaussian blur kernel is 2 The variance of the anisotropic Gaussian blur kernel is chosen with medium probability from [0.2, 2] and without rotation. and The values ​​are selected from [0.2, 2] with equal probability and the rotation angle θ is selected from [0, π] with equal probability. In order to ensure that the size of the blurred image remains unchanged, reflection padding is used to fill the input image.

[0114] In a specific implementation, the purpose of the downsampling operation is to scale the image to the expected ratio, simulating as much as possible the degradation of the image that may occur during the transmission process due to compression and other operations. In order to ensure randomness, the downsampling operation is randomly selected with medium probability among nearest neighbor downsampling, bilinear downsampling and bicubic downsampling. These three methods are also the most commonly used downsampling methods in the field of super-resolution. The neighbor downsampling is that the pixel value of the low-resolution image is the same as that of the nearest neighbor image; bilinear downsampling is to perform single linear interpolation for each direction, and then interpolate, which is recorded as bilinear downsampling; bicubic downsampling is that the value of the function f at the point (x, y) can be obtained by the weighted average of the nearest sixteen sampling points in the rectangular grid. Here, two polynomial interpolation cubic functions are needed, one for each direction.

[0115] In a specific implementation, the noise level is different due to different sources. The noise includes three-dimensional Gaussian noise with random parameters, Poisson noise with random parameters, and JPEG compression noise with a quality factor randomly selected within [0, 90]. Specifically, the present application mainly considers the most commonly used Gaussian noise, Poisson noise and JPEG compression noise in the image field. In terms of Gaussian noise, since the image color channel is three-dimensional, the present invention adopts a three-dimensional Gaussian noise model, which obeys the three-dimensional Gaussian distribution of N(0,Σ), where Σ=σ 2M is a 3×3 covariance matrix. Considering the noise caused by the change in the number of photons received by the CMOS photosensitive element of some cameras or cameras, Poisson noise is introduced to simulate this noise when collecting images. Poisson noise conforms to the Poisson distribution and is proportional to the pixel value on the image, that is, the larger the image pixel value, the greater the frequency of Poisson noise; unlike Gaussian noise, Poisson noise is related to the pixel value of the image and the noise of each pixel is independent. The Poisson parameter is selected with equal probability between the quality factors [2,4]. In terms of JPEG compression noise, JPEG compression is the most commonly used method in the image compression process, so JPEG compression noise is also included in the degradation model. The compression degree is affected by the quality factor in the range of [0,100]. The lower the quality factor, the higher the compression degree and the lower the quality of the compressed image. In order to avoid too high a quality factor leading to too small compression noise, this paper selects with equal probability between the quality factors [0,90] and uses JPEG to compress the input image to simulate JPEG compression noise. In addition, due to the common use of JPEG compression, the last step of the complex degradation model is uniformly set to JPEG compression noise.

[0116] Existing degradation models usually adopt A linear process where y is the degraded image, x is the original image, s is the downsampling, and n is the noise. However, in the real world, the degradation of an image through compression, transmission, editing and other operations is complex and diverse. Specifically, an image from the Internet may have camera blur and sensor noise during the process of being taken by a camera; JPEG compression may be used to store the image; or unknown blur and noise may be introduced during the image degradation process in order to upload it to media software. Therefore, this application proposes a complex degradation model in which the key factors of the degradation model appear randomly and are randomly combined.

[0117] Reference Figure 5, is a structural diagram of a complex degradation model. In order to solve the above problems, the present application proposes a complex degradation model of quadratic nonlinear combination, specifically, two nonlinear combination degradations are applied. Each nonlinear combination degradation is composed of Gaussian blur, downsampling, Gaussian noise, and Poisson noise in random order. Specifically, a fuzzy operation is applied, and the first preset probability generated is 0.8; downsampling is a necessary operation, so the second preset probability generated is 1, and it is randomly selected with medium probability among nearest downsampling, bilinear downsampling, and bicubic downsampling; the third preset probability of generating Gaussian noise is 0.6; the fourth preset probability of generating Poisson noise is 0.2; considering the common use of image JPEG compression, JPEG compression is added at the end of each nonlinear combination degradation, and the fifth preset probability generated by JPEG compression in the two nonlinear combination degradations is 1. Finally, the probability of the first nonlinear combination degradation is 1, and considering that the process of the secondary nonlinear combination degradation is complex and the probability of occurrence is small, the probability of the second nonlinear combination degradation is 0.1. Therefore, for the complex degradation model proposed in this paper, an input image is degraded by the first inevitable nonlinear combination and the possible second nonlinear combination to obtain a low-resolution image that simulates the real degradation. It is worth noting that the two nonlinear combinations are unrelated. The schematic diagram is shown in Figure 2-4 As shown in Figure 2, due to the quadratic nonlinear combination degradation strategy and the randomization of parameters in each operation, the degradation space has been greatly expanded.

[0118] As described in step S140, a second relationship parameter between the super-resolution image generated by the degraded image and the high-resolution image in the registered image pair dataset is determined according to the initial super-resolution model.

[0119] It should be noted that the initial super-resolution model includes a first model and a second model, and the second relationship parameters include a first relationship and a second relationship.

[0120] In one embodiment of the present invention, the specific process of "determining the second relationship parameter between the super-resolution image generated by the degraded image and the high-resolution image in the registration image pair data set based on the initial super-resolution model" in step S140 can be further explained in combination with the following description.

[0121] As described in the following steps, determining a first relationship between the degraded image and the corresponding super-resolution image by using the first model;

[0122] As described in the following steps, a second relationship between the super-resolution image and the corresponding high-resolution image in the registered image pair dataset is determined by the second model.

[0123] In an embodiment of the present invention, the specific process of “determining the first relationship between the degraded image and the corresponding super-resolution image by using the first model” may be further explained in combination with the following description.

[0124] It should be noted that the first model includes a first sub-model and a second sub-model, and the first relationship includes a first sub-relationship and a second sub-relationship.

[0125] As described in the following steps, determining a first sub-relationship between the degraded image and the corresponding image feature through the first sub-model; wherein the first sub-model is composed of a first convolutional layer and a residual dense shrinkage block, and the residual dense shrinkage block includes a fusion network composed of a soft threshold function, a channel attention mechanism and a ResNet;

[0126] As described in the following steps, a second sub-relationship between the image feature and the corresponding super-resolution graph is determined through the second sub-model; wherein the second sub-model is composed of a sub-pixel convolution layer and a second convolution layer.

[0127] It should be noted that, refer to Figure 6 , showing a structural schematic diagram of a super-resolution model provided by an embodiment of the present application, the super-resolution model includes a first model consisting of 1 convolutional layer and 8 residual dense shrinkage blocks (Residual ShrinkageDenseBlock, RSDB), and a sub-pixel convolution upsampling module (second model), RSDB is composed of 4 residual shrinkage networks (Residual Shrinkage Network, RSN) densely connected, and the dense connection alleviates the gradient vanishing problem caused by network deepening.

[0128] As an example, for an input degraded image I LR , whose dimensions are H*W*C, representing the height, width and number of channels of the image respectively. The first model uses the first convolutional layer to obtain I LR Then, 8 consecutive residual dense shrinkage blocks with channel attention mechanism are used to extract I LR The stride of the convolutional layer in the first model is kept as 1, and finally the shallow features are fused with the deep features as the image feature F LR , the feature dimension is H*W*N.

[0129] Specifically, a residual dense shrinkage block is composed of four residual shrinkage networks densely connected. The dense connection is equivalent to an implicit supervision method, which alleviates the gradient vanishing problem caused by the deepening of the network and has an optimization effect on the back propagation gradient. Since the features are repeatedly transmitted between multiple residual shrinkage networks, the features are effectively reused. Among them, the residual shrinkage network combines the soft threshold function, the channel attention mechanism and ResNet. Based on ResNet, for the input x, the original output on the residual path is G(x). The residual shrinkage network is as follows Figure 6 As shown in the RSN in , the residual shrinkage network creates a branch on the residual path. This branch calculates the absolute average value of each channel through channel average pooling (Chanel Average Pool) to generate an absolute average vector with a dimension of 1*1*C0. The vector is sent to two learnable fully connected layers (Full Connected layer, FC) to obtain a result of 1*1*C0, and the number of channels remains unchanged. The above results are then processed by the sigmoid function so that all values ​​in the results are between 0 and 1. The absolute average vector is multiplied with it to obtain the threshold of each channel, and the threshold vector t is obtained. The dimension of the final threshold vector t is 1*1*C0. Finally, the output G(x) of the residual path is multiplied by the threshold vector t to obtain the soft thresholded feature as the feature on the residual path, then the output of the residual path is G(x)*t. Its effect is to reduce the noise and redundant information in the feature space during the feature learning process by using the threshold vector t to soft threshold the information on the residual path.

[0130] In a specific implementation, in the sub-pixel upsampling module, the input image features are converted into sub-pixel features T through a learnable sub-pixel convolutional layer, where the dimension of T is H*W*r 2 C, through the last fixed second convolutional layer, T is mapped to the output super-resolution image I SR Specifically, after obtaining the image feature F LR Finally, the image features are mapped to a super-resolution image I using a second convolution-based learnable model SR , let the scale factor be r, then I SR The dimension should be rH*rW*C. Specifically, for the image feature F LR , the number of channels used is r 2 C with a convolutional layer W and a bias b with a stride of 1 so that the output sub-pixel feature T has a dimension of H*W*r 2 C. Because when the step length When , the convolution layer can expand the height and width of the feature map to r times, then only one sub-pixel convolution is needed to convolve T to further change the dimension to rH*rW*C as I SR , its formula is expressed as follows:

[0131] I SR =PS(W*F LR +b)

[0132]

[0133] Where PS is the dimension H*W*r 2 The operation of periodically filtering the tensor of C to obtain a tensor of dimension rH*rW*C can be understood as a mapping and can also be implemented in the form of convolution.

[0134] As described in step S150, a target model is established according to the second relationship parameter and the initial super-resolution model.

[0135] In one embodiment of the present invention, the specific process of "establishing a target model according to the second relationship parameter and the initial super-resolution model" in step S150 may be further explained in combination with the following description.

[0136] As an example, a loss function is constructed, and the difference between the super-resolution image and the high-resolution image in the registration image pair data set is determined based on the loss function. The initial super-resolution model is trained based on the difference to obtain a target model.

[0137] In a specific implementation, the network model of the entire initial super-resolution model is defined, that is, a low-resolution image I LR To super-resolution image I SR The process is I for the high-resolution image after registration. HR , in order to measure I SR with I HR The difference between the two is that the loss function L of the entire network is composed of pixel loss L MSE And the content loss L content The composition is defined as follows:

[0138] L=λ mse *L MSE +λ content *L content

[0139]

[0140] In the formula, is the super-resolution image I SR The value at position (x,y), For high-resolution images I HR The value at position (x,y), H is the value of the image at position (x, y) in the feature map of the jth convolution output of the i-th layer in the VGG network (convolutional neural network), i,j and W i,j is the height and width of the feature map output by the jth convolution in the i-th layer of the VGG network, λ mse Usually set to 0.2, λ content Set to 1.

[0141] L MSE Will I HR with I SR Pixel-by-pixel comparison is also called pixel loss. However, due to its smoothing effect, the high-frequency information of the image may be too smooth, resulting in loss of image details. Therefore, content loss is added to supplement the image similarity difference. Two similar images maintain similarity in the feature space extracted by the same feature extraction network. Therefore, VGG-16 network is added to I HR and I SR Perform feature extraction and measure the mean square error of each point value on the feature map of each convolution output of each layer one by one, which is called content loss.

[0142] The network training process is as follows:

[0143] (1) Load the degraded image simulating real degradation and the high-resolution image in the registration image pair dataset, and initialize the super-resolution model;

[0144] (2) The batch size is set to 16, and 16 high-resolution and low-resolution image pairs are randomly selected from the above dataset;

[0145] (3) Normalize all images to [0,1], that is, divide the value of each pixel by 255;

[0146] (4) 16 low-resolution images are input into the initial super-resolution model, and 16 super-resolution images are output through forward propagation. The difference between the output image and the corresponding high-resolution image is measured through the loss function; wherein, the network parameters are updated using the Adam back-propagation algorithm through the loss value to minimize the loss, and the hyperparameters of Adam are = 0.9, = 0.999;

[0147] (5) Repeat steps (2) to (4).

[0148] Specifically, the learning rate is initialized to 10 in the network model. -4 , in every 10 5 After 5*10 iterations, the learning rate is halved. 3 The model is saved once for each iteration, and the target model is finally selected as the model with the smallest loss calculated for the image set in the validation set.

[0149] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0150] Reference Figure 7 , shows a model training device for super-resolution of real complex scene images provided by an embodiment of the present application, wherein the device is used to perform super-resolution processing on real complex scene images;

[0151] Specifically include:

[0152] The acquisition module 710 is used to acquire a sample image and an initial super-resolution model; wherein the sample image includes a high-resolution image and a low-resolution image;

[0153] A registration module 720, configured to determine key points with size invariance in the sample image, and perform one-to-one pairing of the key points in the high-resolution image and the low-resolution image to generate a registered image pair data set;

[0154] A degradation module 730, configured to determine a first relationship parameter between the low-resolution image and the degraded image in the registered image pair data set according to a complex degradation model;

[0155] A super-resolution module 740 is used to determine, based on the initial super-resolution model, a second relationship parameter between the super-resolution image generated by the degraded image and the high-resolution image in the registered image pair data set;

[0156] The training module 750 is used to establish a target model according to the second relationship parameter and the initial super-resolution model.

[0157] In one embodiment of the present invention, the registration module 720 includes:

[0158] A key point determination submodule, configured to construct a Gaussian difference pyramid based on the sample image, and determine a key point with scale invariance in the sample image based on the Gaussian difference pyramid; wherein the key point includes a position coordinate of the key point and scale information of the key point;

[0159] A main direction determination submodule, used to determine the main direction of the key point, and generate a key point description vector according to the key point and the main direction of the key point;

[0160] The registration image pair data set generation module is used to perform one-to-one pairing of key points of the high-resolution image and the low-resolution image according to the key point description vector to generate a registration image pair data set.

[0161] In one embodiment of the present invention, the key point determination submodule includes:

[0162] A pixel value determination unit, used to determine the pixel value of the image detection point in the Gaussian difference pyramid;

[0163] A first pixel value set determining unit, used to determine a first pixel value set of eight pixel points adjacent to the image detection point;

[0164] A second pixel value set determination unit, used to determine a second pixel value set of a pixel point at a corresponding position of an adjacent image in the same group as the image detection point in the Gaussian difference pyramid and eight neighboring pixel points;

[0165] A key point calculation unit, configured to determine that the image detection point is the key point when the pixel values ​​are all greater than or all less than the first pixel value set and the second pixel value set; wherein the key point is as shown in the following formula:

[0166] X=(x,y,σ)

[0167] Where X is the key point; (x, y) is the position coordinate of the key point; σ is the scale information of the key point.

[0168] In one embodiment of the present invention, the main direction determination submodule includes:

[0169] A direction information determination unit, used to determine the direction information of each pixel of the key point within a preset radius; wherein the direction information includes the amplitude and argument of the pixel;

[0170] A histogram establishing unit, used for establishing a histogram of a preset direction according to the direction information;

[0171] A weighting unit, configured to perform Gaussian weighting on the direction information according to the distance from each pixel to the key point within a preset radius, and obtain the highest column in the histogram as the main direction of the key point;

[0172] A key point description vector generating unit is used to generate the key point description vector according to the position coordinates, the scale information and the main direction.

[0173] In one embodiment of the present application, the degradation module 730 includes:

[0174] A blur submodule, used for degrading the low-resolution image in the registered image pair data set according to a blur kernel generated by a first preset probability; wherein the blur kernel includes an isotropic Gaussian blur kernel and an anisotropic Gaussian blur kernel;

[0175] A downsampling submodule, configured to degrade the low-resolution image in the registered image pair data set according to the downsampling generated by the second preset probability; wherein the downsampling includes nearest neighbor downsampling, bilinear downsampling and bicubic downsampling;

[0176] The noise submodule is used to degrade the low-resolution image in the registration image pair data set according to the Gaussian noise generated by the third preset probability, the Poisson noise generated by the fourth preset probability, and the JPEG compression noise generated by the fifth preset probability.

[0177] In an embodiment of the present application, the initial super-resolution model includes a first model and a second model, and the second relationship parameter includes a first relationship and a second relationship; the super-resolution module 740 includes:

[0178] a first relationship determination submodule, configured to determine a first relationship between the degraded image and the corresponding super-resolution image by using the first model;

[0179] The second relationship determination submodule is used to determine a second relationship between the super-resolution image and the corresponding high-resolution image in the registration image pair data set through the second model.

[0180] In an embodiment of the present application, the first model includes a first sub-model and a second sub-model, the first relationship includes a first sub-relationship and a second sub-relationship; and the first relationship determination submodule includes:

[0181] A first sub-relationship determination unit, configured to determine a first sub-relationship between the degraded image and the corresponding image feature through the first sub-model; wherein the first sub-model is composed of a first convolutional layer and a residual dense shrinkage block, and the residual dense shrinkage block includes a fusion network composed of a soft threshold function, a channel attention mechanism and a ResNet;

[0182] A second sub-relationship determination unit is used to determine a second sub-relationship between the image feature and the corresponding super-resolution graphic through the second sub-model; wherein the second sub-model is composed of a sub-pixel convolution layer and a second convolution layer.

[0183] Reference Figure 8 , showing a computer device of a real complex scene image super-resolution model training method of the present invention, which may specifically include the following:

[0184] The computer device 12 is in the form of a general-purpose computing device, and the components of the computer device 12 may include but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0185] The bus 18 represents one or more of several types of bus 18 structures, including a memory bus 18 or memory controller, a peripheral bus 18, an accelerated graphics port, a processor or a local bus 18 using any of a variety of bus 18 architectures. These architectures include, by way of example, but are not limited to, an Industry Standard Architecture (ISA) bus 18, a Micro Channel Architecture (MAC) bus 18, an Enhanced ISA bus 18, an Audio Video Electronics Standards Association (VESA) local bus 18, and a Peripheral Component Interconnect (PCI) bus 18.

[0186] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0187] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write to non-removable, non-volatile magnetic media (commonly referred to as a "hard drive"). Although Figure 8 Not shown, a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical medium) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory may include at least one program product having a set (e.g., at least one) of program modules 42, which are configured to perform the functions of various embodiments of the present invention.

[0188] A program / utility 40 having a set (at least one) of program modules 42 may be stored in, for example, a memory, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules 42, and program data, each of which or some combination may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of the embodiments described herein.

[0189] The computer device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, cameras, etc.), one or more devices that enable an operator to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the computer device 12 may also communicate with one or more networks (e.g., local area networks (LANs)), wide area networks (WANs), and / or public networks (e.g., the Internet) via a network adapter 20. Figure 8 As shown, the network adapter 20 communicates with other modules of the computer device 12 via the bus 18. It should be understood that although Figure 8 Not shown, other hardware and / or software modules may be used in conjunction with the computer device 12, including but not limited to: microcode, device drivers, redundant processing units 16, external disk drive arrays, RAID systems, tape drives, and data backup storage systems 34, etc.

[0190] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing a model training method for super-resolution of real complex scene images provided in an embodiment of the present invention.

[0191] That is, when the processing unit 16 executes the above program, it realizes: by acquiring a sample image and an initial super-resolution model; wherein the sample image includes a high-resolution image and a low-resolution image; determining the key points with size invariance in the sample image, and pairing the key points in the high-resolution image and the low-resolution image one by one to generate a registered image pair data set; determining the correspondence between the low-resolution image and the degraded image in the registered image pair data set according to a complex degradation model; determining the correspondence between the super-resolution image generated by the degraded image and the high-resolution image in the registered image pair data set according to the initial super-resolution model. This application combines the high- and low-resolution image pairs taken by the camera with a complex degradation model for the first time, simulates a real degraded data set to cover a wide degradation space, improves the generalization of the network model, and proposes a super-resolution model based on a residual shrinkage network and dense connections, which has the characteristics of strong generalization and wide application, and can be applied to natural image super-resolution in a variety of real scenes.

[0192] In an embodiment of the present invention, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a model training method for super-resolution of real complex scene images as provided in all embodiments of the present application.

[0193] That is, when the program is executed by the processor, it is implemented as follows: by obtaining a sample image and an initial super-resolution model; wherein the sample image includes a high-resolution image and a low-resolution image; determining the key points with size invariance in the sample image, and pairing the key points in the high-resolution image and the low-resolution image one by one to generate a registered image pair data set; determining the correspondence between the low-resolution image and the degraded image in the registered image pair data set based on a complex degradation model; determining the correspondence between the super-resolution image generated by the degraded image and the high-resolution image in the registered image pair data set based on the initial super-resolution model. This application combines the high- and low-resolution image pairs taken by a camera with a complex degradation model for the first time, simulating a real degraded data set to cover a wide degradation space, improving the generalization of the network model, and proposing a super-resolution model based on a residual shrinkage network and dense connections, which has the characteristics of strong generalization and wide application, and can be applied to natural image super-resolution in a variety of real scenes.

[0194] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, device, or device.

[0195] Computer-readable signal media may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0196] The computer program code for performing the operation of the present invention can be written in one or more programming languages ​​or a combination thereof, including object programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. The program code can be executed entirely on the operator's computer, partially on the operator's computer, as an independent software package, partially on the operator's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the operator's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet). The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other.

[0197] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0198] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0199] The above is a detailed introduction to the model training method and device for super-resolution of real complex scene images provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A model training method for super-resolution of real complex scene images, the model training method is applied to super-resolution processing of real complex scene images, characterized in that: The method comprises: Acquire a sample image and an initial super-resolution model; wherein the sample image includes a high-resolution image and a low-resolution image; Determining key points with size invariance in the sample image, and pairing the key points in the high-resolution image with the key points in the low-resolution image one by one to generate a registration image pair data set; Determine a first relationship parameter between a low-resolution image and a degraded image in the registered image pair data set according to a complex degradation model; specifically, degrade the low-resolution image in the registered image pair data set according to a blur kernel generated by a first preset probability; wherein the blur kernel includes an isotropic Gaussian blur kernel and an anisotropic Gaussian blur kernel; degrade the low-resolution image in the registered image pair data set according to downsampling generated by a second preset probability; wherein the downsampling includes nearest neighbor downsampling, bilinear downsampling and bicubic downsampling; degrade the low-resolution image in the registered image pair data set according to Gaussian noise generated by a third preset probability, Poisson noise generated by a fourth preset probability and JPEG compression noise generated by a fifth preset probability; Determining, according to the initial super-resolution model, a second relationship parameter between a super-resolution image generated by the degraded image and a high-resolution image in the registration image pair data set; A target model is established according to the second relationship parameter and the initial super-resolution model.

2. The method according to claim 1, characterized in that The step of determining the key points with size invariance in the sample image, and pairing the key points in the high-resolution image and the low-resolution image one by one to generate a registered image pair data set includes: Constructing a Gaussian difference pyramid based on the sample image, and determining a key point with scale invariance in the sample image based on the Gaussian difference pyramid; wherein the key point includes the position coordinates of the key point and the scale information of the key point; Determine the main direction of the key point, and generate a key point description vector according to the key point and the main direction of the key point; The key points of the high-resolution image and the low-resolution image are matched one by one according to the key point description vectors to generate a registered image pair data set.

3. The method according to claim 2, characterized in that The step of constructing a Gaussian difference pyramid based on the sample image and determining key points with scale invariance in the sample image based on the Gaussian difference pyramid includes: Determining pixel values ​​of image detection points in the Gaussian difference pyramid; Determine a first pixel value set of 8 pixel points adjacent to the image detection point; Determine a second pixel value set of a pixel point at a corresponding position of an adjacent image in the same group as the image detection point in the Gaussian difference pyramid and eight adjacent pixel points; When the pixel values ​​are all greater than or all less than the first pixel value set and the second pixel value set, the image detection point is determined to be the key point; wherein the key point is as shown in the following formula: X=(x,y,σ) Where X is the key point; (x, y) is the position coordinate of the key point; σ is the scale information of the key point.

4. The method according to claim 3, characterized in that: The step of determining the main direction of the key point and generating a key point description vector according to the key point and the main direction of the key point comprises: Determine the direction information of each pixel of the key point within a preset radius; wherein the direction information includes the amplitude and argument of the pixel; Establishing a histogram of a preset direction according to the direction information; Gaussian weighting is performed on the direction information according to the distance from each pixel to the key point within a preset radius, and the highest column in the histogram is obtained as the main direction of the key point; The key point description vector is generated according to the position coordinates, the scale information and the main direction.

5. The method according to claim 4, characterized in that The initial super-resolution model includes a first model and a second model, and the second relationship parameter includes a first relationship and a second relationship; the step of determining the second relationship parameter between the super-resolution image generated by the degraded image and the high-resolution image in the registration image pair data set according to the initial super-resolution model includes: Determining a first relationship between the degraded image and the corresponding super-resolution image by using the first model; A second relationship between the super-resolution image and the corresponding high-resolution image in the registered image pair dataset is determined by the second model.

6. The method according to claim 5, characterized in that The first model includes a first sub-model and a second sub-model, and the first relationship includes a first sub-relationship and a second sub-relationship; the step of determining the first relationship between the degraded image and the corresponding super-resolution image by using the first model includes: Determining a first sub-relationship between the degraded image and the corresponding image feature through the first sub-model; wherein the first sub-model is composed of a first convolutional layer and a residual dense shrinkage block, and the residual dense shrinkage block includes a fusion network composed of a soft threshold function, a channel attention mechanism and a ResNet; The second sub-relationship between the image feature and the corresponding super-resolution graph is determined by the second sub-model; wherein the second sub-model is composed of a sub-pixel convolution layer and a second convolution layer.

7. A model training device for super-resolution of real complex scene images, the device is used to perform super-resolution processing on real complex scene images, characterized in that ; The device comprises: An acquisition module, used to acquire a sample image and an initial super-resolution model; wherein the sample image includes a high-resolution image and a low-resolution image; A registration module, used to determine key points with size invariance in the sample image, and to pair the key points in the high-resolution image with the key points in the low-resolution image one by one to generate a registered image pair data set; A degradation module, used to determine a first relationship parameter between a low-resolution image and a degraded image in the registered image pair data set according to a complex degradation model; specifically, degrade the low-resolution image in the registered image pair data set according to a blur kernel generated by a first preset probability; wherein the blur kernel includes an isotropic Gaussian blur kernel and an anisotropic Gaussian blur kernel; degrade the low-resolution image in the registered image pair data set according to downsampling generated by a second preset probability; wherein the downsampling includes nearest neighbor downsampling, bilinear downsampling and bicubic downsampling; degrade the low-resolution image in the registered image pair data set according to Gaussian noise generated by a third preset probability, Poisson noise generated by a fourth preset probability and JPEG compression noise generated by a fifth preset probability; A super-resolution module, used for determining, based on the initial super-resolution model, a second relationship parameter between a super-resolution image generated by the degraded image and a high-resolution image in the registration image pair data set; A training module is used to establish a target model based on the second relationship parameter and the initial super-resolution model.

8. A computer device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the model processing method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the model processing method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Super-resolution image reconstruction method and system based on degradation model

    CN113538245A

  • Light field super-resolution three-dimensional reconstruction method and system

    CN113870433A