Image restoration model training method, device and electronic device
By determining the number of reference pixel points and search area based on feature images in the image restoration model, a collection of reference pixel points with high similarity is obtained, and feature aggregation is performed, the problem of unmet reconstruction requirements of deep learning super-resolution in different regions is solved, and an efficient image restoration effect is achieved.
Patent Information
- Application Number
- CN202510012308.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-06
AI Technical Summary
In the prior art, the reconstruction requirements of deep learning super-resolution in different regions have not been met, resulting in poor high-resolution image effects.
In the image restoration model, feature images are acquired based on the initial image, target pixel points and their search areas are determined, reference pixel points are determined based on the feature image, reference pixel points sets with high similarity are obtained, and feature aggregation is performed through the cross attention algorithm to generate an output image.
Efficient reconstruction for different regions is achieved, the clarity and fineness of image restoration is improved, and the reconstruction needs of different regions by deep learning super-resolution is met.
Smart Images

Figure CN119399081B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to an image restoration model training method, device and electronic equipment. Background Art
[0002] The image acquisition system is affected by the camera equipment or the acquisition environment, and will acquire blurred low-resolution images. Due to the lack of clarity, low-resolution images seriously interfere with subsequent image processing. Therefore, it is necessary to restore low-resolution images to high-resolution images.
[0003] In the prior art, deep learning super-resolution is usually used to process low-resolution images in order to obtain high-resolution images with good effects. In deep learning super-resolution, whether a convolutional network or a moving window transformer is used for processing, a fixed number of pixels are configured in a fixed area around each pixel to achieve super-resolution reconstruction. The method of using a unified pixel configuration in the prior art does not meet the reconstruction requirements of deep learning super-resolution for different regions, so the high-resolution images obtained in this way are poor. Therefore, how to meet the reconstruction requirements of deep learning super-resolution images for different regions has become an urgent problem to be solved. Summary of the invention
[0004] The embodiments of the present application provide an image restoration model training method, device, system and computer-readable storage medium to at least solve the problem in the related art that the reconstruction requirements of deep learning super-resolution images for different regions cannot be met.
[0005] In a first aspect, an embodiment of the present application provides an image restoration model training method, characterized by comprising:
[0006] Acquire a feature image according to an initial image;
[0007] Determine a target pixel point and a search area for the target pixel point in the initial image, and determine the number of reference pixels for reconstructing the target pixel point based on the feature image;
[0008] In the search area, obtaining the similarity between each pixel and the target pixel, and obtaining a reference pixel set in descending order of similarity according to the number of reference pixels;
[0009] Acquire output features according to the reference pixel set and the target pixel, and acquire an output image according to the output features and the feature image;
[0010] The initial image is used as input data, a loss function is constructed using the output image and the feature image, and the generative network model is trained until the loss function is less than or equal to a preset threshold, thereby obtaining an image restoration model.
[0011] In one embodiment, determining a target pixel point and a search area for the target pixel point in the initial image includes:
[0012] Taking the target pixel point as the geometric center, obtaining a local search area in the initial image according to a preset size;
[0013] Taking the target pixel point as the starting point, sampling the initial image according to a preset step length to obtain a global search area;
[0014] The search area is determined according to the local search area and the global search area.
[0015] In one embodiment, the feature image includes a first feature image and a second feature image, and the reference pixel point set satisfies the following configuration:
[0016]
[0017] in, D F is the number of reference pixels, D max is the reference pixel number threshold, norm is the maximum normalization function, F is the first feature image, F ↓↑ is the second feature image, and “*” is the ceiling function.
[0018] In one embodiment, the loss function is constructed based on the output image and the feature image:
[0019] Based on the output image and the feature image, obtaining a content loss function, wherein the content loss function represents the detail richness of the output image;
[0020] Based on a preset discriminant loss function and the content loss function, obtaining a loss function;
[0021] The specific configuration of the loss function is:
[0022]
[0023] in, l is the loss function, l g is the content loss function, l d is the preset discriminant loss function.
[0024] In one embodiment, the content loss function is specifically configured as follows:
[0025]
[0026] in, l g is the content loss function, F is the first eigenimage, F ↓↑ is the second feature image, I HR is the image to be processed, y is the output image ,mse ( I HR ,y ) is the mean square error between the output image and the feature image.
[0027] In one embodiment, the reference pixel set satisfies the following configuration:
[0028]
[0029] in, p x is a set of reference pixels, u is the pixel in the search area, S x is the search area, sim ( F x ,F u ) is the cosine similarity function, F x is the feature of the target pixel, F u is the feature of the pixel in the search area, D FX yes Refer to the number of pixels.
[0030] In one embodiment, the step of acquiring output features according to the reference pixel set and the target pixel, and acquiring an output image according to the output features and the feature image, includes:
[0031] splicing the features in the reference pixel set to obtain reference features;
[0032] Based on the reference feature and the feature of the target pixel point, feature aggregation is performed by a cross attention algorithm to obtain the output feature;
[0033] Performing upsampling processing on the output features and the feature image to obtain the output image;
[0034] The upsampling process is specifically configured as follows:
[0035]
[0036] in, y For the output image, Upsample is the upsampling function, F is the feature image, F agg is the output feature.
[0037] In one embodiment, the step of acquiring a feature image according to an initial image includes:
[0038] Acquire a plurality of the images to be processed, perform degradation processing on the images to be processed, and obtain the initial images;
[0039] Extracting image features from the initial image to obtain the first feature image;
[0040] Downsampling and downsampling processing are performed on the first feature image to obtain a second feature image.
[0041] In a second aspect, an embodiment of the present application provides an image restoration device, comprising:
[0042] A feature image acquisition module is used to acquire a feature image according to an initial image;
[0043] A determination module, configured to determine a target pixel and a search area for the target pixel in the initial image, and determine the number of reference pixels for reconstructing the target pixel based on the feature image;
[0044] A reference pixel point set acquisition module is used to acquire the similarity between each pixel point and the target pixel point in the search area, and acquire the reference pixel point set in descending order of similarity according to the number of reference pixel points;
[0045] An output image acquisition module is used to acquire output features according to the reference pixel point set and the target pixel point, and acquire an output image according to the output features and the feature image;
[0046] The image restoration model module is used to use the initial image as input data, construct a loss function with the output image and the feature image, train the generative network model until the loss function is less than or equal to a preset threshold, and obtain an image restoration model.
[0047] In a third aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the image restoration model training method as described in the first aspect above is implemented.
[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image restoration model training method as described in the first aspect above.
[0049] An image restoration model training method, device, and electronic device provided in the embodiments of the present application have at least the following technical effects.
[0050] The required number of reference pixels for the target pixel is determined by the feature image, and the same number of pixels with high similarity to the target pixel are obtained in the search area of the target pixel. In the above manner, different target pixels in the initial image will be determined by the feature image with different numbers of reference pixels, so that according to the number of parameter pixels, the target pixel will be mapped to multiple new pixels, which can improve the clarity of the image restoration. The pixels in the reference pixel set are all pixels with high similarity to the target pixel, so that the rich details in the initial image can be retained and the fineness of the image restoration can be improved. In the above manner, different pixels are allocated to the target pixel as needed, so as to meet the reconstruction requirements of different areas of the deep learning super-resolution map.
[0051] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0053] Figure 1 is a flowchart of an image restoration model training method according to an exemplary embodiment;
[0054] Figure 2 is a block diagram of an image restoration model training device according to an exemplary embodiment;
[0055] Figure 3 is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0057] Obviously, the drawings described below are only some examples or embodiments of the present application. For ordinary technicians in this field, the present application can also be applied to other similar scenarios based on these drawings without creative work. In addition, it can also be understood that although the efforts made in this development process may be complicated and lengthy, for ordinary technicians in this field related to the content disclosed in this application, some changes in design, manufacturing or production based on the technical content disclosed in this application are just conventional technical means, and should not be understood as insufficient content disclosed in this application.
[0058] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0059] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the" and the like involved in this application do not indicate a quantity limitation, and may indicate the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0060] In a first aspect, the present application provides an image restoration model training method. Figure 1 is a flowchart of an image restoration model training method according to an exemplary embodiment. Figure 1 As shown, the image restoration model training method includes:
[0061] Step S101, obtaining a feature image according to an initial image.
[0062] The initial image is subjected to feature extraction through a convolutional layer to obtain a first feature image. The size of the first feature image is H×W×C, where H is the height of the feature image, i.e., the number of pixels of the feature image in the vertical direction, W is the width of the feature image, i.e., the number of pixels of the feature image in the horizontal direction, and C is the number of color information of each pixel in the channel number image. The first feature image is downsampled to reduce the size of the first feature image to one-half, and then the reduced first feature image is upsampled to the original size to obtain a second feature image. After the first feature image is downsampled and downsampled, some detail information will be lost in the second feature image. When the initial image is downsampled to the first feature image, the total number of pixels of the first feature image is reduced compared to the initial image, i.e., the pixel information of the initial image is lost. The first feature image is upsampled to the size of the initial image to obtain the second feature image, and the pixels in the second feature image are also interpolated based on the lost pixel information. Since the pixels in the flat area of the initial image are similar, the flat area of the second feature image obtained by interpolation of similar pixels is basically consistent with the flat area of the initial image, and the pixel value does not change much. However, the difference between the pixels in the detail-rich area of the initial image is large. After losing some pixel information, the pixel values of the second feature image obtained by interpolation of other pixels in the detail-rich area are very different from the pixel values of the detail-rich area of the initial image. The greater the difference between the pixel values of the second feature image and the pixel values of the initial image, the richer the details in the detail-rich area of the initial image.
[0063] The first feature image and the second feature image after downsampling and upsampling highlight the areas with rich details in the initial image, providing a basis for subsequent image restoration.
[0064] Step S102: determining a target pixel and a search area for the target pixel in the initial image, and determining the number of reference pixels for reconstructing the target pixel based on the feature image.
[0065] In the initial image, each pixel is processed as a target pixel, and the processing structure of each pixel is as follows:
[0066] Determine the search area of the target pixel in the initial image, and the search area includes a local search area and a global search area. In the initial image, take the target pixel as the geometric center, and construct a local search area in the neighborhood of the target pixel according to a preset size. For example, the preset size is 5×5, and the 5×5 neighborhood with the target pixel as the geometric center is used as the local search area. In the initial image, take the target pixel as the starting point, sample in the initial image according to the preset step size, and use all the sampled pixels as the global search area. For example, the preset step size is 2 pixels, and stride sampling is performed in the initial image according to 2 pixels, and all the sampled pixels are used as the global search area. The union of the local search area and the global search area is used as the search area.
[0067] The search area obtained based on the local search area and the global search area avoids searching all the pixels of the initial image, which reduces the amount of calculation for the subsequent calculation of the pixels and the target pixels. According to the feature image, the number of reference pixels for reconstructing the target pixel is determined. The reference pixel set meets the following configuration:
[0068]
[0069] in, D F is the number of reference pixels, D max is the reference pixel number threshold, norm is the maximum normalization function, F is the first feature image, F ↓↑ is the second feature image, and “*” is the ceiling function.
[0070] It should be noted that the input value is normalized to the range of 0 to 1 through the maximum normalization function to ensure that different features have similar sizes. The reference pixel number threshold is set manually. The larger the reference pixel number is set, the more complex the calculation. The upward forensics function can make the reference pixel number range from 0 to the reference pixel number threshold.
[0071] The search area of the target pixel provides a pixel sampling area for reconstructing the target pixel, and the pixel sampling area is reduced from all pixels of the initial image to the search area. Sampling pixels in the search area increases the sampling rate. The number of reference pixels of the target pixel determines the number of sampled pixels required to reconstruct the target pixel according to the richness of the details, so that the number of different sampling pixels can be determined according to the richness of the details of the target pixel, and enough pixels are provided for the target pixel, so that the target pixel can be better reconstructed, and the reconstructed image can be clearer.
[0072] Step S103: in the search area, obtain the similarity between each pixel and the target pixel, and obtain a reference pixel set in descending order of similarity according to the number of reference pixels.
[0073] In the search area, the cosine similarity between each pixel in the search area and the target pixel is calculated to obtain the similarity between each pixel and the target pixel. Based on the order of similarity from large to small, the pixels with large similarity are selected as the elements of the reference pixel set until the number of pixels in the reference pixel set is consistent with the number of reference pixels. Among them, the reference pixel set meets the following configuration:
[0074]
[0075] in, p x is a set of reference pixels, u is the pixel in the search area, S x is the search area, sim ( F x ,F u ) is the cosine similarity function, F x is the feature of the target pixel, F u is the feature of the pixel in the search area, D Fx is the number of reference pixels.
[0076] The reference pixel set is used to collect pixel points similar to the target pixel point, and each pixel point in the reference pixel set is assigned to the target pixel point, providing enough similar pixel points for reconstructing the target pixel point, thereby making the reconstructed target pixel point more accurate and improving the accuracy of the final image restoration.
[0077] Step S104: Obtain output features according to the reference pixel set and the target pixel, and obtain an output image according to the output features and the feature image.
[0078] Obtaining output features specifically includes the following steps:
[0079] Step S100: splicing the features in the reference pixel point set to obtain reference features.
[0080] The features of each pixel in the reference pixel set are concatenated to obtain the reference features. D Fx pixels, and the characteristic size of each pixel is 1×C. DFx Features concatenated into D Fx × C The matrix of D Fx × C The matrix of is multiplied by the C×HW feature of the input image to obtain D Fx ×HW similarity diagram. D Fx The similarity graph of ×HW is matrix multiplied with the transposed features of HW×C to obtain D Fx ×C's reference characteristics.
[0081] Step S200: Based on the reference features and the features of the target pixel points, feature aggregation is performed through a cross-attention algorithm to obtain output features.
[0082] Through the cross attention algorithm, the features of the target pixel are used as input data and the reference features are used as queries to complete the modeling features. The obtained features are aggregated through several layers of features, and the aggregation results are passed through a single convolution layer to adjust the aggregation results and obtain the output features.
[0083] Based on the output features obtained in step S100 to step S0200, up-sampling processing is performed on the output features and the feature image to obtain an output image.
[0084] Based on the deep learning super-resolution framework, upsampling is performed on the output features and the first feature image to obtain an output image. Optionally, the upsampling process uses a PixelUnshuffle operation.
[0085] The specific configuration of upsampling processing is:
[0086]
[0087] in, y For the output image, Upsample is the upsampling function, F is the first feature image, F agg is the output feature.
[0088] In step S104, output features are obtained by feature aggregation, and an output image with clearer resolution and more details is obtained based on the output features and the feature image.
[0089] Step S105: Use the initial image as input data, construct a loss function with the output image and the feature image, train the generative network model until the loss function is less than or equal to a preset threshold, and obtain an image restoration model.
[0090] Constructing a loss function based on the output image and the feature image includes the following steps:
[0091] Step S500: Based on the output image and the feature image, a content loss function is obtained, where the content loss function represents the detail richness of the output image.
[0092] The image to be processed is processed to obtain an initial image. The image to be processed is an image obtained by a camera device, and the image to be processed is a high-resolution image. Compared with the image to be processed, the output image has loss content of lost pixels. The deviation between the output image and the image to be processed is obtained to construct a content loss function. Optionally, the method for calculating the deviation between the output image and the image to be processed includes mean square error and mean absolute error. Preferably, the method for calculating the deviation is mean square error.
[0093] Based on the output image and the deviation between the output image and the image to be processed, a content loss function is constructed. The content loss function characterizes the detail richness of the output image and is used to focus on learning detail-rich areas. The content loss function specifically includes:
[0094]
[0095] in, l g is the content loss function, F is the first eigenimage, F ↓↑ is the second feature image, I HR is the image to be processed, y is the output image ,mse ( I HR ,y ) is the mean square error between the output image and the feature image.
[0096] Through the content loss function, the difference between the detail-rich areas of the output image and the image to be processed can be continuously reduced, so that the detail-rich areas of the output image can maintain sufficient details and clarity. The difference between the first feature image and the second feature image represents the richness of the details. The larger the difference, the greater the loss value of the content loss function, which makes the gradient of back propagation larger, so that the model learns faster, so as to achieve the purpose of focusing on learning detail-rich areas.
[0097] Step S600, obtaining a loss function based on a preset discrimination loss function and a content loss function;
[0098] The preset discriminant loss function represents the cross entropy loss of classification, which is used to discriminate the difference between the output image and the image to be processed. The preset discriminant loss function and the content loss function are used as loss functions to realize gradient backpropagation during model training.
[0099] The specific configuration of the loss function is:
[0100]
[0101] in, l is the loss function, l g is the content loss function, l d is the preset discriminant loss function.
[0102] Continuing to refer to step S105, the generative network model is trained with the initial image as input data, and the output image is obtained through steps S101 to S104. The loss value of the model is obtained through the loss function. If the loss value of the model is less than or equal to the preset threshold, the image restoration model is obtained. In the image restoration model, if the loss value of the model is less than or equal to the preset threshold, the output image can be used as the restored image.
[0103] In the image restoration model, the generative network model includes a generator and a discriminator. The process of obtaining the output image through steps S101 to S104 is equivalent to the generator, which is used to generate the output image. The discriminator, that is, adding a preset discriminant loss function, discriminates the output image in the generator, discriminates the generated output image and the real image to be processed, so that the output image can successfully deceive the discriminator to achieve a result that is indistinguishable from the real thing.
[0104] In addition, the initial image in step S101 is a low-resolution image. The initial image can only be obtained by performing degradation processing on several images to be processed. The images to be processed are high-resolution images collected from actual scenes by a camera device. The images to be processed are cropped, and the size of each image to be processed is aligned by padding to construct a set of images to be processed. Before each step of training, the images to be processed are read from the set of images to be processed, and blur processing, upsampling and downsampling, noise processing and JPEG compression are randomly combined to perform degradation processing on the images to be processed. The results of the degradation processing are downsampled to obtain the initial image.
[0105] Through the above processing method, a low-resolution initial image is obtained, and there is a mapping relationship between the low-resolution initial image and the high-resolution image to be processed. Whether the image reconstructed from the initial image reaches a high resolution is ultimately based on the image to be processed as a reference standard.
[0106] In one embodiment, the QR code can store data, and the data includes pointing to a website, contact information, and text, so the QR code is an indispensable content in modern society. The QR code is scanned by a mobile phone camera, and the stored data is obtained by decoding the QR code. However, in some actual scenarios, the QR code scanned by the mobile phone camera accounts for too small a proportion of the camera's field of view, and the influence of the camera's focal length makes the acquired QR code blurred, resulting in decoding failure. For blurred QR codes, the bicubic interpolation method or the moving window transformer SwinIR method is generally used for image restoration to obtain a super-resolution image. However, the images restored by these two methods still have the problem of blurring, and the decoding rate of the QR code is not high. Using the above-mentioned image restoration model, a clearer QR code image can be obtained, thereby improving the decoding rate of the QR code.
[0107] The low-resolution image of the two-dimensional code is processed by the bicubic interpolation method, the moving window transformer SwinIR method and the image restoration model. The two-dimensional code image obtained by the bicubic interpolation method is blurred, and the two-dimensional code image obtained by the moving window transformer SwinIR method is over-sharpened. The two-dimensional code image obtained by the image restoration model in this application has high definition and restores the details. By comparing the results of the three methods, it can be obtained that the image restoration model has the best restoration effect. And after laboratory comparison, in 500 frames of blurred two-dimensional code images, the bicubic interpolation method is used to restore and decode, and the decoding rate is 48.8%; the moving window transformer SwinIR method is used to restore and decode, and the decoding rate is 56.6%; the image restoration model is used to restore and decode, and the decoding rate is 100%.
[0108] In summary, the image restoration model training method provided by the embodiment of the present application determines the search area and the number of reference pixels based on the target pixel points, determines the number of pixels required for reconstruction according to different target pixel points, and allocates pixels with high similarity to the target pixel points according to the number of pixels for feature aggregation, thereby achieving the effect of on-demand reconstruction for different target pixel points, with more pixels allocated to areas with rich details and fewer pixels allocated to flat areas, thereby improving the clarity of the image restored in areas with rich details and solving the reconstruction requirements of deep learning super-resolution images for different areas. In addition, the search area is composed of a local search area and a global search area, which avoids calculating the similarity of all pixels, thereby achieving the effect of reducing the amount of calculation.
[0109] In a second aspect, an embodiment of the present application provides an image restoration model training device. Figure 2 FIG. 1 is a block diagram of an image restoration model training device according to an exemplary embodiment. Figure 2 As shown, the image restoration model training device includes:
[0110] A feature image acquisition module is used to acquire a feature image according to an initial image;
[0111] A determination module, used to determine the target pixel and the search area of the target pixel in the initial image, and determine the number of reference pixels for reconstructing the target pixel based on the feature image;
[0112] A reference pixel point set acquisition module is used to acquire the similarity between each pixel point and the target pixel point in the search area, and acquire the reference pixel point set in descending order of similarity according to the number of reference pixel points;
[0113] An output image acquisition module is used to acquire output features according to a reference pixel set and a target pixel, and to acquire an output image according to the output features and a feature image;
[0114] The image restoration model module is used to use the initial image as input data, construct a loss function with the output image and feature image, train the generative network model until the loss function is less than or equal to a preset threshold, and obtain an image restoration model.
[0115] In summary, the image restoration model training device provided by the present application determines the required number of reference pixels of the target pixel through the feature image, and obtains the same number of pixels with high similarity to the target pixel in the search area of the target pixel. In the above manner, different target pixels in the initial image will have different numbers of reference pixels determined by the feature image, so that according to the number of parameter pixels, the target pixel will be mapped to multiple new pixels, which can improve the clarity of the image restoration. The pixels in the reference pixel set are all pixels with high similarity to the target pixel, so that the rich details in the initial image can be retained and the fineness of the image restoration can be improved. In the above manner, different pixels are allocated to the target pixel as needed, so as to meet the reconstruction requirements of different areas of the deep learning super-resolution map.
[0116] It should be noted that the image restoration model training device provided in this embodiment is used to implement the above-mentioned implementation mode, and the description has been made no further. As used above, the terms "module", "unit", "sub-unit", etc. can implement a combination of software and / or hardware for a predetermined function. Although the device described in the above embodiment is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0117] In a third aspect, an embodiment of the present application provides an electronic device, Figure 3 FIG. 1 is a block diagram of an electronic device according to an exemplary embodiment. Figure 3 As shown, the electronic device may include a processor 81 and a memory 82 storing computer program instructions.
[0118] Specifically, the processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0119] Among them, the memory 82 may include a large-capacity memory for data or instructions. By way of example and not limitation, the memory 82 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 82 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 82 may be inside or outside the data processing device. In a specific embodiment, the memory 82 is a non-volatile memory. In a specific embodiment, the memory 82 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. Under appropriate circumstances, the RAM can be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM can be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0120] The memory 82 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81 .
[0121] The processor 81 implements any one of the image restoration model training methods in the above embodiments by reading and executing computer program instructions stored in the memory 82 .
[0122] In one embodiment, the image restoration model training device may further include a communication interface 83 and a bus 80. Figure 3 As shown, the processor 81, the memory 82, and the communication interface 83 are connected via a bus 80 and communicate with each other.
[0123] The communication interface 83 is used to implement communication between the modules, devices, units and / or equipment in the embodiment of the present application. The communication port 83 can also implement data communication with other components such as: external devices, image / data acquisition equipment, databases, external storage, and image / data processing workstations.
[0124] The bus 80 includes hardware, software, or both, and couples the components of the image restoration model training device to each other. The bus 80 includes, but is not limited to, at least one of the following: a data bus, an address bus, a control bus, an expansion bus, and a local bus. By way of example and not limitation, bus 80 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, bus 80 may include one or more buses. Although embodiments of the present application describe and illustrate a particular bus, the present application contemplates any suitable bus or interconnect.
[0125] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the image restoration model training method provided in the first aspect.
[0126] The readable storage medium may include but is not limited to: a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device or any suitable combination of the above.
[0127] In a possible implementation, the present invention can also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of the image restoration model training method provided in the first aspect.
[0128] The program code for executing the present invention may be written in any combination of one or more programming languages, and may be executed entirely on a user device, partially on a user device, as an independent software package, partially on a user device and partially on a remote device, or entirely on a remote device.
[0129] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0130] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for training an image restoration model, characterized in that: include: Acquire a number of images to be processed, perform degradation processing on the images to be processed, and obtain initial images; Extracting image features from the initial image to obtain a first feature image; Performing downsampling and upsampling processing on the first feature image to obtain a second feature image; Determine a target pixel point and a search area for the target pixel point in the initial image, and determine the number of reference pixels for reconstructing the target pixel point based on the first feature image and the second feature image, wherein the number of reference pixels satisfies the following formula: in, D F is the number of reference pixels, D max is the reference pixel number threshold, norm is the maximum normalization function, F is the first feature image, F ↓↑ is the second feature image, “*” is the upward rounding function; In the search area, obtaining the similarity between each pixel and the target pixel, and obtaining a reference pixel set in descending order of similarity according to the number of reference pixels; Acquire output features according to the reference pixel set and the target pixel, and acquire an output image according to the output features and the feature image; The initial image is used as input data, a loss function is constructed using the output image and the feature image, and the generative network model is trained until the loss function is less than or equal to a preset threshold, thereby obtaining an image restoration model.
2. The image restoration model training method according to claim 1, characterized in that: The step of determining a target pixel point and a search area for the target pixel point in the initial image includes: Taking the target pixel point as the geometric center, obtaining a local search area in the initial image according to a preset size; Taking the target pixel point as the starting point, sampling the initial image according to a preset step length to obtain a global search area; The search area is determined according to the local search area and the global search area.
3. The image restoration model training method according to claim 1, characterized in that: The step of constructing a loss function using the output image and the feature image comprises: Based on the output image and the feature image, obtaining a content loss function, wherein the content loss function represents the detail richness of the output image; Based on a preset discriminant loss function and the content loss function, obtaining a loss function; The loss function satisfies the following configuration: in, l is the loss function, l g is the content loss function, l d is the preset discriminant loss function.
4. The image restoration model training method according to claim 3, characterized in that: The content loss function is specifically configured as follows: in, l g is the content loss function, F is the first eigenimage, F ↓↑ is the second feature image, I HR is the image to be processed, y is the output image , mse ( I HR ,y ) is the mean square error between the output image and the feature image.
5. The image restoration model training method according to claim 1, characterized in that: The reference pixel set meets the following configuration: in, p x is a set of reference pixels, u is the pixel in the search area, S x is the search area, sim ( F x ,F u ) is the cosine similarity function, F x is the feature of the target pixel, F u is the feature of the pixel in the search area, D FX is the number of reference pixels.
6. The image restoration model training method according to claim 5, characterized in that: The step of acquiring output features according to the reference pixel set and the target pixel, and acquiring an output image according to the output features and the feature image, comprises: splicing the features in the reference pixel set to obtain reference features; Based on the reference feature and the feature of the target pixel point, feature aggregation is performed by a cross attention algorithm to obtain the output feature; Performing upsampling processing on the output features and the feature image to obtain the output image; The upsampling process is specifically configured as follows: in, y For the output image, Upsample is the upsampling function, F is the first feature image, F agg is the output feature.
7. An image restoration device, characterized in that: include: A first image acquisition module is used to acquire a number of images to be processed, and perform degradation processing on the images to be processed to obtain initial images; A second image acquisition module is used to extract image features from the initial image to obtain a first feature image; A third image acquisition module is used to perform downsampling and upsampling processing on the first feature image to obtain a second feature image; Determine a target pixel point and a search area for the target pixel point in the initial image, and determine the number of reference pixels for reconstructing the target pixel point based on the first feature image and the second feature image, wherein the number of reference pixels satisfies the following formula: in, D F is the number of reference pixels, D max is the reference pixel number threshold, norm is the maximum normalization function, F is the first feature image, F ↓↑ is the second feature image, “*” is the upward rounding function; A determination module, configured to obtain the similarity between each pixel and the target pixel in the search area, and obtain a reference pixel set in descending order of similarity according to the number of reference pixels; A reference pixel point set acquisition module is used to acquire output features according to the reference pixel point set and the target pixel point, and acquire an output image according to the output features and the feature image; The image restoration model module is used to use the initial image as input data, construct a loss function with the output image and the feature image, train the generative network model until the loss function is less than or equal to a preset threshold, and obtain an image restoration model.
8. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the image restoration model training method as described in any one of claims 1 to 6 when executing the computer program.
Citation Information
Patent Citations
Image identification model training method and device, equipment and medium
CN113505820A
Agricultural aerial image processing method and system based on artificial intelligence
CN113538290A