Image Super-Resolution
Through multi-level attention information and cross-resolution feature integration technology, the problem of inaccurate texture migration in existing image super-resolution technologies is solved, and clearer and more realistic high-resolution images are generated.
Patent Information
- Application Number
- CN202010414770.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-05-15
AI Technical Summary
Existing image super-resolution techniques make it difficult to accurately migrate the texture features of the reference image, resulting in the generated high-resolution images being blurred and unreal at complex textures.
Through multi-level attention information, the texture features in the reference image are accurately migrated into the input image, combined with feature integration across different resolution scales, enhance feature expression capabilities and generate clearer and more realistic images.
Improves the accuracy of texture feature search and migration during image super-resolution, reduces texture blur and distortion, and generates clearer, more realistic high-resolution images.
Smart Images

Figure CN113674146B_ABST
Abstract
Description
Background Art
[0001] In the field of image processing, converting the resolution of an image is a common requirement. Image super-resolution (SR) refers to an image processing process of generating a high-resolution image with natural and clear textures based on a low-resolution image. Image super-resolution is a very important issue in the field of image enhancement. In recent years, thanks to the powerful learning ability of deep learning technology, this issue has made significant progress. Image super-resolution technology has been widely applied in various fields, such as digital zoom in digital camera photography, material enhancement in game remakes (e.g., texture material enhancement), etc. On the other hand, image super-resolution technology has further promoted the development of other issues in the field of computer vision, such as medical imaging, surveillance imaging, satellite imaging, etc. Currently, the solutions for image super-resolution mainly include single-image super-resolution (SISR) technology and reference-image-based super-resolution (RefSR) technology. Summary of the Invention
[0002] According to an implementation of the present disclosure, a solution for image processing is proposed. In this solution, first information and second information are determined based on the texture features of an input image and a reference image. The first information at least indicates a second pixel block in the reference image that is most relevant to a first pixel block in the input image according to the texture features, and the second information at least indicates the degree of relevance between the first pixel block and the second pixel block. A migration feature map with a target resolution is determined based on the first information and the reference image. The migration feature map includes a feature block corresponding to the first pixel block, and the feature block includes the texture features of the second pixel block. The input image is transformed into an output image with a target resolution based on the migration feature map and the second information. The output image embodies the texture features of the reference image. In this solution, using information related to texture features at the pixel block level can make the search and migration of texture features more accurate, so as to reduce texture blurring and texture distortion. In this way, this solution can efficiently and accurately migrate the texture of the reference image and obtain a clearer and more realistic image processing result.
[0003] The Summary of the Invention section is provided to introduce a selection of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention section is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Brief Description of the Drawings
[0004] Figure 1 A block diagram of a computing device capable of implementing multiple implementations of the present disclosure is shown;
[0005] Figure 2 An architecture diagram of a system for image processing according to an implementation of the present disclosure is shown;
[0006] Figure 3 A block diagram of a texture transformer according to some implementations of the present disclosure is shown;
[0007] Figure 4 A block diagram of cross-scale feature integration according to some implementations of the present disclosure is shown; and
[0008] Figure 5 A flowchart of a method for image processing according to an implementation of the present disclosure is shown.
[0009] In these figures, the same or similar reference signs are used to denote the same or similar elements. Detailed implementation
[0010] The present disclosure will now be described with reference to several example implementations. It should be understood that these implementations are described only to enable those of ordinary skill in the art to better understand and thus implement the present disclosure, and do not imply any limitation on the scope of the present disclosure.
[0011] As used herein, the term "comprising" and its variants are to be construed as open-ended terms meaning "including but not limited to". The term "based on" is to be construed as "at least partially based on". The terms "one implementation" and "an implementation" are to be construed as "at least one implementation". The term "another implementation" is to be construed as "at least one other implementation". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included hereinafter.
[0012] As used herein, a "neural network" is capable of processing inputs and providing corresponding outputs, and generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications typically include many hidden layers, thus increasing the depth of the network. The layers of a neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of a neural network includes one or more nodes (also referred to as processing nodes or neurons), and each node processes the input from the previous layer. A CNN is a type of neural network that includes one or more convolutional layers for performing convolutional operations on respective inputs. A CNN can be used in various scenarios and is particularly suitable for processing image or video data. In this article, the terms "neural network", "network", and "neural network model" may be used interchangeably.
[0013] Example environment
[0014] Figure 1 A block diagram of a computing device 100 capable of implementing multiple implementations of the present disclosure is shown. It should be understood thatFigure 1 The computing device 100 shown is merely exemplary and should not impose any limitation on the functions and scope of the implementations described in this disclosure. As Figure 1 shown, the computing device 100 includes a computing device 100 in the form of a general-purpose computing device. The components of the computing device 100 may include, but are not limited to, one or more processors or processing units 110, a memory 120, a storage device 130, one or more communication units 140, one or more input devices 150, and one or more output devices 160.
[0015] In some implementations, the computing device 100 may be implemented as various user terminals or service terminals with computing capabilities. The service terminal may be a server, a large computing device, etc. provided by various service providers. The user terminal may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablets, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that the computing device 100 can support any type of user interface (such as "wearable" circuits, etc.).
[0016] The processing unit 110 may be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 100. The processing unit 110 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0017] The computing device 100 generally includes multiple computer storage media. Such media may be any available media accessible to the computing device 100, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 120 may be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The memory 120 may include a conversion module 122, and these program modules are configured to perform the functions of various implementations described herein. The conversion module 122 may be accessed and run by the processing unit 110 to implement the corresponding functions.
[0018] The storage device 130 can be a removable or non-removable medium and can include a machine-readable medium that can be used to store information and / or data and can be accessed within the computing device 100. The computing device 100 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 1 , a disk drive for reading from or writing to a removable, non-volatile disk and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces.
[0019] The communication unit 140 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 100 can be implemented in a single computing cluster or multiple computer machines that can communicate via a communication connection. Thus, the computing device 100 can operate in a networked environment using a logical connection to one or more other servers, personal computers (PCs), or another general network node.
[0020] The input device 150 can be one or more of various input devices such as a mouse, keyboard, trackball, voice input device, etc. The output device 160 can be one or more output devices such as a display, speaker, printer, etc. The computing device 100 can also communicate with one or more external devices (not shown) as needed via the communication unit 140, external devices such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the computing device 100, or communicate with any device that enables the computing device 100 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0021] In some implementations, in addition to being integrated on a single device, some or all of the various components of computing device 100 may also be provided in the form of a cloud computing architecture. In a cloud computing architecture, these components may be remotely located and may work together to implement the functions described in this disclosure. In some implementations, cloud computing provides computing, software, data access, and storage services, which do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various implementations, cloud computing uses appropriate protocols to provide services over a wide area network, such as the Internet. For example, a cloud computing provider provides applications over a wide area network, and they can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture, as well as the corresponding data, may be stored on a server at a remote location. The computing resources in a cloud computing environment may be consolidated at a remote data center location or they may be distributed. The cloud computing infrastructure may provide services through a shared data center, even though they appear as a single access point for the user. Thus, the components and functions described herein may be provided from a service provider at a remote location using a cloud computing architecture. Alternatively, they may also be provided from a conventional server, or they may be installed directly or otherwise on a client device.
[0022] Computing device 100 may be used to implement image processing in various implementations of this disclosure. As Figure 1 shown, computing device 100 may receive a reference image 172 and an input image 171 to be transformed to a target resolution via an input device 150. Generally, the input image 171 has a resolution lower than the target resolution and may be referred to as a low-resolution image. The resolution of the reference image 172 may be equal to or greater than the target resolution. Computing device 100 may transform the input image 171 into an output image 180 having the target resolution based on the reference image 172. The output image 180 embodies the texture features of the reference image 172. For example, as Figure 1 shown, a region 181 of the output image 180 may embody the texture features migrated from the reference image 172. The output image 180 may also be referred to as a high-resolution image or a super-resolution (SR) image. In this document, "embody the texture features of a reference image" or similar expressions are to be interpreted as "embody at least part of the texture features of a reference image".
[0023] As mentioned above, the solutions for image super-resolution mainly include single-image super-resolution (SISR) techniques and reference-image-based super-resolution (RefSR) techniques. In traditional SISR, a deep convolutional neural network is generally trained to fit the restoration operation from a low-resolution image to a high-resolution image. This process involves a typical one-to-many problem because the same low-resolution image may correspond to multiple feasible high-resolution images. Therefore, the model trained in the traditional one-to-one fitting manner tends to fit the average of multiple feasible results, which results in the final generated high-resolution image still being relatively blurred at complex textures and it is difficult to obtain satisfactory results. In recent years, some methods have introduced generative adversarial training to constrain the restored image to be in the same distribution as the real high-resolution image. These methods can alleviate the above problems to a certain extent, but they also introduce the problem of the generation of some unrealistic textures.
[0024] To address the above problems, the RefSR technique has been proposed. In RefSR, a high-resolution image similar to the input image is used to assist the entire super-resolution restoration process. The introduction of the high-resolution reference image transforms the image super-resolution problem from texture restoration / generation to texture search and transfer, resulting in a significant improvement in the visual effect of the super-resolution results. However, the existing RefSR solutions are still difficult to achieve satisfactory image super-resolution results. For example, in some cases, there are inaccurate or even incorrect texture search and transfer.
[0025] Some of the problems in the current image super-resolution solutions are discussed above. According to the implementation of the present disclosure, a solution for image processing is provided, aiming to solve one or more of the above problems and other potential problems. In this solution, multi-level attention information is utilized to transfer the texture in the reference image to the input image to generate an output image with the target resolution. The multi-level attention information can indicate the correlation of the pixel blocks in the reference image and the pixel blocks in the input image according to the texture features. By using the multi-level attention information, the search and transfer of texture features can be made more accurate, thereby reducing texture blurring and texture distortion. In some implementations, feature integration across different resolution scales can also be achieved. In this implementation, by fusing the texture features of the reference image at different resolution scales, the feature expression ability is enhanced, which helps to generate more realistic images. Additionally, in some implementations, the texture extractor used to extract texture features from the reference image and the input image can be trained for the image super-resolution task. In this implementation, more accurate texture features can be extracted to further facilitate the generation of a clear and realistic output image.
[0026] The following further describes various example implementations of this solution in detail with reference to the accompanying drawings.
[0027] System architecture
[0028] Figure 2 shows an architectural diagram of a system 200 for image processing according to an implementation of the present disclosure. The system 200 can be implemented in Figure 1 a computing device 100. For example, in some implementations, the system 200 can be implemented as Figure 1 at least a part of the image processing module 122 of the computing device 100, that is, implemented as a computer program module. As Figure 2 shown, the system 200 generally can include a texture transfer subsystem 210 and a backbone subsystem 230. It should be understood that the structure and function of the system 200 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure. Implementations of the present disclosure can also be implemented in different structures and / or functions.
[0029] As Figure 2 shown, the input of the system 200 includes an input image 171 to be transformed to a target resolution and a reference image 172. The input image 171 is generally a low-resolution image with a resolution lower than the target resolution. The reference image 172 can be obtained in any suitable manner. In some implementations, the resolution of the reference image 172 can be equal to or greater than the target resolution. The output of the system 200 is an output image 180 that is transformed to the target resolution. Compared with the input image 171, the output image 180 is not only enlarged, but also embodies the texture features migrated from the reference image 172.
[0030] The texture transfer subsystem 210 can transfer the texture features of the reference image 172 to the input image 171. The texture transfer subsystem 210 at least includes a texture transformer 220-1, which can be configured to transfer the texture features of the reference image 172 to the input image 171 at the target resolution and output a synthetic feature map with the target resolution. As used herein, the term "synthetic feature map" refers to a feature map that includes the image features of the input image 171 itself and the texture features migrated from the reference image 172. Hereinafter, the texture transformer 220-1 can sometimes also be referred to as the texture transformer corresponding to the target resolution.
[0031] In some implementations, the texture transfer subsystem 210 can include a stacked texture transformer corresponding to multiple resolutions. As Figure 2As shown, in addition to the texture transformer 220-1, the texture migration subsystem 210 may further include at least one of the texture transformers 220-2 and 220-3. The texture transformers 220-2 and 220-3 may be configured to migrate the texture features of the reference image 172 to the input image 171 at a resolution lower than the target resolution, and output a synthetic feature map with the corresponding resolution. For example, the texture transformer 220-2 may correspond to a first resolution lower than the target resolution, and the texture transformer 220-3 may correspond to a second resolution lower than the first resolution.
[0032] The multiple resolutions corresponding to the stacked texture transformers may include the target resolution and any suitable resolution. In one implementation, the multiple resolutions corresponding to the stacked texture transformers may include the target resolution, the initial resolution of the input image 171, and one or more resolutions between the target resolution and the initial resolution. In such an implementation, the resolution of the input image may be increased step by step to further optimize the quality of the output image, as will be described below with reference to Figure 4 As an example, if the target resolution is to magnify the input image 171 by 4 times (4X), the texture transformer 220-2 may correspond to the resolution of magnifying the input image 171 by 2 times (2X), and the texture transformer 220-3 may correspond to the initial resolution of the input image 171. Hereinafter, the texture transformers 220-1, 220-2, and 220-3 may be collectively referred to as the texture transformer 220. The implementation of the texture transformer 220 will be described below with reference to Figure 3 to describe.
[0033] The backbone subsystem 230 may extract the image features of the input image 171, including but not limited to the color features, texture features, shape features, spatial relationship features, etc. of the input image 171. The backbone subsystem 230 may provide the extracted image features to the texture migration subsystem 210, so that the texture transformer 220 can generate a synthetic feature map. The backbone subsystem 230 may also process the synthetic feature map generated by the texture transformer 220 to obtain the output image 180. For example, the backbone subsystem 230 may include a convolutional neural network to extract image features and process the synthetic feature map. In some implementations, the convolutional neural network may include a residual network.
[0034] In an implementation where the texture migration subsystem 210 includes a stacked texture transformer, the backbone subsystem 230 may also include a cross-scale feature integration (CSFI) module (not shown). The CSFI module may be configured to exchange feature information between different resolution scales. For example, the CSFI module may transform the synthetic feature map output by the texture transformer 220-2 and provide it to the texture transformer 220-1. Additionally, the CSFI module may also transform the synthetic feature map output by the texture transformer 220-1 and further combine the transformed synthetic feature map with the synthetic feature map output by the texture transformer 220-2. Additionally, the CSFI module may also combine the synthetic feature maps output by different texture transformers. The operation of the CSFI module will be described below with reference to Figure 4 to describe the operation of the CSFI module.
[0035] Texture transformer
[0036] Figure 3 FIG. 300 is a block diagram of a texture transformer 220 according to some implementations of the present disclosure. In Figure 3 this example, the texture transformer 220 includes a texture extractor 301, a correlation embedding module 302, a hard attention module 303, and a soft attention module 304. The texture extractor 301 may be configured to extract the texture features of the input image 171 and the reference image 172. The correlation embedding module 302 may be configured to determine a first information 321 and a second information 322 based on the extracted texture features, which indicate the correlation of the pixel blocks of the input image 171 and the reference image 172 according to the texture features. The first information 321 may also be referred to as hard attention information and is exemplarily shown as a hard attention map in Figure 3 The second information 322 may also be referred to as soft attention information and is exemplarily shown as a soft attention map in Figure 3 The hard attention module 303 may be configured to determine a migration feature map 314 based on the first information 321 and the reference image 172. The feature blocks in the migration feature map 314 include the texture features of the pixel blocks in the reference image 172. As
[0037] shown in Figure 3 , the migration feature map 314 may be represented as T. The soft attention module 304 may be configured to generate a synthetic feature map 316 based on the second information 322, the migration feature map 314, and the feature map 315 from the backbone subsystem 230. The synthetic feature map 316 may be provided to the backbone subsystem 230 to be further processed into the output image 180. As Figure 3 shown in
[0038] The working principle of the texture transformer 220 will be described in detail below.
[0039] In the following, the original reference image 172 and the input image 171 may be represented as Ref and LR respectively. In the RefSR task, texture extraction for the reference image Ref is necessary. In some implementations, the original input image 171 and the reference image 172 may be applied to the texture extractor 301.
[0040] In some implementations, before extracting the texture features of the reference image 172 and the input image 171 using the texture extractor 301, the reference image 172 and the input image 171 may be preprocessed. As Figure 3 shown, in some implementations, the input image 171 may be upsampled by a predetermined factor to obtain a preprocessed input image 371, which may be represented as LR↑. For example, the preprocessed input image 371 may have a target resolution. The reference image 172 may be sequentially downsampled and upsampled by the predetermined factor to obtain a preprocessed reference image 372, which may be represented as Ref↓↑. As an example, 4X upsampling (e.g., bicubic interpolation) may be applied to the input image 171 to obtain the preprocessed input image 371. The reference image 172 may be sequentially applied with downsampling and upsampling with the same factor of 4X (e.g., bicubic interpolation) to obtain the preprocessed reference image 372.
[0041] In this implementation, the preprocessed input image 371 and the preprocessed reference image 372 may be in the same distribution. This can facilitate the texture extractor 301 to more accurately extract the texture features in the input image 171 and the reference image 172.
[0042] As Figure 3 shown, the preprocessed input image 371, the preprocessed reference image 372, and the reference image 172 may be applied to the texture extractor 301. The texture extractor 301 may perform texture feature extraction in any suitable manner.
[0043] For the extraction of texture features, the current main method is to input the image to be processed into a pre-trained classification model (e.g., the VGG network in computer vision) to extract some intermediate shallow features as the texture features of the image. However, this method has some drawbacks. First, the training objective of a classification model such as the VGG network is image category labels oriented by semantics, and there is a large difference between its high-level semantic information and low-level texture information. Second, for different tasks, the required texture information to be extracted is different, and using a pre-trained and fixed-weight VGG network lacks flexibility.
[0044] In view of this, in some implementations, instead of using a pre-trained classification model, a learnable texture extractor 301 can be implemented. A neural network (which can also be referred to as the first neural network herein) can be utilized to implement the texture extractor 301. For example, the neural network can be a shallow convolutional neural network, and during the training process of the texture transformer 220, the parameters of the neural network can also be continuously updated. In this way, the trained texture extractor 301 can be more suitable for texture feature extraction for the RefSR task, and thus can capture more accurate texture features from the reference image and the input image. Such a texture extractor can extract the texture information most suitable for the image generation task, thereby providing a good foundation for subsequent texture search and migration. This can further promote the generation of high-quality results. The training of the texture extractor 301 will be described below in conjunction with the design of the loss function.
[0045] As Figure 3 shown, the texture extractor 301 can generate feature maps 311, 312, 313 representing texture features. The feature maps 311, 312, 313 respectively correspond to the pre-processed input image 371, the pre-processed reference image 372, and the reference image 172. If LTE represents the texture extractor 301, and Q, K, and V respectively represent the feature maps 311, 312, and 313, then the process of extracting texture features can be expressed as:
[0046] Q = LTE(LR↑) (1)
[0047] K = LTE(Ref↓↑) (2)
[0048] V = LTE(Ref) (3)
[0049] where LTE(·) represents the output of the texture extractor 301. The feature map 311 (which can also be referred to as the first feature map herein) corresponds to the query Q, which represents the texture features extracted from the low-resolution input image (or the pre-processed input image) for texture search; the feature map 312 (which can also be referred to as the second feature map herein) corresponds to the key K, which represents the texture features extracted from the high-resolution reference image (or the pre-processed reference image) for texture search; the feature map 313 corresponds to the value V, which represents the texture features extracted from the original reference image for texture migration.
[0050] Each of the input image 171, the reference image 172, the preprocessed input image 371, and the preprocessed reference image 372 may include a plurality of pixel blocks. Each pixel block may include a set of pixels, and there may be overlapping pixels between different pixel blocks. The pixel block may be square (e.g., including 3x3 pixels), rectangular (e.g., including 3x6 pixels), or any other suitable shape. Accordingly, each of the feature maps 311-313 may include a plurality of feature blocks. Each feature block may correspond to a pixel block and include the texture features of the pixel block. For example, in Figure 3 the example, the feature map 311 represented as Q may include a plurality of feature blocks corresponding to a plurality of pixel blocks in the preprocessed input image 371. Similarly, the feature map 312 represented as K may include a plurality of feature blocks corresponding to a plurality of pixel blocks in the preprocessed reference image 372; the feature map 313 represented as V may include a plurality of feature blocks corresponding to a plurality of pixel blocks in the original reference image 172. The feature block may have the same size as the corresponding pixel block.
[0051] As Figure 3 shown, the feature map 311 and the feature map 312 may be used by the correlation embedding module 302. The correlation embedding module 302 may be configured to determine the correlation between the reference image 172 and the input image 171 in terms of texture features by estimating the similarity between the feature map 311 as the query Q and the feature map 312 as the key K. Specifically, the correlation embedding module 302 may extract feature blocks from the feature map 311 and the feature map 312 respectively, and then calculate the correlation between the feature blocks in the feature map 311 and the feature map 312 pairwise in the form of an inner product. The larger the inner product, the stronger the correlation between the corresponding two feature blocks. Accordingly, there are more transferable high-frequency texture features. On the contrary, the smaller the inner product, the weaker the correlation between the corresponding two feature blocks. Accordingly, there are fewer transferable high-frequency texture features.
[0052] In some implementations, a sliding window may be applied to the feature map 311 and the feature map 312 respectively to determine the feature blocks to be considered. The feature blocks from the feature map 311 (e.g., patch) may be denoted as q i (i ∈ [1, H LR ×W LR ), and the feature blocks from the feature map 312 (e.g., patch) may be denoted as k j (j ∈ [1, H Ref ×W Ref ), where H LR ×W LR represents the number of pixel blocks in the input image 171, and H Ref ×W RefIndicates the number of pixel blocks in the reference image 172. Then, for any pair of feature blocks q i and k j the correlation r i,j between them can be represented by the inner product of the normalized feature blocks and as:
[0053]
[0054] As an example, a sliding window having the same size as the pixel blocks in the input image 171 can be applied to the feature map 311 to locate the feature block q i in the feature map 311. For a given feature block q i , a sliding window having the same size as the pixel blocks in the reference image 172 can be applied to the feature map 312 to locate the feature block k j in the feature map 312. Then, the correlation between the feature blocks q i and k j can be calculated based on Equation (4). In this way, the correlation between any pair of feature blocks q i and k j can be determined.
[0055] The correlation embedding module 302 can further determine the first information 321 based on the correlation. The first information 321 can also be referred to as hard attention information and is exemplarily shown as the hard attention map H in Figure 3 . The first information 321 or the hard attention information can be used to transfer texture features from the reference image 172. For example, the first information 321 can be used to transfer the texture features of the reference image 172 from the feature map 313 as the value V. The traditional attention mechanism adopts the weighted sum of the features from V for each feature block. However, this operation causes a blurring effect and thus cannot transfer high-resolution texture features. Therefore, in the attention mechanism of the present disclosure, the local correlation according to pixel blocks is considered.
[0056] For example, the correlation embedding module 302 can generate a hard attention map H such as i,j shown in Figure 3 based on the correlation r i (i ∈ [1, H LR × W LR ) can be calculated from the following formula:
[0057]
[0058] In view of the correspondence between the pixel blocks and the feature blocks mentioned above, the i-th element h imay correspond to the i-th pixel block in the input image 171, and the value of h i may represent the index of the pixel block in the reference image 172 that is most relevant (i.e., has the highest degree of relevance) to the i-th pixel block according to the texture feature. For example, Figure 3 the value of the element 325 shown in is 3, which means that the pixel block in the upper left corner of the input image 171 is most relevant to the pixel block with index 3 in the reference image according to the texture feature.
[0059] Therefore, it can be understood that for the pixel blocks in the input image 171, what the first information 321 indicates is the position of the pixel block in the reference image 172 that is most relevant according to the texture feature. Additionally, although in the above example, the first information 321 indicates the most relevant pixel block in the reference image 172 for each pixel block in the input image 171, this is only illustrative. In some implementations, the first information 321 may indicate the most relevant pixel block in the reference image 172 for one or some pixel blocks in the input image 171.
[0060] The relevance embedding module 302 may further determine the second information 322 based on the degree of relevance. The second information 322 may also be referred to as soft attention information, and is exemplarily shown as the soft attention map S in Figure 3 . The second information 322 or the soft attention information may be used to weight the texture features from the reference image 172 for fusion into the input image 171.
[0061] For example, the relevance embedding module 302 may generate a soft attention map S such as i,j shown based on the degree of relevance r Figure 3 . The i-th element s i in the soft attention map S can be calculated from the following formula:
[0062]
[0063] It can be understood that the i-th element h i in the hard attention map H and the i-th element s i in the soft attention map S correspond. Given the correspondence between the pixel blocks and the feature blocks mentioned above, the i-th element s i in the soft attention map S may correspond to the i-th pixel block in the input image 171, and the value of s i may represent the degree of relevance between the i-th pixel block in the input image 171 and the pixel block with index h i in the reference image 172. For example, Figure 3The value of the element 326 shown in is 0.4, which indicates that the correlation between the pixel block in the upper left corner of the input image 171 and the pixel block indexed 3 (as indicated by the hard attention map H) in the reference image 172 is 0.4.
[0064] The hard attention module 303 can be configured to determine the migrated feature map 314 based on the first information 321 and the feature map 313 of the reference image 172. The migrated feature map 314 can be denoted as T. Each feature block of the migrated feature map 314 corresponds to a pixel block in the input image 171 and includes the texture features (e.g., especially high-frequency texture features) of the pixel block in the reference image 172 that is most relevant to that pixel block. The resolution of the migrated feature map 314 can be related to the resolution scale corresponding to the texture transformer 220. In the texture transformer 220-1 corresponding to the target resolution, the migrated feature map 314 has the target resolution. In an implementation where the texture migration subsystem 210 includes stacked texture transformers, in the texture transformer 220-2 corresponding to the first resolution, the migrated feature map 314 can have the first resolution; in the texture transformer 220-3 corresponding to the second resolution, the migrated feature map 314 can have the second resolution.
[0065] As an example, in order to obtain the migrated feature map 314 including the texture features migrated from the reference image 172, the hard attention information 321 can be used as an index to apply an index selection operation to the feature blocks of the feature map 313. For example, the value of the i-th feature block in the migrated feature map 314 denoted as T can be calculated from the following formula:
[0066]
[0067] where t i represents the value of the i-th feature block in T, which is the value of the feature block indexed h i in the feature map 313 denoted as V. It should be understood that the feature block indexed h i in the feature map 313 represents the features of the pixel block indexed h i in the reference image 172. Therefore, the i-th feature block in the migrated feature map 314 includes the features of the pixel block indexed h i in the reference image 172.
[0068] The above describes an example process for generating the migrated feature map 314 and the second information 322. As Figure 3As shown, the backbone subsystem 230 may provide the feature map 315 for the input image 171 to the texture transformer 220. The feature map 315, which may be denoted as F, includes at least the image features of the input image 171, including but not limited to color features, texture features, shape features, spatial relationship features, etc. In some implementations, the feature map 315 may include the image features of the input image 171. For example, the backbone subsystem 230 may utilize a trained neural network (which may also be referred to herein as the second neural network) to extract the image features of the original input image 171 and feed them to the texture transformer 220 as the feature map 315. In some implementations, in addition to the image features of the input image 171, the feature map 315 may also include texture features migrated from the reference image 172. For example, in an implementation where the texture migration subsystem 210 includes a stacked texture transformer, the feature map 315 may also include texture features migrated from the reference image 172 by another texture transformer at another resolution. For example, the synthesized feature map output by the texture transformer 220-2 may be provided to the texture transformer 220-1 as the feature map 315 after being transformed to the target resolution. This implementation will also be described below with reference to Figure 4 this.
[0069] In some implementations, the soft attention module 304 may directly apply the second information 322 (e.g., the shown soft attention map S) to the migrated feature map 314. Then, the migrated feature map 314 to which the second information 322 has been applied may be fused into the feature map 315 to generate the synthesized feature map 316.
[0070] In some implementations, as Figure 3 shown, the feature map 315 from the backbone subsystem 230 may be fed to the soft attention module 304. Instead of directly applying the second information 322 to the migrated feature map 314, the soft attention module 304 may apply the second information 322 to the combination of the migrated feature map 314 and the feature map 315 as the output of the soft attention module 304. The output of the soft attention module 304 may be further added to the feature map 315 from the backbone subsystem 230 to generate the synthesized feature map 316 as the output of the texture transformer 220. In this way, the image features of the input image 171 can be better utilized.
[0071] As an example, the above operation of generating the synthesized feature map 316 based on the feature map 315 (denoted as F), the migrated feature map 314 (denoted as T), and the soft attention map S may be expressed as:
[0072] F out = F + Conv(Concat(F, T)) ⊙ S (8)
[0073] where F outDenote the synthetic feature map 316. Conv and Concat denote the convolutional layer and the concatenate operation by channel respectively. X⊙Y denotes the element-wise multiplication between graph X and Y.
[0074] As can be seen from the above, the relevance indicated by the soft attention map S is applied as a weight to the combination of the image features and the transferred texture features. In this way, the texture features with strong relevance in the reference image 172 can be assigned relatively large weights and thus enhanced; meanwhile, the texture features with weak relevance in the reference image 172 can be assigned relatively small weights and thus suppressed. Accordingly, the finally obtained output image 180 will tend to embody the texture features with strong relevance in the reference image 172.
[0075] The above reference Figure 3 describes a texture transformer according to some implementations of the present disclosure. In such a texture transformer, a multi-level attention mechanism including hard attention and soft attention is utilized. Utilizing the hard attention mechanism can localize the transfer of texture features, and utilizing the soft attention mechanism can make the localized texture features be utilized more accurately and pertinently. Therefore, the texture transformer utilizing the multi-level attention mechanism can efficiently and accurately transfer relevant texture features from the reference image to the low-resolution input image. In some implementations, the learnable texture extractor 301 can extract the texture features of the reference image and the input image more accurately, thereby further promoting the accuracy of image super-resolution.
[0076] As Figure 3 shown, the synthetic feature map 316 generated by the texture transformer 220 can be provided to the backbone subsystem 230. In an implementation where the texture transfer subsystem 210 includes a texture transformer 220-1, the second neural network (e.g., a convolutional neural network) in the backbone subsystem 230 can process the synthetic feature map 316 into an output image 180. In an implementation where the texture transfer subsystem 210 includes a stacked texture transformer, the backbone subsystem 230 can fuse the synthetic feature maps generated by different texture transformers. An example of such an implementation will be described below with reference to Figure 4 to describe an example of such an implementation.
[0077] The present disclosure has been mainly described in the context of super-resolution tasks. However, it should be understood that the image processing scheme according to the present disclosure is not limited to super-resolution tasks. For example, in the case where the target resolution is equal to or lower than the original resolution of the input image, the texture transformer described above can also be used to transfer texture features from the reference image.
[0078] Cross-scale feature integration
[0079] As mentioned above, in an implementation where the texture transfer subsystem 210 includes a stacked texture transformer, the backbone subsystem 230 can implement cross-scale feature integration. Figure 4 FIG. 400 is a schematic block diagram showing cross-scale feature integration according to some implementations of the present disclosure. As Figure 4 shown, the texture transfer subsystem 210 may include stacked texture transformers 220-1, 220-2, and 220-3. The texture transformer 220-1 may correspond to the target resolution, i.e., the texture transformer 220-1 may transfer the texture features of the reference image 172 at the scale of the target resolution. Thus, the synthesized feature map 441 generated by the texture transformer 220-1 may have the target resolution. The texture transformer 220-2 may correspond to a first resolution lower than the target resolution, i.e., the texture transformer 220-2 may transfer the texture features of the reference image 172 at the scale of the first resolution. Thus, the synthesized feature map 421 generated by the texture transformer 220-2 may have the first resolution. The texture transformer 220-3 may correspond to a second resolution lower than the first resolution, i.e., the texture transformer 220-3 may transfer the texture features of the reference image 172 at the scale of the second resolution. Thus, the synthesized feature map 411 generated by the texture transformer 220-3 may have the second resolution.
[0080] In Figure 4 the example, the target resolution is intended to magnify the input image 171 by 4 times (4X), the first resolution is intended to magnify the input image 171 by 2 times (2X), and the second resolution is the original resolution (1X) of the input image 171.
[0081] As mentioned above, the backbone subsystem 230 may include a second neural network to extract the image features of the input image 171 and process the synthesized feature maps 316 generated by the texture transformers 220. In some implementations, the second neural network may include a residual network. The use of a residual network can improve the accuracy of image processing by increasing the network depth. A residual network generally may include a plurality of residual blocks (RBs). Figure 4 FIG. 450 shows a plurality of residual blocks included in the backbone subsystem 230.
[0082] As Figure 4As shown, in some implementations, the backbone subsystem 230 may further include a CSFI module 460, which is configured to exchange feature information at different resolution scales. The CSFI module 460 may fuse synthetic feature maps at different resolution scales. For each resolution scale, the CSFI module 460 may transform the synthetic feature maps at other resolution scales to this resolution (e.g., upsampling or downsampling), and then perform a concatenation operation on the transformed synthetic feature maps from other resolution scales and the synthetic feature map at this resolution scale in the channel dimension to obtain a new synthetic feature map at this resolution scale. Next, a convolutional layer (e.g., the residual block 450) may map this new synthetic feature map to the original number of channels.
[0083] The following describes an example process of feature integration with reference to Figure 4 . The backbone subsystem 230 (e.g., the second neural network of the backbone subsystem 230) may process the input image 171 having a second resolution (i.e., the original resolution of the input image 171) to obtain a feature map (not shown) including the image features of the input image 171. This feature map may be provided to the texture transformer 220-3 corresponding to the second resolution as the Figure 3 feature map F shown. The texture transformer 220-3 may then generate a synthetic feature map 411 having the second resolution. Thus, the synthetic feature map 411 may include the image features of the input image and the texture features migrated from the reference image 172 at the scale of the second resolution.
[0084] The residual block 450 may process the synthetic feature map 411 to obtain an updated synthetic feature map 412. This synthetic feature map 412 may be transformed to the first resolution, e.g., by upsampling or pixel shuffle. The transformed synthetic feature map 412 may be provided to the texture transformer 220-2 corresponding to the first resolution as the Figure 3 feature map F shown. The texture transformer 220-2 may then generate a synthetic feature map 421 having the first resolution. Thus, the synthetic feature map 421 may include the image features of the input image, the texture features migrated from the reference image 172 at the scale of the second resolution, and the texture features migrated from the reference image 172 at the scale of the first resolution.
[0085] The CSFI module 460 can downsample the synthesized feature map 421 with the first resolution to the second resolution, and combine (e.g., concatenate in the channel dimension) the downsampled synthesized feature map 421 with the synthesized feature map 412 with the second resolution into a new synthesized feature map 413. The residual block 450 can process the synthesized feature map 413, e.g., apply one or more convolution operations to the synthesized feature map 413 to obtain a synthesized feature map 414 mapped to the original number of channels.
[0086] Similarly, the CSFI module 460 can upsample the synthesized feature map 412 with the second resolution to the first resolution, and combine (e.g., concatenate in the channel dimension) the upsampled synthesized feature map 412 with the synthesized feature map 421 with the first resolution into a new synthesized feature map 422. The residual block 450 can process the synthesized feature map 422, e.g., apply one or more convolution operations to the synthesized feature map 422 to obtain a synthesized feature map 423 mapped to the original number of channels.
[0087] Similar to the synthesized feature map 412, the synthesized feature map 423 can be transformed to the target resolution, e.g., by upsampling or pixel rearrangement. The transformed synthesized feature map 423 can be provided to the texture transformer 220-1 corresponding to the target resolution as the Figure 3 feature map F as shown. The texture transformer 220-1 can then generate a synthesized feature map 441 with the target resolution. Thus, the synthesized feature map 441 can include the image features of the input image, the texture features migrated from the reference image 172 at the scale of the second resolution, the texture features migrated from the reference image 172 at the scale of the first resolution, and the texture features migrated from the reference image 172 at the scale of the target resolution.
[0088] As Figure 4 shown, the CSFI module 460 can combine (e.g., concatenate in the channel dimension) the synthesized feature map 441 with the target resolution, the synthesized feature map 423 upsampled to the target resolution, and the synthesized feature map 414 upsampled to the target resolution into a new synthesized feature map 442. The residual block 450 can process the synthesized feature map 442, e.g., apply one or more convolution operations to the synthesized feature map 442 to obtain a synthesized feature map 443 mapped to the original number of channels.
[0089] Similarly, the CSFI module 460 can combine (e.g., concatenate in the channel dimension) the synthesized feature map 423 with the first resolution, the synthesized feature map 414 upsampled to the first resolution, and the synthesized feature map 441 downsampled to the first resolution into a new synthesized feature map 424. The residual block 450 can process the synthesized feature map 424, e.g., apply one or more convolutional operations to the synthesized feature map 424 to obtain a synthesized feature map 425 mapped to the original number of channels.
[0090] Similarly, the CSFI module 460 can combine (e.g., concatenate in the channel dimension) the synthesized feature map 414 with the second resolution, the synthesized feature map 423 downsampled to the second resolution, and the synthesized feature map 441 downsampled to the second resolution into a new synthesized feature map 415. The residual block 450 can process the synthesized feature map 415, e.g., apply one or more convolutional operations to the synthesized feature map 415 to obtain a synthesized feature map 416 mapped to the original number of channels.
[0091] The backbone subsystem 230 can then combine the synthesized feature map 443 with the target resolution, the synthesized feature map 425 with the first resolution, and the synthesized feature map 416 with the second resolution to obtain the output image 180. For example, the synthesized feature map 443, the synthesized feature map 425 upsampled to the target resolution, and the synthesized feature map 416 upsampled to the target resolution can be concatenated by channel and then convolved to obtain the final output image 180.
[0092] In this implementation, the texture features migrated from the stacked texture transformers can be exchanged across different resolution scales. In this way, the reference image features of different granularities can be fused into different scales, thereby enhancing the feature expression ability of the network. Therefore, cross-scale feature integration can further improve the quality of the output image based on the multi-level attention mechanism. In addition, it should also be understood that cross-scale feature integration can achieve more powerful feature expression without significantly increasing the number of parameters and the amount of computation. For example, in some implementations, the attention information can be calculated once in one texture transformer and then shared among all texture transformers.
[0093] It should be understood that Figure 4 the multiple resolution scales shown are for illustration only and are not intended to limit the scope of the present disclosure. In addition, it should also be understood that the texture migration subsystem 210 may also include more or fewer (e.g., two) texture transformers.
[0094] Training loss function
[0095] The training of a system 200 for image processing according to an implementation of the present disclosure is described below. The input images and reference images used during the training process may be referred to as training input images and training reference images, respectively. The true high-resolution image corresponding to the training input image may be referred to as the ground truth image, and the image output by the system 200 during the training process may be referred to as the training output image.
[0096] The loss function for training the system 200 may include three parts, namely, a reconstruction loss function, an adversarial training loss function, and a perceptual loss function. For example, the total loss function may be expressed by the following formula:
[0097]
[0098] where and are the reconstruction loss function, the adversarial training loss function, and the perceptual loss function, respectively.
[0099] Reconstruction loss function can be expressed as:
[0100]
[0101] where I SR represents the result of super-resolution of the training input image, i.e., the training output image; I HR represents the ground truth image; (C, H, W) is the size of the ground truth image.
[0102] In this example, L1 is selected as the reconstruction loss function. Compared with L2, L1 can obtain a clearer output image. The reconstruction loss function ensures the identity between the output image and the input image by requiring the output image to be pixel-by-pixel consistent with the input image.
[0103] Generative adversarial networks have been proven to be effective in generating clear and visually realistic images. The adversarial training loss function for generative adversarial networks may include two parts represented by equations (11) and (12):
[0104]
[0105]
[0106] where G represents the network included in the system 200. D is a discriminator network introduced to train G, and its goal is to distinguish the output of network G from real images. Networks G and D can be alternately trained so that network D cannot distinguish the output of network G (i.e., the training output image) from real images. indicates is the output image of network G (i.e., the training output image) and the distribution of represents the probability that x is a real image determined by network D. Thus, the term included in both equations (11) and (12) represents the expectation that the training output image determined by network D is a real image. indicates that x is a real image and the distribution of x is D(x) represents the probability that x is a real image determined by network D. Thus, the term in equation (11) represents the expectation that x is a real image determined by network D. The term in equation (11) represents the gradient penalty, where the distribution is based on the distribution and is determined, indicates is an image that follows the distribution and represents the gradient of network D at .
[0107] The adversarial training loss function requires the output image to have the same distribution as the input image, so as to generate clearer and more realistic textures. Therefore, the adversarial training loss function can make the visual effect of the output image realistic enough.
[0108] The perceptual loss function is a special "reconstruction loss" imposed in the feature space of a specific pre-trained network. The perceptual loss function is beneficial to generating clearer and more realistic textures. In some implementations, the perceptual loss function may include the following two parts:
[0109]
[0110] The first part of the perceptual loss function shown in equation (13) is the traditional perceptual loss, where represents the feature map of the i-th layer of VGG19, and (C i , H i , W i ) represents the shape of the feature map of this layer. I SR represents the training output image; I HR represents the ground truth image.
[0111] The second part of the perceptual loss function shown in equation (13) is the transfer perceptual loss, which is designed for the texture feature extraction task. represents the texture feature map extracted from the j-th layer of the texture extractor (LTE) 301, and (C j , Hj , W j ) represents the shape of this layer. T represents Figure 3 the migrated feature map 314 shown in
[0112] Specifically, during the training process, the training reference image and the training input image can be applied to the texture extractor 301 to generate the migrated feature map T during training. The obtained training output image can be applied to the j-th layer of the texture extractor 301 to generate the texture feature map of the j-th layer Then, training can be carried out similarly to the conventional perceptual loss function.
[0113] This migrated perceptual loss function constrains the output image to have texture features similar to the migrated texture features. Therefore, the use of the migrated perceptual loss function can promote the more effective migration of the texture in the reference image.
[0114] It should be understood that in implementation, one or more of the reconstruction loss function, adversarial training loss function, and perceptual loss function described above can be used according to requirements. For example, in an implementation where it is desired to obtain an output image with better visual quality, the adversarial training loss function can be used. In an implementation where requirements are placed on the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) of the output image, the adversarial training loss function can be not used. In some implementations, only the reconstruction loss function can also be used.
[0115] Figure 5 A flowchart of a method 500 for image processing according to some implementations of the present disclosure is shown. The method 500 can be implemented by the computing device 100, for example, it can be implemented at the image processing module 122 in the memory 120 of the computing device 100.
[0116] As Figure 5 shown, at block 510, the computing device 100 determines first information and second information based on the texture features of the input image and the reference image. The first information at least indicates the second pixel block in the reference image that is most relevant to the first pixel block according to the texture features, and the second information at least indicates the degree of correlation between the first pixel block and the second pixel block. At block 520, the computing device 100 determines a migrated feature map with a target resolution based on the first information and the reference image. The migrated feature map includes a feature block corresponding to the first pixel block and the feature block includes the texture features of the second pixel block. At block 530, the computing device 100 transforms the input image into an output image with a target resolution based on the migrated feature map and the second information. The output image embodies the texture features of the reference image.
[0117] In some implementations, determining the first information and the second information includes: generating a first feature map representing the texture features of the input image and a second feature map representing the texture features of the reference image by applying the reference image and the input image to a trained first neural network; and determining the first information and the second information based on the first feature map and the second feature map.
[0118] In some implementations, determining the first information and the second information includes: determining the correlation between a first feature block corresponding to a first pixel block in the first feature map and multiple feature blocks in the second feature map; selecting a second feature block with the highest correlation with the first feature block from the multiple feature blocks; determining the first information based on the first pixel block and a second pixel block corresponding to the second feature block in the reference image; and determining the second information based on the correlation between the second feature block and the first feature block.
[0119] In some implementations, applying the reference image and the input image to the first neural network includes: preprocessing the input image by upsampling the input image by a predetermined factor; preprocessing the reference image by downsampling and then upsampling the reference image by a predetermined factor; and applying the preprocessed input image and the preprocessed reference image to the first neural network.
[0120] In some implementations, transforming the input image into an output image includes: obtaining a first synthetic feature map with a first resolution for the input image, the first resolution being lower than the target resolution, and the first synthetic feature map including the image features of the input image and the texture features migrated from the reference image; transforming the first synthetic feature map to the target resolution; generating a second synthetic feature map with the target resolution for the input image based on the second information, the migrated feature map, and the transformed first synthetic feature map; and determining the output image based on the second synthetic feature map.
[0121] In some implementations, determining the output image based on the second synthetic feature map includes: combining the second synthetic feature map and the transformed first synthetic feature map into a third synthetic feature map with the target resolution; transforming the second synthetic feature map to the first resolution; combining the first synthetic feature map and the transformed second synthetic feature map into a fourth synthetic feature map with the first resolution; and determining the output image based on the third synthetic feature map and the fourth synthetic feature map.
[0122] In some implementations, the input image has a second resolution lower than the first resolution, and obtaining the first synthesized feature map includes: extracting image features of the input image using a second neural network; generating a fifth synthesized feature map with the second resolution for the input image based on first information, second information, and the image features, where the fifth synthesized feature map includes texture features migrated from a reference image; transforming the fifth synthesized feature map to the first resolution; and determining the first synthesized feature map based on the transformed fifth synthesized feature map.
[0123] In some implementations, method 500 may further include: determining a training migration feature map based on a training reference image, a training input image, and a first neural network; generating a third feature map representing the texture features of a training output image by applying the training output image with a target resolution to the first neural network; and training the first neural network by minimizing the difference between the training migration feature map and the third feature map.
[0124] It can be seen from the above description that the image processing solution according to the implementations of the present disclosure can make the search and migration of texture features more accurate, thereby reducing texture blur and texture distortion. Additionally, the feature expression ability is enhanced through feature integration across different resolution scales to facilitate the generation of more realistic images.
[0125] Some example implementations of the present disclosure are listed below.
[0126] In one aspect, the present disclosure provides a computer-implemented method. The method includes: determining first information and second information based on the textures of an input image and a reference image, where the first information at least indicates a second pixel block in the reference image that is most relevant to a first pixel block in the input image according to texture features, and the second information at least indicates the degree of correlation between the first pixel block and the second pixel block; determining a migration feature map with a target resolution based on the first information and the reference image, where the migration feature map includes a feature block corresponding to the first pixel block and the feature block includes the texture features of the second pixel block; and transforming the input image into an output image with the target resolution based on the migration feature map and the second information, where the output image embodies the texture features of the reference image.
[0127] In some implementations, determining the first information and the second information includes: generating a first feature map representing the texture features of the input image and a second feature map representing the texture features of the reference image by applying the reference image and the input image to a trained first neural network; and determining the first information and the second information based on the first feature map and the second feature map.
[0128] In some implementations, determining the first information and the second information includes: determining a relevance between a first feature block corresponding to the first pixel block in the first feature map and a plurality of feature blocks in the second feature map; selecting, from the plurality of feature blocks, a second feature block having the highest relevance to the first feature block; determining the first information based on the first pixel block and a second pixel block corresponding to the second feature block in the reference image; and determining the second information based on the relevance between the second feature block and the first feature block.
[0129] In some implementations, applying the reference image and the input image to the first neural network includes: preprocessing the input image by upsampling the input image by a predetermined factor; preprocessing the reference image by sequentially downsampling and upsampling the reference image by the predetermined factor; and applying the preprocessed input image and the preprocessed reference image to the first neural network.
[0130] In some implementations, transforming the input image into the output image includes: obtaining a first synthetic feature map of the input image having a first resolution lower than the target resolution, the first synthetic feature map including image features of the input image and texture features migrated from the reference image; transforming the first synthetic feature map to the target resolution; generating, based on the second information, the migrated feature map, and the transformed first synthetic feature map, a second synthetic feature map of the input image having the target resolution; and determining the output image based on the second synthetic feature map.
[0131] In some implementations, determining the output image based on the second synthetic feature map includes: combining the second synthetic feature map and the transformed first synthetic feature map into a third synthetic feature map having the target resolution; transforming the second synthetic feature map to the first resolution; combining the first synthetic feature map and the transformed second synthetic feature map into a fourth synthetic feature map having the first resolution; and determining the output image based on the third synthetic feature map and the fourth synthetic feature map.
[0132] In some implementations, the input image has a second resolution lower than the first resolution, and obtaining the first composite feature map includes: extracting the image features of the input image by using a second neural network; generating a fifth composite feature map with the second resolution for the input image based on the first information, the second information, and the image features, where the fifth composite feature map includes texture features migrated from the reference image; transforming the fifth composite feature map to the first resolution; and determining the first composite feature map based on the transformed fifth composite feature map.
[0133] In some implementations, the second neural network includes a residual network.
[0134] In another aspect, the present disclosure provides an electronic device. The electronic device includes: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, which when executed by the processing unit cause the device to perform actions, the actions including: determining first information and second information based on an input image and a reference image, where the first information at least indicates a second pixel block in the reference image that is most relevant to a first pixel block in the input image according to texture features, and the second information at least indicates the degree of relevance between the first pixel block and the second pixel block; generating a migrated feature map with a target resolution based on the first information and the reference image, where the migrated feature map includes a feature block corresponding to the first pixel block and the feature block includes the texture features of the second pixel block; and transforming the input image into an output image with the target resolution based on the migrated feature map and the second information, where the output image reflects the texture features of the reference image.
[0135] In some implementations, determining the first information and the second information includes: generating a first feature map representing the texture features of the input image and a second feature map representing the texture features of the reference image by applying the reference image and the input image to a trained first neural network; and determining the first information and the second information based on the first feature map and the second feature map.
[0136] In some implementations, determining the first information and the second information includes: determining the degree of relevance between a first feature block corresponding to the first pixel block in the first feature map and multiple feature blocks in the second feature map; selecting a second feature block with the highest degree of relevance to the first feature block from the multiple feature blocks; determining the first information based on the first pixel block and the second pixel block corresponding to the second feature block in the reference image; and determining the second information based on the degree of relevance between the second feature block and the first feature block.
[0137] In some implementations, applying the reference image and the input image to the first neural network includes: preprocessing the input image by upsampling the input image by a predetermined factor; preprocessing the reference image by downsampling and then upsampling the reference image by the predetermined factor; and applying the preprocessed input image and the preprocessed reference image to the first neural network.
[0138] In some implementations, transforming the input image into the output image includes: obtaining a first composite feature map for the input image having a first resolution lower than the target resolution, the first composite feature map including the image features of the input image and texture features migrated from the reference image; transforming the first composite feature map to the target resolution; generating a second composite feature map for the input image having the target resolution based on the second information, the migrated feature map, and the transformed first composite feature map; and determining the output image based on the second composite feature map.
[0139] In some implementations, determining the output image based on the second composite feature map includes: combining the second composite feature map and the transformed first composite feature map into a third composite feature map having the target resolution; transforming the second composite feature map to the first resolution; combining the first composite feature map and the transformed second composite feature map into a fourth composite feature map having the first resolution; and determining the output image based on the third composite feature map and the fourth composite feature map.
[0140] In some implementations, the input image has a second resolution lower than the first resolution, and obtaining the first composite feature map includes: extracting the image features of the input image using a second neural network; generating a fifth composite feature map for the input image having the second resolution based on the first information, the second information, and the image features, the fifth composite feature map including texture features migrated from the reference image; transforming the fifth composite feature map to the first resolution; and determining the first composite feature map based on the transformed fifth composite feature map.
[0141] In some implementations, the method further includes determining a training migrated feature map based on a training reference image, a training input image, and the first neural network; generating a third feature map representing the texture features of the training output image by applying the training output image having the target resolution to the first neural network; and training the first neural network by minimizing the difference between the training migrated feature map and the third feature map.
[0142] In another aspect, the present disclosure provides a computer program product tangibly stored in a non-transitory computer storage medium and including machine-executable instructions that, when executed by a device, cause the device to perform the methods of the above aspects.
[0143] In another aspect, the present disclosure provides a computer-readable medium having machine-executable instructions stored thereon that, when executed by a device, cause the device to perform the methods of the above aspects.
[0144] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example, and not limitation, the types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0145] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0146] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0147] In addition, although the operations are depicted in a particular order, this should be understood as requiring that the operations be performed in the particular order shown or in sequential order, or requiring that all illustrated operations be performed to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features described in the context of separate implementations may also be implemented combinatorially in a single implementation. Conversely, the various features described in the context of a single implementation may also be implemented separately or in any suitable sub-combination in multiple implementations.
[0148] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A computer-implemented method, comprising: Determining first information and second information based on texture features of an input image and a reference image, wherein the first information at least indicates a second pixel block in the reference image that is most relevant to a first pixel block in the input image according to the texture features, and the second information at least indicates the degree of relevance between the first pixel block and the second pixel block; Determining a migration feature map with a target resolution based on the first information and the reference image, wherein the migration feature map includes a feature block corresponding to the first pixel block and the feature block includes the texture features of the second pixel block; And Transforming the input image into an output image with the target resolution based on the migration feature map and the second information, wherein the output image embodies the texture features of the reference image; Wherein transforming the input image into the output image includes: Obtaining a first composite feature map with a first resolution for the input image, the first resolution being lower than the target resolution, and the first composite feature map including image features of the input image and texture features migrated from the reference image; Transforming the first composite feature map to the target resolution; Generating a second composite feature map with the target resolution for the input image based on the second information, the migration feature map, and the transformed first composite feature map; and Determining the output image based on the second composite feature map.
2. The method according to claim 1, wherein determining the first information and the second information includes: Generating a first feature map representing the texture features of the input image and a second feature map representing the texture features of the reference image by applying the reference image and the input image to a trained first neural network; And Determining the first information and the second information based on the first feature map and the second feature map.
3. The method according to claim 2, wherein determining the first information and the second information includes: Determining the degree of relevance between a first feature block corresponding to the first pixel block in the first feature map and a plurality of feature blocks in the second feature map; Selecting a second feature block with the highest degree of relevance to the first feature block from the plurality of feature blocks; Determining the first information based on the first pixel block and the second pixel block in the reference image corresponding to the second feature block; and Determining the second information based on the degree of relevance between the second feature block and the first feature block.
4. The method according to claim 2, wherein applying the reference image and the input image to the first neural network includes: Preprocessing the input image by upsampling the input image by a predetermined factor; Preprocessing the reference image by sequentially downsampling and upsampling the reference image by the predetermined factor; And Applying the preprocessed input image and the preprocessed reference image to the first neural network.
5. The method according to claim 1, wherein determining the output image based on the second synthesized feature map includes: Combining the second synthesized feature map and the transformed first synthesized feature map into a third synthesized feature map having the target resolution; Transforming the second synthesized feature map to the first resolution; Combining the first synthesized feature map and the transformed second synthesized feature map into a fourth synthesized feature map having the first resolution; And Determining the output image based on the third synthesized feature map and the fourth synthesized feature map.
6. The method according to claim 1, wherein the input image has a second resolution lower than the first resolution, and obtaining the first synthesized feature map includes: Extracting the image features of the input image by using a second neural network; Generating a fifth synthesized feature map having the second resolution for the input image based on the first information, the second information, and the image features, the fifth synthesized feature map including texture features migrated from the reference image; Transforming the fifth synthesized feature map to the first resolution; And Determining the first synthesized feature map based on the transformed fifth synthesized feature map.
7. The method according to claim 2, further comprising: Determining a training migration feature map based on a training reference image, a training input image, and the first neural network; Generating a third feature map representing the texture features of the training output image by applying the training output image having the target resolution to the first neural network; And Training the first neural network by minimizing the difference between the training migration feature map and the third feature map.
8. An electronic device, comprising: A processing unit; And A memory coupled to the processing unit and containing instructions stored thereon, which when executed by the processing unit cause the device to perform actions, the actions including: Determining first information and second information based on the texture features of an input image and a reference image, the first information at least indicating a second pixel block in the reference image that is most relevant to a first pixel block in the input image according to the texture features, and the second information at least indicating the degree of relevance between the first pixel block and the second pixel block; Determining a migration feature map having a target resolution based on the first information and the reference image, the migration feature map including a feature block corresponding to the first pixel block and the feature block including the texture features of the second pixel block; and Transforming the input image into an output image having the target resolution based on the migration feature map and the second information, the output image embodying the texture features of the reference image; Wherein transforming the input image into the output image includes: Obtaining a first synthesized feature map having a first resolution for the input image, the first resolution being lower than the target resolution, and the first synthesized feature map including the image features of the input image and the texture features migrated from the reference image; Transforming the first synthesized feature map to the target resolution; Generate a second synthesized feature map with the target resolution for the input image based on the second information, the migration feature map, and the transformed first synthesized feature map; and Determine the output image based on the second synthesized feature map.
9. The apparatus according to claim 8, wherein determining the first information and the second information includes: Generating a first feature map representing the texture features of the input image and a second feature map representing the texture features of the reference image by applying the reference image and the input image to a trained first neural network; And Determining the first information and the second information based on the first feature map and the second feature map.
10. The apparatus according to claim 9, wherein determining the first information and the second information includes: Determining the correlation between the first feature block corresponding to the first pixel block in the first feature map and multiple feature blocks in the second feature map; Selecting a second feature block with the highest correlation with the first feature block from the multiple feature blocks; Determining the first information based on the first pixel block and the second pixel block corresponding to the second feature block in the reference image; and Determining the second information based on the correlation between the second feature block and the first feature block.
11. The apparatus according to claim 9, wherein applying the reference image and the input image to the first neural network includes: Preprocessing the input image by upsampling the input image by a predetermined factor; Preprocessing the reference image by sequentially downsampling and upsampling the reference image by the predetermined factor; And Applying the preprocessed input image and the preprocessed reference image to the first neural network.
12. The apparatus according to claim 8, wherein determining the output image based on the second synthesized feature map includes: Combining the second synthesized feature map and the transformed first synthesized feature map into a third synthesized feature map with the target resolution; Transforming the second synthesized feature map to the first resolution; Combining the first synthesized feature map and the transformed second synthesized feature map into a fourth synthesized feature map with the first resolution; And Determining the output image based on the third synthesized feature map and the fourth synthesized feature map.
13. The apparatus according to claim 8, wherein the input image has a second resolution lower than the first resolution, and obtaining the first synthesized feature map includes: Extracting the image features of the input image by using a second neural network; Generating a fifth synthesized feature map with the second resolution for the input image based on the first information, the second information, and the image features, the fifth synthesized feature map including texture features migrated from the reference image; Transforming the fifth synthesized feature map to the first resolution; And Determining the first synthesized feature map based on the transformed fifth synthesized feature map.
14. The device according to claim 9, wherein the action further comprises: determining a training transfer feature map based on a training reference image, a training input image, and the first neural network; generating a third feature map representing the texture features of the training output image by applying the training output image with the target resolution to the first neural network; and training the first neural network by minimizing the difference between the training transfer feature map and the third feature map.
15. A computer program product tangibly stored in a non-transitory computer storage medium and comprising machine-executable instructions that, when executed by a device, cause the device to perform actions, the actions comprising: determining first information and second information based on the texture features of an input image and a reference image, the first information at least indicating a second pixel block in the reference image that is most relevant to a first pixel block in the input image according to the texture features, and the second information at least indicating the degree of relevance between the first pixel block and the second pixel block; determining a transfer feature map with a target resolution based on the first information and the reference image, the transfer feature map including a feature block corresponding to the first pixel block and the feature block including the texture features of the second pixel block; and transforming the input image into an output image with the target resolution based on the transfer feature map and the second information, the output image embodying the texture features of the reference image; wherein transforming the input image into the output image comprises: obtaining a first composite feature map with a first resolution for the input image, the first resolution being lower than the target resolution, and the first composite feature map including the image features of the input image and the texture features migrated from the reference image; transforming the first composite feature map to the target resolution; generating a second composite feature map with the target resolution for the input image based on the second information, the transfer feature map, and the transformed first composite feature map; and determining the output image based on the second composite feature map.
16. The computer program product according to claim 15, wherein determining the first information and the second information comprises: generating a first feature map representing the texture features of the input image and a second feature map representing the texture features of the reference image by applying the reference image and the input image to a trained first neural network; and determining the first information and the second information based on the first feature map and the second feature map.
17. The computer program product according to claim 16, wherein applying the reference image and the input image to the first neural network comprises: preprocessing the input image by upsampling the input image by a predetermined factor; preprocessing the reference image by downsampling and then upsampling the reference image by the predetermined factor in sequence; and Apply the preprocessed input image and the preprocessed reference image to the first neural network.