Image Enhancement System and Image Enhancement Method Based on Non-Local Features

Through an end-to-end trainable guidance system, the feature extraction block, non-local feature generator and enhancement block are used to solve the problems of low computing efficiency and insufficient accuracy in the prior art, and low-cost efficient image enhancement is achieved, especially in portrait segmentation applications to improve processing effect.

CN115965543BActive Publication Date: 2025-08-05BLACK SESAME TECH (CHONGQING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211461327.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-12-08
Filing Date
2022-11-21
Publication Date
2025-08-05
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

The prior art has problems of low computational efficiency and inaccuracy in image enhancement, especially in super-resolution, denoising, multi-frame system image enhancement or video enhancement tasks. There is a lack of end-to-end trainable guidance systems or methods, making it difficult to effectively deal with difficult-to-process areas such as hair and hands.

Method used

Using an end-to-end trainable boot system, including feature extraction blocks, non-local feature generators, and non-local feature enhancement blocks, low-level image problems are handled through non-local feature concepts, non-local feature merging blocks are used to reveal the relationship between feature pixels, and features are reconstructed through conditional graphs to achieve enhancement.

Benefits of technology

It realizes efficient image enhancement effect at low computing cost, improves the accuracy and stability of image processing, especially in portrait segmentation applications, which reduces the learning space dimension and speeds up the convergence speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115965543B_ABST
    Figure CN115965543B_ABST
Patent Text Reader

Abstract

This invention discloses a system and method for image enhancement based on non-local features. The invention comprises an end-to-end trainable guidance system, including a feature extraction block, a non-local feature generator, and a non-local feature enhancement block. The system aims to leverage the concept of non-local features to address low-level image problems. The invention also employs a non-local feature merging block to correct for translational features, further improving the non-local features. Finally, the corrected features are reconstructed to form an enhanced image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to systems and methods for enhancing images, and more particularly, to image processing based on non-local features. Background Art

[0002] When processing images, similar pixels or features can be used to achieve denoising, deblurring, and super-resolution. However, finding closer similar features requires more computation. Traditional methods are inefficient and imprecise. Recently, deep learning models have demonstrated excellent performance in many image processing tasks that require pixel-by-pixel relationships, such as super-resolution, denoising, and multi-frame image and video enhancement.

[0003] Adobe's patent publication number 9,087,390 discloses a technique for upscaling an image sequence. Furthermore, the patent discloses generating an upsampled frame based on an original frame in a sequence of multiple frames. While the patent provides for image upscaling to introduce noise or amplify existing noise in an image, an end-to-end trainable guidance system or method is still lacking.

[0004] Another US patent application, filed under patent number 20190156210 and owned by Facebook, discloses techniques for image and video analysis using machine learning within a network environment, particularly relating to hardware and software for intelligent assistant systems. While this invention includes machine learning and is an improvement over previous patents, it still fails to achieve low-cost and accurate image enhancement.

[0005] Another Chinese patent application, patent publication number 109360156, filed by Shanghai Jiao Tong University, provides a single-image rain removal method or system based on image segmentation using a generative adversarial network. This invention addresses the aforementioned shortcomings of the prior art by providing a single-image rain removal method based on image segmentation using a generative adversarial network to address issues such as restoration of single images captured under various rainy conditions. However, this patent still lacks the ability to enhance multiple frames, as the system primarily focuses on removing rain lines from images.

[0006] The present invention is directed to a system and method for image enhancement. More specifically, the present invention performs image processing based on non-local features. Furthermore, to improve non-local performance and leverage the capabilities of deep learning networks, the present invention proposes an end-to-end trainable guidance system comprising a feature extraction block, a non-local feature generator, and a non-local feature enhancement block for processing low-level image problems using the concept of non-local features. This system can flexibly enhance images by creating non-local features for multi-frame or single-frame systems. Compared to other non-local methods based on deep learning, this system achieves enhanced images at a significantly lower computational cost.

[0007] Therefore, to overcome the shortcomings of existing technologies, such as image processing of difficult-to-process human body parts like hair and hands, a hierarchical hybrid loss is needed to replace the traditional segmentation loss. The hierarchical hybrid loss is reflected in the designed weights. Finally, to customize portrait segmentation applications and reduce the dimensionality of the learning space, a unique data augmentation strategy was devised to evenly distribute the training data, resulting in more stable performance and faster convergence. In view of the above inventions, the art needs a system to overcome or alleviate the above shortcomings of the existing technologies.

[0008] Currently, there are numerous methods and systems developed in the prior art that are suitable for various purposes. Furthermore, while these inventions may be suitable for the specific purposes for which they are intended, they are not suitable for the purposes of the present invention as described above. Therefore, there is a need for an advanced image enhancement system that performs image enhancement based on non-local features. Summary of the Invention

[0009] This paper proposes an end-to-end trainable guidance system, which includes a feature extraction block, a non-local feature generator and a non-local feature enhancement block, aiming to utilize the concept of non-local features to handle low-level image problems.

[0010] The image can be sent to the Feature Extraction Block (FEB) for feature extraction. The abstract feature set generates non-local features through the Non-local Feature Generator (NLFG). NLFG translates the features in nine directions through manually designed shifts to create non-local conditions. The Non-local Feature Enhancement Block (NLFEB) then uses these non-local features to perform image enhancement operations. The Non-Local Feature Merge Block (NLFMB) model is introduced in NLFEB to reveal the relationship between feature pixels. NLFMB can correct the translated features and further improve the non-local features. Finally, the next model can reconstruct the corrected features using a suitable condition map to achieve unique enhancement purposes.

[0011] The present invention also provides an image enhancement system for processing images based on non-local features. The image enhancement system includes a feature extraction module for receiving an image. The feature extraction module includes a processing unit and an extraction unit. The processing unit processes at least one frame of the image to generate multiple feature merging layers. The processing unit connects at least one of the multiple feature merging layers with a condition map to form one or more merged feature maps. The extraction unit extracts multiple feature extraction layers from the one or more merged feature maps, and extracts multiple features from the multiple feature extraction layers.

[0012] The non-local feature generator includes a shift unit and a padding unit. The shift unit applies nine different directional shifts to multiple features to form multiple feature translation layers. The padding unit fixes the shifts across the multiple feature translation layers by performing padding and cropping operations to form one or more translation feature maps.

[0013] The non-local feature enhancement module includes a merging unit, a reconstruction unit, and a connection unit. The merging unit merges one or more translation feature maps to form one or more non-local merged feature maps. The reconstruction unit constructs multiple reconstruction layers based on the one or more non-local merged feature maps. The connection unit connects the multiple reconstruction layers with the conditional map to form an enhanced image.

[0014] The primary objective of the present invention is to provide a system that can flexibly enhance images by creating non-local features for either multi-frame or single-frame systems. Compared to other non-local methods based on deep learning, this system achieves enhanced images at a significantly lower computational cost. The system provides a non-local feature generator for generating features that contain shifts between features extracted from feature extraction blocks. Furthermore, the system exploits non-local behavior by incorporating non-local features, reducing computational cost relative to other deep learning methods.

[0015] Another object of the present invention is to provide a non-local feature merging block (NLFMB) model introduced in the non-local feature enhancement module to reveal the relationship between feature pixels.

[0016] Another object of the present invention is to provide a non-local feature merging block to correct the translation features and further improve the non-local features, and to provide a reconstruction unit to reconstruct the corrected features.

[0017] Another object of the present invention is to provide a condition map, which may be a noise level map for denoising, or a sharpness weight for sharpening.

[0018] Another object of the present invention is to provide a non-local feature generator that creates nine sets of features in nine directions by translating features. In addition, the translation in nine directions is provided by using appropriate shifts.

[0019] Another object of the present invention is to provide a non-local feature enhancement block, including a deep learning block, to avoid large motions between features.

[0020] Another object of the present invention is to provide a deep learning block, which can be any of a deformable convolutional network, a self-attention mechanism, and a three-dimensional convolutional network. A DCN reveals relationships between features, distorting them for feature registration. A self-attention mechanism focuses on pixel relationships.

[0021] The objects and aspects of the present invention will become more apparent from the following detailed description taken in conjunction with the accompanying drawings, which illustrate, by way of example, the features of embodiments according to the present invention. To achieve the above and related objects, the present invention may be embodied in the form shown in the accompanying drawings, but attention should be paid to the fact that the drawings are illustrative only and that the specific structures illustrated and described may be varied within the scope of the appended claims.

[0022] Although the present invention has been described above through various exemplary embodiments and implementations, it should be understood that the applicability of the various features, aspects, and functions described in one or more individual embodiments is not limited to the specific embodiments described therein, but rather, may be applied alone or in various combinations to one or more other embodiments of the present invention, whether or not such embodiments are described, and whether or not such features are presented as part of such embodiments. Therefore, the breadth and scope of the present invention should not be limited by any of the exemplary embodiments described above. In some cases, the use of words and phrases such as "one or more," "at least," "but not limited to," or other similar phrases to expand the scope should not be understood to require the use of a narrower scope when such expanding phrases are not present. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The objects and features of the present invention will become more apparent from the following description and claims, taken in conjunction with the accompanying drawings. It should be understood that these drawings illustrate only typical embodiments of the present invention and, therefore, should not be construed as limiting the scope of the present invention. The present invention is described and explained with additional features and details by means of the following drawings.

[0024] Figure 1A An image enhancement system according to the present invention is shown.

[0025] Figure 1B is a schematic diagram of an image enhancement system according to the present invention.

[0026] Figure 2A The feature extraction module of the image enhancement system according to the present invention is shown.

[0027] Figure 2B It shows the single-frame feature extraction in the feature extraction module according to the present invention.

[0028] Figure 2C It shows multi-frame feature extraction in the feature extraction module according to the present invention.

[0029] Figure 3A A non-local feature generator of the feature enhancement system according to the present invention is shown.

[0030] Figure 3B Schematic diagram of a non-local feature generator according to the present invention.

[0031] Figure 3C The padding and cropping operations in the non-local feature generator according to the present invention are shown.

[0032] Figure 4A A non-local feature enhancement module of an image enhancement system is shown.

[0033] Figure 4BFIG. 4 is a schematic diagram showing a non-local feature enhancement module according to the present invention.

[0034] Figure 5A An image enhancement method according to the present invention is shown.

[0035] Figure 5B Another image enhancement method based on non-local features is shown. DETAILED DESCRIPTION

[0036] When processing images, similar pixels or similar features can be used to achieve denoising, deblurring, super-resolution, etc. However, if you want to find closer similar features, you need to perform more calculations. Especially in the traditional way, it is not only inefficient but also inaccurate. Recently, deep learning models have been able to achieve good performance in many image enhancement tasks such as super-resolution, denoising, and deblurring that require pixel relationships to solve problems. In order to improve non-local performance and utilize the capabilities of deep learning networks, the present invention proposes an end-to-end trainable guidance system, which includes a feature extraction block, a non-local feature generator, and a non-local feature enhancement block, which aims to use the concept of non-local features to process low-level image problems.

[0037] The image can be sent to the Feature Extraction Block (FEB) for feature extraction. The abstract feature set generates non-local features through the Non-local Feature Generator (NLFG). NLFG translates the features in nine directions through manually designed shifts to create non-local conditions. The Non-local Feature Enhancement Block (NLFEB) then uses these non-local features to perform image enhancement operations. The Non-Local Feature Merge Block (NLFMB) model is introduced in NLFEB to reveal the relationship between feature pixels. NLFMB can correct the translated features and further improve the non-local features. Finally, the next model can reconstruct the corrected features using a suitable conditional map to achieve unique enhancement purposes. The system will be described in detail below.

[0038] Figure 1AAn image enhancement system 100 according to the present invention is shown. The image enhancement system 100 is used to enhance an image and includes a feature extraction module 200, a non-local feature generator 300, and a non-local feature enhancement module 400. The feature extraction module 200 is used to receive an image and extract features to obtain details that are helpful for enhancement. The image enhancement system 100 regards a single frame or multiple frames of an image as input to facilitate image merging and feature extraction. The recommended network can include many popular blocks to be considered for image-level motion estimation. The image enhancement system 100 includes a deformable convolutional network instead of a traditional convolutional network, which overcomes the disadvantage that traditional convolutional networks cannot handle deformable objects and features.

[0039] The feature extraction module 200 also includes a processing unit and an extraction unit. The processing unit processes one or more image frames to generate one or more feature merging layers. The processing unit connects the one or more feature merging layers with a condition map to form one or more merged feature maps. The extraction unit extracts multiple feature extraction layers from the one or more merged feature maps. The extraction unit extracts multiple features from the multiple feature extraction layers.

[0040] The non-local feature generator 300 translates multiple features to form one or more translated feature maps. The NLFG creates nine sets of features in nine directions by translating the features after the feature extraction block. The translations in the nine directions should be appropriately offset. Large motion between translated features can be considered non-local behavior in the temporal dimension, as features in the same region between translated features may be similar to each other, as they originate from the original features sent by the feature extraction block.

[0041] To achieve this, several manually designed large shifts should be selected at the beginning of inference, such as 9, 15, 21, etc. Ultimately, the network does not need to perform additional computations to search for non-local pixels or non-local features, which means that denoising is more efficient.

[0042] The non-local feature enhancement module 400 merges one or more translation feature maps to form one or more non-local merged feature maps. The non-local feature enhancement module includes a reconstruction unit and a connection unit. The reconstruction unit constructs multiple reconstruction layers based on the one or more non-local merged feature maps. The connection unit connects the multiple reconstruction layers with the condition map to form an enhanced image.

[0043] By repeating the intrinsic pattern in nine directions, a non-local feature merging block (NLFMB) is obtained from the spatial and temporal dimensions. NLFMB also suppresses irrelevant information during processing, such as motion ghosting caused by the NLFG block, which produces large motion between features. Several deep learning blocks or networks are recommended to address this issue, including Deformable Convolutional Network V2, self-attention blocks, and Three Dimensional Convolutional Networks (3DCNNs).

[0044] After the non-local features are merged, the reconstruction model can be designed to have the same popular CNN model as the single-frame feature extraction part mentioned. Multiple types of conditional maps can be used with the merged features into the reconstruction model, depending on the task to be run. For example, if you want to use these blocks for noise reduction, you can make a conditional map based on the noise level coefficient. Alternatively, if the task is super-resolution, the conditional map can be a priori degradation kernel (e.g., a bicubic downsampling kernel) to guide model reconstruction.

[0045] Non-local features require a merging block to suppress irrelevant features, as nine directions of non-local features are generated. This system uses a channel attention block to determine the direction the non-local feature network prefers to maintain. To overcome ghosting and artifacts in certain areas caused by NLFG, a deformable convolution block is added after the attention block to extract useful information from the feature map. This network utilizes a non-local feature generator to achieve high image denoising quality while reducing computational cost.

[0046] Figure 1B An image enhancement system according to the present invention is shown. The present invention discloses a deep learning model design with a non-local feature block. An image can be sent as input 102 to a feature extraction block (FEB) for feature extraction. The abstract feature set is passed through a non-local feature generator (NLFG) to generate non-local features. The NLFG translates features in nine directions using manually designed shifts to create non-local conditions. The non-local feature enhancement block (NLFEB) then utilizes these non-local features to perform image enhancement operations.

[0047] The Non-Local Feature Merging Block (NLFMB) model is introduced in NLFEB to reveal the relationship between feature pixels. NLFMB can correct the translation features and further improve the non-local features at various stages of 24, h, w (104); 48, h / 2, w / 2 (106); and 96, h / 4, w / 4 (108) to create three-dimensional convolutional features (110). Finally, the next model can reconstruct the corrected features as output (112) using a suitable conditional map to achieve unique enhancement purposes.

[0048] The network is based on a standard U-Net with the following components. First, a non-local feature generator (NLFG) block is introduced to extract non-local features at each resolution level in the U-Net encoder. Creating non-local features in the encoder is chosen because the encoder can preserve more high-frequency details than the decoder. The decoder can then be responsible for the denoising task of extracting low-frequency regions specific to non-local features. As mentioned above, non-local features require a merging block to suppress irrelevant features, as nine directions of non-local features are created. The system uses a channel attention block to determine the direction in which the non-local feature network should maintain. To overcome ghosting and artifacts in certain areas caused by NLFG, a deformable convolution block is added after the attention block to extract useful information from the feature map. Leveraging the non-local feature generator, the network achieves high image denoising quality while reducing computational cost.

[0049] Figure 2A The feature extraction module 200 of the image enhancement system is shown. The feature extraction module 200 is configured to receive an image and includes a processing unit 202 and an extraction unit 204. The processing unit 202 processes at least one frame of the image to generate a plurality of feature merging layers. The processing unit 202 connects at least one feature merging layer from the plurality of feature merging layers with a condition map to form one or more merged feature maps.

[0050] The extraction unit 204 extracts multiple feature extraction layers from the one or more merged feature maps. The extraction unit 204 extracts multiple features from the multiple feature extraction layers.

[0051] The system captures details to aid image enhancement. It can be used in both multi-frame and single-frame conditions. The system manually selects the appropriate condition for use.

[0052] Multi-frame feature extraction: Increased sampling provides more information for image enhancement, so multi-frame systems often leverage this advantage for image enhancement tasks such as denoising, deblurring, and super-resolution. See the multi-frame feature extraction procedure. While multi-frame systems can generally extract more information than single-frame systems, relative motion between frames is always present, making this a crucial consideration when merging frames before feature extraction.

[0053] Single-frame feature extraction: The feature extraction model in multi-frame feature extraction (MFE) can be reused by the single-frame feature extraction pipeline. If the system takes a single frame as input, many popular CNN models or CNN blocks have the ability to efficiently extract features.

[0054] Figure 2BThe single-frame feature extraction block in the feature extraction module according to the present invention is shown. Single-frame feature extraction: The feature extraction model in the MFE can be reused by the single-frame feature extraction pipeline. If the system uses a single frame as input 206, many popular CNN models or CNN blocks have the ability to efficiently extract features. In addition to popular blocks such as ResNet, MobileNet, and NASNet, stacked residual blocks or back-projection models are good options for achieving a large receptive field.

[0055] Condition map 212 may be connected to certain layers in feature extraction layer 208 to form extracted features 210. Condition map 212 is additional information for reference, for example, a noise level map for denoising, or a sharpness weight for sharpening.

[0056] Figure 2C The multi-frame feature extraction block in the feature extraction module according to the present invention is shown. Multi-frame feature extraction: Increased sampling provides more information for image enhancement, so multi-frame systems often take advantage of this advantage in image enhancement tasks such as denoising, deblurring, and super-resolution. Referring to the multi-frame feature extraction procedure, it receives multi-frame input 214 and outputs extracted features 210, which can be similar to or different from single-frame feature extraction.

[0057] While multi-frame systems may generally extract more information than single-frame systems, relative motion between frames always exists, which is a key consideration when merging frames prior to feature extraction. In other words, if the system selects multiple frames as input 214, good motion estimation can aid image merging to form a feature merging layer 216, which in turn aids feature extraction by merging feature maps 218 to form a feature extraction layer 208. The proposed network can include many popular blocks to consider. For image-level motion estimation, utilizing a deformable convolutional network instead of a traditional convolutional network can overcome the shortcomings of traditional convolutional networks in their inability to handle deformable objects and features.

[0058] The condition map 212 may be connected to some specific layers as additional information for reference. For example, the condition map 212 may be a noise level map for denoising, or a sharpness weight for sharpening.

[0059] Figure 3A The non-local feature generator of the feature enhancement system is shown. The non-local feature generator 300 includes a shift unit 302 and a padding unit 304. The shift unit 302 applies nine different directional shifts to multiple features to form multiple feature translation layers. The padding unit 304 fixes the shifts on the multiple feature translation layers by performing padding and cropping operations to form one or more translation feature maps.

[0060] Figure 3BA non-local feature generator according to the present invention is shown. The non-local feature generator 300 can reduce computational costs by finding spatial similarities in the time dimension. The non-local feature generator (NLFG) 300 creates nine sets of features in nine directions by translating the features after the feature extraction block. The translation in the nine directions should have appropriate shifts to form a feature translation layer 308 based on the extracted features 306, and finally form a translation feature map 310. Large movements between translated features can be regarded as non-local behavior in the time dimension, because features in the same area between translated features may be similar to each other, and they come from the original features sent by the feature extraction block. To achieve this goal, several manually designed shifts should be selected at the beginning of inference, such as 9, 15, 21, etc.

[0061] Ultimately, the network does not need to perform additional computations to search for non-local pixels or non-local features, which means that denoising is more efficient.

[0062] Figure 3C The padding and cropping operations in the non-local feature generator according to the present invention are shown. If the manually designed shift is named k, the padding sizes of the four sides should be [k, k, k, k], and the upper left points of the patches in the nine directions should be [–k, -k], [-k, 0], [-k, k], [0, -k], [0, 0], [0, k], [k, -k], [k, 0], [k, k], thereby forming various cropping and padding patterns (312a-312d), and then forming a translation feature 314.

[0063] It is important to note that this block does not contain any trainable weights and only caches the translated feature data to automatically complete backpropagation in popular deep learning training architectures (such as Tensor Flow, PyTorch, etc.). This block can save computational costs compared to other non-local methods based on deep learning. Unlike this method, the non-local method is used for image processing. The non-local block requires three flattening operations and dot products on each dimension (height, width, channel), which means that the computational cost will increase significantly when the feature size increases slightly.

[0064] Figure 4A A non-local feature enhancement module for an image enhancement system is shown. The non-local feature enhancement module 400 includes a merging unit 402, a reconstruction unit 404, and a connection unit 406. The merging unit 402 merges one or more translation feature maps to form one or more non-local merged feature maps. The reconstruction unit 404 constructs multiple reconstruction layers based on the one or more non-local merged feature maps. The connection unit 406 connects the multiple reconstruction layers with the condition map to form an enhanced image.

[0065] After the non-local features are merged, the reconstruction model can be designed to have the same popular CNN model as the single-frame feature extraction part mentioned. Multiple types of conditional maps can be used with the merged features into the reconstruction model, depending on the task to be run. For example, if you want to use these blocks for noise reduction, you can make a conditional map based on the noise level coefficient. Alternatively, if the task is super-resolution, the conditional map can be a priori degradation kernel (e.g., a bicubic downsampling kernel) to guide model reconstruction.

[0066] Figure 4B The non-local feature enhancement module according to the present invention is shown. Although the translation between features may cause similar features to overlap with each other, resulting in non-local behavior, ghosting and artifacts may also be generated in certain areas when different features overlap. In order to overcome this shortcoming, the present invention proposes a non-local feature merging block (NLFMB), which receives a translation feature map 408 as input, and the NLFMB 410 merges the translation feature map 408 to form a merged feature map 412, and then forms a reconstruction layer 416. The conditional map 414 is applied to the merged feature map 412 to form the enhanced image 418. In this block, non-local features can be obtained from the spatial dimension and the temporal dimension by repeating the inherent pattern in 9 directions. At the same time, the NLFMB can also suppress irrelevant information during processing, such as the motion ghost caused by the NLFG block to produce large motion between features. It is recommended to use a deep learning block or a deep learning network to solve this problem. These deep learning blocks or deep learning networks include but are not limited to:

[0067] Deformable Convolutional Network (DCN) V2: DCN has a unique ability to reveal implicit relationships between features and can be used to warp features for feature registration, rather than using traditional pixel-level algorithms to estimate motion. The trained offsets serve as a flow map between features in feature space. Diversity within the same location of each feature ensures that non-local features are present after registration.

[0068] The trainable mask introduced in DCN V2 can suppress "bad" features caused by outliers in the trainable offset, especially in motion areas. In this application method, DCN V2 can enhance the acquisition of non-local features, which uses the trainable offset to find better non-local feature locations and warp them backward. At the same time, DCN V2 can also reduce the involvement of irrelevant features, such as large local motion.

[0069] Self-Attention Block: The self-attention mechanism has become popular in recent years. Unlike DCN, the self-attention block focuses on the pixel-wise relationship between two features without any distortion. Instead, it provides an attention weight map for each feature. This is more like a connection than a handoff to the DCN, helping the network find useful information. There are two types of attention mechanisms: spatial attention and temporal attention. In some cases, such as video enhancement, it is necessary to consider both dimensions simultaneously.

[0070] 3D Convolutional Network (3D-CNN): If self-attention has a temporal paradigm, 3D-CNN can also participate in feature merging operations. 3D-CNN can obtain more information from the sequence and find its inherent characteristics. In some cases, features can be combined in a third dimension as input to the model to extract temporal and spatial features from the sequence. By designing the network with a sufficiently large receptive field, it can fully cover the sequence, thereby outputting features that take into account the entire sequence information.

[0071] After the non-local features are merged, the reconstruction model can be designed to have the same popular CNN model as the single-frame feature extraction part mentioned. Multiple types of conditional maps can be used with the merged features into the reconstruction model, depending on the task to be run. For example, in order to make these blocks perform noise reduction, a conditional map can be made based on the noise level coefficient. Alternatively, if the task is super-resolution, the conditional map can be a priori degradation kernel (e.g., a bicubic downsampling kernel) to guide model reconstruction.

[0072] Figure 5A An image enhancement method 500A is illustrated. The method 500A includes: 502: a feature extraction module receives an image, processes at least one frame of the image, and generates a plurality of feature merging layers; 504: connecting at least one of the plurality of feature merging layers with a condition map to form one or more merged feature maps; and 506: extracting a plurality of feature extraction layers from the one or more merged feature maps to extract a plurality of features from the plurality of feature extraction layers.

[0073] 508: Shifting the plurality of features via the non-local feature generator to form one or more shifted feature maps. 510: Merging the one or more shifted feature maps to form one or more non-local merged feature maps. 512: Reconstructing a plurality of reconstruction layers based on the one or more non-local merged feature maps. 514: The non-local feature enhancement module concatenates the plurality of reconstruction layers with the conditional map to form an enhanced image.

[0074] Figure 5BA method 500B for image enhancement based on non-local features is shown. Method 500B includes: 516: a feature extraction module receives an image, processes at least one frame of the image, and generates multiple feature merging layers; 518: connecting at least one feature merging layer from the multiple feature merging layers to a condition map to form one or more merged feature maps; and 520: extracting multiple feature extraction layers from the one or more merged feature maps to extract multiple features from the multiple feature extraction layers.

[0075] 522: Generate nine shifts in different directions on multiple features to form multiple feature translation layers. 524: The non-local feature generator fixes the shifts on the multiple feature translation layers by performing padding and cropping operations to form one or more translation feature maps. 526: Merge the one or more translation feature maps to form one or more non-local merged feature maps. 528: Reconstruct multiple reconstruction layers based on the one or more non-local merged feature maps. 530: The non-local feature enhancement module connects the multiple reconstruction layers with the conditional map to form an enhanced image.

[0076] While various embodiments of the present invention have been described above, it should be understood that these embodiments are presented by way of example only and not limitation. Similarly, the accompanying drawings may depict example architectures or other configurations of the present invention to facilitate understanding of the features and functionality that may be included in the present invention. The present invention is not limited to the example architectures or configurations depicted in the accompanying drawings, and various alternative architectures and configurations may be utilized to achieve desired features.

[0077] Although the present invention has been described above through various exemplary embodiments and implementations, it should be understood that the applicability of the various features, aspects, and functions described in one or more individual embodiments is not limited to the specific embodiments described, but rather, may be applied alone or in various combinations to one or more other embodiments of the present invention, whether or not such embodiments are described, and whether or not such features are presented as part of such embodiments. Therefore, the breadth and scope of the present invention should not be limited by any of the above exemplary embodiments.

[0078] In some cases, the use of words and phrases such as "one or more," "at least," "but not limited to," or other similar phrases to broaden the scope should not be understood to require a narrower scope without such broadening phrases.

Claims

1. An image enhancement system for enhancing an image, comprising: A feature extraction module is configured to receive an image, wherein the feature extraction module includes: a processing unit, wherein the processing unit processes at least one frame of the image to generate a plurality of feature merging layers, and the processing unit further connects at least one feature merging layer of the plurality of feature merging layers with a condition map to form one or more merged feature maps; and an extraction unit, wherein the extraction unit extracts a plurality of feature extraction layers from the one or more merged feature maps, and further extracts a plurality of features from the plurality of feature extraction layers; a non-local feature generator, wherein the non-local feature generator translates the plurality of features to form one or more translated feature maps; and A non-local feature enhancement module, wherein the non-local feature enhancement module merges the one or more translation feature maps to form one or more non-local merged feature maps, wherein the non-local feature enhancement module includes: a reconstruction unit, wherein the reconstruction unit constructs a plurality of reconstruction layers based on the one or more non-local merged feature maps; and A connection unit, wherein the connection unit connects the plurality of reconstruction layers and the condition map to form an enhanced image.

2. The image enhancement system according to claim 1, wherein: The non-local feature merging block (NLFMB) model within the non-local feature enhancement module identifies the relationship between pixels to form the enhanced image.

3. The image enhancement system according to claim 2, wherein: The non-local feature merging block NLFMB checks and corrects the one or more translation feature maps before forming the one or more non-local merging feature maps.

4. The image enhancement system according to claim 1, wherein: The condition map is a noise level map used for denoising.

5. The image enhancement system according to claim 1, wherein: The condition map enhances the image by applying sharpness weights to the multiple feature merging layers to form the one or more merged feature maps.

6. The image enhancement system according to claim 1, wherein: The non-local feature generator checks spatial similarity through the temporal dimension to reduce computational cost.

7. The image enhancement system according to claim 1, wherein: The non-local feature generator generates nine different directions on the multiple features to form the multiple feature translation layers.

8. The image enhancement system according to claim 1, wherein: The non-local feature enhancement module includes a deep learning block to avoid large shifts between the multiple feature extraction layers.

9. The image enhancement system according to claim 8, wherein: The deep learning block is based on any one of a deformable convolutional network (DCN), a self-attention mechanism, and a three-dimensional convolutional network.

10. The image enhancement system according to claim 9, wherein: The DCN identifies a relationship between the one or more non-local merged feature maps for registration.

11. The image enhancement system of claim 9, wherein the self-attention mechanism examines pixel relationships of the one or more non-local pooled feature maps.

12. The image enhancement system according to claim 9, wherein: The self-attention mechanism includes spatial attention and temporal attention.

13. The image enhancement system according to claim 9, wherein: The three-dimensional convolutional network merges the one or more translational feature maps to identify at least one translational feature map to form the one or more non-local merged feature maps.

14. The image enhancement system according to claim 1, wherein: The non-local feature generator extracts the multiple features in the U-Net encoder.

15. An image enhancement system for processing images based on non-local features, wherein: The image enhancement system comprises: A feature extraction module is configured to receive an image, wherein the feature extraction module comprises: a processing unit, wherein the processing unit processes at least one frame of the image to generate a plurality of feature merging layers, and the processing unit further connects at least one feature merging layer of the plurality of feature merging layers with a condition map to form one or more merged feature maps; and an extraction unit, wherein the extraction unit extracts a plurality of feature extraction layers from the one or more merged feature maps, and further extracts a plurality of features from the plurality of feature extraction layers; and a non-local feature generator, wherein the non-local feature generator comprises: a shifting unit, wherein the shifting unit applies shifts in nine different directions to the plurality of features to form a plurality of feature translation layers; and a filling unit, wherein the filling unit fixes the shifts on the plurality of feature translation layers by performing padding and cropping operations to form one or more translation feature maps; A non-local feature enhancement module, wherein the non-local feature enhancement module includes: a merging unit, wherein the merging unit merges the one or more translation feature maps to form one or more non-local merged feature maps; a reconstruction unit, wherein the reconstruction unit constructs a plurality of reconstruction layers based on the one or more non-local merged feature maps; and A connection unit, wherein the connection unit connects the plurality of reconstruction layers and the condition map to form an enhanced image.

16. An image enhancement method, comprising: receiving an image, processing at least one frame of the image to generate a plurality of feature merging layers; Connecting at least one feature merging layer of the plurality of feature merging layers to a condition map to form one or more merged feature maps; extracting a plurality of feature extraction layers from the one or more merged feature maps to extract a plurality of features from the plurality of feature extraction layers; translating the plurality of features to form one or more translated feature maps; merging the one or more translated feature maps to form one or more non-local merged feature maps; reconstructing a plurality of reconstruction layers based on the one or more non-local merged feature maps; as well as The plurality of reconstruction layers are connected to the condition map to form an enhanced image.

17. A method for image enhancement based on non-local features, comprising: receiving an image, processing at least one frame of the image to generate a plurality of feature merging layers; Connecting at least one feature merging layer of the plurality of feature merging layers to a condition map to form one or more merged feature maps; extracting a plurality of feature extraction layers from the one or more merged feature maps to extract a plurality of features from the plurality of feature extraction layers; generating shifts in nine different directions on the plurality of features to form a plurality of feature translation layers; Fixing the shifts on the plurality of feature translation layers by performing padding and cropping operations to form one or more translation feature maps; merging the one or more translated feature maps to form one or more non-local merged feature maps; reconstructing a plurality of reconstruction layers based on the one or more non-local merged feature maps; as well as The plurality of reconstruction layers are connected to the condition map to form an enhanced image.

18. A computer usable medium having computer program logic for enabling at least one processor in a computer system to enhance an image via a software platform, the computer program logic comprising: receiving an image, processing at least one frame of the image to generate a plurality of feature merging layers; Connecting at least one feature merging layer of the plurality of feature merging layers to a condition map to form one or more merged feature maps; extracting a plurality of feature extraction layers from the one or more merged feature maps to extract a plurality of features from the plurality of feature extraction layers; translating the plurality of features to form one or more translated feature maps; merging the one or more translated feature maps to form one or more non-local merged feature maps; reconstructing a plurality of reconstruction layers based on the one or more non-local merged feature maps; as well as The plurality of reconstruction layers are connected to the condition map to form an enhanced image.

19. A computer usable medium having computer program logic for enabling at least one processor in a computer system to enhance an image based on non-local features via a software platform, the computer program logic comprising: receiving an image, processing at least one frame of the image to generate a plurality of feature merging layers; Connecting at least one feature merging layer of the plurality of feature merging layers to a condition map to form one or more merged feature maps; extracting a plurality of feature extraction layers from the one or more merged feature maps to extract a plurality of features from the plurality of feature extraction layers; generating shifts in nine different directions on the plurality of features to form a plurality of feature translation layers; Fixing the shifts on the plurality of feature translation layers by performing padding and cropping operations to form one or more translation feature maps; merging the one or more translated feature maps to form one or more non-local merged feature maps; reconstructing a plurality of reconstruction layers based on the one or more non-local merged feature maps; as well as The plurality of reconstruction layers are connected to the condition map to form an enhanced image.

Citation Information

Patent Citations

  • Contourlet information hiding method based on images

    CN103093464A

  • Image enhancement method based on residual self-attention and generative adversarial network

    CN112561838A