Blind image restoration method and system
By adopting a highly robust blind image repair network model in blind image repair, including a damaged area positioning subnet and a high-rootability repair subnet, the problems of inaccurate positioning of damaged areas and limited repair performance in the prior art are solved, and high-precision damaged area positioning and high-quality image repair are achieved.
Patent Information
- Application Number
- CN202510423729.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing blind image repair methods have inaccuracy in the positioning of damaged areas, resulting in limited repair performance, especially when the damage information is not a single constant value.
A highly robust blind image repair network model is adopted, including a damaged area positioning subnet and a highly robust repair subnet. The damaged area positioning subnet is positioned through multiple ring residual modules and global local feature analysis modules; the high-rootability repair subnet performs image repair through the adaptive feature aggregation module and the transposed convolution module.
Improves the accuracy of damaged area positioning and image repair performance, generates high-quality repair images, which can still be effective when the damaged information is complex.
Smart Images

Figure CN119941584A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer vision technology, and specifically relates to a blind image restoration method and system. Background Art
[0002] Image restoration aims to fill damaged images and restore them to their original state. The restored images must have reasonable semantics and natural textures. Most existing research on image restoration focuses on the situation where the damaged areas are known. However, in many practical applications, the damaged areas are unknown and random, i.e., blind image restoration, such as old photo restoration, cultural relic restoration, and medical image quality enhancement.
[0003] At present, the common problem faced by existing blind inpainting methods is that if the damaged area is not accurately located, such as missing or misdetecting damaged areas, it will directly affect the subsequent repair process, resulting in erroneous or missed repairs, etc., thus affecting the image restoration quality and having poor robustness. Since the position, size, and shape of the damaged area are all random and uncertain, existing methods still have limitations in achieving high-precision damaged area positioning, especially when the damage information in the damaged area is not a single constant value, which in turn affects the overall performance of the blind image inpainting method. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide a blind image repair method and system, which can solve the technical problems of inaccurate positioning of damaged areas and limited repair performance in the prior art.
[0005] In order to solve the above technical problems, this application is implemented as follows: In a first aspect, an embodiment of the present application provides a blind image restoration method, the method comprising: Acquire an image data set, and preprocess the image data set to obtain a damaged image data set; Constructing a highly robust blind image restoration network model, wherein the highly robust blind image restoration network model includes a damaged region positioning subnetwork and a highly robust restoration subnetwork; Training the highly robust blind image restoration network model according to part of the data in the damaged image dataset, and testing the trained highly robust blind image restoration network model according to another part of the data; The target damaged image is input into the tested high robustness blind image restoration network model for processing to obtain a target restored image of the target damaged image.
[0006] As an optional implementation of the first aspect of the present application, the target damaged image is input into the tested robust blind image restoration network model for processing to obtain a target restored image of the target damaged image; specifically: The target damaged image is input into the highly robust blind image restoration network model, and the damaged region positioning subnetwork performs positioning processing on the target damaged image to obtain a damaged region positioning image; The highly robust repair subnetwork repairs the target damaged image according to the damaged area positioning image to obtain the target repaired image.
[0007] As an optional implementation of the first aspect of the present application, the damaged area positioning subnetwork includes multiple annular residual modules and a global local feature analysis module, and the damaged area positioning subnetwork performs positioning processing on the target damaged image to obtain a damaged area positioning image; specifically: The target damaged image is input into the damaged area positioning sub-network, and the multiple annular residual modules process the target damaged image to obtain a rough positioning feature map; The global-local feature analysis module processes the rough positioning feature map and the target damaged image to obtain the damaged area positioning image.
[0008] As an optional implementation of the first aspect of the present application, the global-local feature analysis module includes a global feature extraction layer, a local feature extraction layer and a dynamic fusion attention layer; the global-local feature analysis module processes the coarse positioning feature map and the target damaged image to obtain the damaged area positioning image; specifically: The coarse positioning feature map is input into the global local feature analysis module, and channel segmentation is performed on the coarse positioning feature map to obtain a first segmentation feature and a second segmentation feature; The global feature extraction layer performs global feature extraction processing on the first segmentation feature to obtain a global feature, and the local feature extraction layer performs local feature extraction processing on the second segmentation feature to obtain a local feature; Performing element-wise addition on the global features and the local features to obtain global local features, and inputting the global local features into the dynamic fusion attention layer for processing to obtain a global local weight matrix, and solving a complementary matrix of the global local weight matrix; Performing element-wise multiplication processing on the global-local weight matrix and the global feature to obtain a global attention feature, and performing element-wise multiplication processing on the complementary matrix and the local feature to obtain a local attention feature; The global attention features, local attention features, global features and local features are added element-wise to obtain global local attention features, and the global local attention features, the coarse positioning feature map and the target damaged image are channel-connected and mapped to obtain the damaged area positioning image.
[0009] As an optional implementation manner of the first aspect of the present application, the global feature extraction layer performs global feature extraction processing on the first segmentation feature to obtain a global feature, specifically: Mapping the first segmentation features into a key matrix, a query matrix and a value matrix, wherein the key matrix is mapped by group convolution, and the value matrix is mapped by convolution; Performing channel connection on the key matrix and the query matrix to obtain a connection matrix, and performing two consecutive convolution processes on the connection matrix to obtain a connection attention weight; The connection attention weight is element-wise multiplied by the value matrix to obtain an attention value matrix, and the attention value matrix is element-wise added to the key matrix to obtain the global feature.
[0010] As an optional implementation manner of the first aspect of the present application, the local feature extraction layer performs local feature extraction processing on the second segmentation feature to obtain a local feature, specifically: Performing multiple stacked convolution processes on the second segmentation feature to obtain a deep feature, and performing convolution processes and activation processes on the deep feature in sequence to obtain a deep activation feature; Performing element-wise multiplication on the depth activation feature and the second segmentation feature to obtain a depth activation segmentation feature, and performing element-wise addition on the depth activation segmentation feature and the second segmentation feature to obtain the local feature; Among them, the process of stacked convolution processing is: convolution processing is performed on the features, and depth-separable convolution processing is performed on the features after the convolution processing, and then element-wise addition is performed on the features after the convolution processing and the features processed by the depth-separable convolution processing.
[0011] As an optional implementation of the first aspect of the present application, the global-local features are input into a dynamic fusion attention layer for processing to obtain a global-local weight matrix, specifically: Performing global average pooling, convolution processing, activation processing, and convolution processing on the global and local features in sequence to obtain a first weight matrix; Performing global average pooling and global maximum pooling on the global and local features respectively to obtain average pooling features and maximum pooling features; Perform channel connection on the average pooling feature and the maximum pooling feature to obtain a connection pooling feature, and perform convolution processing on the connection pooling feature to obtain a second weight matrix; Performing element-wise addition of the first weight matrix and the second weight matrix to obtain an added weight matrix, and performing channel connection of the added weight matrix and the global feature to obtain a concatenated weight matrix; The concatenated weight matrix is sequentially subjected to channel scrambling, group convolution and activation processing to obtain the global-local weight matrix.
[0012] As an optional implementation of the first aspect of the present application, the highly robust restoration subnetwork includes a plurality of different convolution modules, a plurality of different transposed convolution modules, a plurality of identical residual modules, a transposed convolution output module and an adaptive feature fusion module; the number of the convolution modules is one more than the number of the transposed convolution modules; The highly robust repair sub-network repairs the target damaged image according to the damaged area positioning image to obtain the target repaired image; specifically: The plurality of convolution modules process the target damaged image step by step according to the damaged area positioning image to obtain a plurality of convolution features of different scales; The plurality of residual modules process the convolution features output by the last convolution module step by step, and output the residual convolution features through the last residual module; The adaptive feature aggregation module processes the multiple convolutional features respectively through multiple re-encoding layers to obtain multiple enhanced coding features, and performs channel connection on the multiple enhanced coding features to obtain adaptive fusion features; The adaptive fusion feature is transferred to all transposed convolution modules through skip connections, and the residual convolution feature and the adaptive fusion feature are processed by the first transposed convolution module to obtain the transposed feature; All transposed convolution modules except the first transposed convolution module process the transposed features step by step according to the adaptive fusion features to obtain target features; The transposed convolution output module processes the target features to obtain the target repaired image.
[0013] As an optional implementation of the first aspect of the present application, the re-encoding layer processes the convolutional features to obtain the enhanced coding features; specifically: Performing horizontal average pooling and vertical average pooling on the convolutional features to obtain horizontal pooling features and vertical pooling features, respectively, and performing channel connection on the horizontal pooling features and the vertical pooling features to obtain a connection average pooling feature; The connected average pooling feature is sequentially subjected to convolution processing, batch normalization processing, and nonlinear activation processing to obtain an activated average pooling feature; The activated average pooling feature is split along the horizontal axis and the vertical axis to obtain a horizontal axis activated pooling feature and a vertical axis activated pooling feature, and the horizontal axis activated pooling feature and the vertical axis activated pooling feature are sequentially subjected to convolution processing and activation processing to obtain a horizontal axis weight and a vertical axis weight; The horizontal axis weight, the vertical axis weight and the convolution feature are element-wise multiplied to obtain a weighted convolution feature, the weighted convolution feature is subjected to a plurality of convolution processes with different dilation rates to obtain a plurality of enhanced coding sub-features with different receptive fields, and a plurality of different enhanced coding sub-features are channel-connected to obtain the enhanced coding feature.
[0014] In a second aspect, an embodiment of the present application provides a blind image restoration system, the system comprising: Acquisition module: acquires an image data set, preprocesses the image data set, and obtains a damaged image data set; Construction module: constructing a highly robust blind image restoration network model, wherein the highly robust blind image restoration network model includes a damaged region positioning subnetwork and a highly robust restoration subnetwork; Training and testing module: training the highly robust blind image restoration network model according to part of the data in the damaged image data set, and testing the trained highly robust blind image restoration network model according to another part of the data; Prediction module: input the target damaged image into the tested high robustness blind image restoration network model for processing to obtain a target restoration image of the target damaged image.
[0015] In the embodiments of the present application, compared with the prior art, the following beneficial effects are achieved: (1) The damaged area localization subnetwork analyzes the deep semantic information of the target damaged image and captures the semantic inconsistency between the damaged area and the known area to locate the damaged area. The global-local feature analysis module analyzes the semantic information of the target damaged image in detail from two different levels, namely the global and local levels, to effectively improve the positioning accuracy of the target damaged area. Even when the damage information is not a single constant value or is similar to the known information, the damaged area can still be accurately located. (2) The highly robust inpainting sub-network enhances multiple convolutional features through multiple re-encoding layers in the adaptive feature aggregation module. The enhanced convolutional features are then aggregated together through channel connections and passed to the decoding stage, thereby effectively improving the image inpainting performance and generating high-quality inpainted images. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flow chart of a blind image restoration method provided by some embodiments of the present application; Figure 2 is a structural diagram of a highly robust blind image restoration network model of a blind image restoration method provided by some embodiments of the present application; Figure 3is a structural diagram of a global-local feature analysis module of a blind image restoration method provided by some embodiments of the present application; Figure 4 is a structural diagram of a global feature extraction layer of a blind image restoration method provided by some embodiments of the present application; Figure 5 is a structural diagram of a local feature extraction layer of a blind image restoration method provided by some embodiments of the present application; Figure 6 It is a structural diagram of a dynamic fusion attention layer of a blind image restoration method provided by some embodiments of the present application; Figure 7 It is a structural diagram of a re-encoding layer of a blind image restoration method provided by some embodiments of the present application. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0018] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.
[0019] In the following, in conjunction with the accompanying drawings, a blind image restoration method and system provided by the embodiment of the present application are described in detail through specific embodiments and their application scenarios.
[0020] Example A blind image restoration method comprises the following steps: S100: Acquire an image data set, and preprocess the image data set to obtain a damaged image data set; Furthermore, the face dataset CelebAMask-HQ and the street view dataset Cityscapes are used as the true image datasets, and the face dataset FFHQ and the street view dataset Paris StreetView with the same attributes are used as damage information. The corresponding damaged images are obtained by following the damage rules, thereby constructing a damaged image dataset.
[0021] Specifically, the damage rule is expressed by the following formula: , in, represents the true value image, represents a binary matrix, Used to indicate the damaged area (where 1 indicates the damaged area and 0 indicates the known area), It is a known condition only during training. From the irregular damaged area dataset proposed by Liu et al., Indicates damage information. represents element-wise multiplication, Indicates a corrupted image.
[0022] S200: Construct a highly robust blind image restoration network model, which includes a damaged area positioning sub-network and a highly robust restoration sub-network; It should be noted that the damaged area positioning subnetwork includes multiple annular residual modules and a global-local feature analysis module; the global-local feature analysis module includes a global feature extraction layer, a local feature extraction layer and a dynamic fusion attention layer; the high robustness repair subnetwork includes multiple different convolution modules, multiple different transposed convolution modules, multiple identical residual modules, a transposed convolution output module and an adaptive feature fusion module; the number of convolution modules is one more than the number of transposed convolution modules; the adaptive feature aggregation module includes multiple recoding layers, and the number of recoding layers is the same as the number of convolution modules; in this embodiment, the number of annular residual modules is 9, and they are symmetrically arranged according to the number of channels; the number of convolution modules is 4, the number of recoding layers is 4, the number of transposed convolution modules is 3, and the number of channels of the recoding layer corresponds one to one to the number of channels of the convolution module.
[0023] S300: training a highly robust blind image restoration network model according to part of the data in the damaged image dataset, and testing the trained highly robust blind image restoration network model according to another part of the data; It should be noted that there are 30,000 images in the CelebAMask-HQ dataset, of which 27,176 images are used as training images and 2,824 images are used as test images. In the FFHQ dataset with the same attributes, the last 2,824 images are selected as test damage information, and the remaining 67,176 images are used as the source of loss information during the training process. For the Cityscapes dataset, this embodiment selects a version of 5,000 high-quality pixel-level annotation images, of which 4,500 are training images and 500 are test images. During the training process, the last 500 images of the original training images in the Paris StreetView dataset are used as the source of test damage information for the Cityscapes dataset, and the other images in the training set and the original test set images, a total of 14,500 images, are used as the source of damage information for the training set. For the irregular damaged area dataset proposed by Liu et al., 12,000 irregular damaged area images are selected. The damaged area dataset is divided according to the damage ratio to obtain six sub-damaged area datasets with damage ratios of 0-10%, 10-20%, 20-30%, 30-40%, 40-50%, and 50-60%. Each sub-damaged area dataset has 2,000 damaged area images, of which the first 1,750 are selected as training images and the remaining 250 are selected as test images.
[0024] S400: Input the target damaged image into the tested high robustness blind image restoration network model for processing to obtain a target restoration image of the target damaged image.
[0025] It should be noted that S400 is specifically: S410: inputting the target damaged image into the highly robust blind image restoration network model, and the damaged region positioning sub-network performs positioning processing on the target damaged image to obtain a damaged region positioning image; S420: The highly robust repair sub-network repairs the target damaged image according to the damaged area positioning image to obtain a target repaired image.
[0026] Furthermore, the deep semantic information in the target loss image is first analyzed through the damage region localization network (DRLN), and the semantically inconsistent regions are located in the damage region localization image. The damage region localization image is then input into the highly robust image inpainting network (HRN), which combines the damage region localization image with the target loss image to repair the target loss image and obtain the target repaired image.
[0027] It should be noted that the damaged area positioning subnetwork in S410 performs positioning processing on the target damaged image to obtain a damaged area positioning image, specifically: S411: the target damaged image is input into the damaged area positioning sub-network, and multiple annular residual modules process the target damaged image to obtain a rough positioning feature map; S412: The global-local feature analysis module processes the rough positioning feature map and the target damaged image to obtain a damaged area positioning image.
[0028] Furthermore, first, the damage region localization network (DRLN) amplifies the differences between different semantic contents in the target damaged image through 9 annular residual modules, roughly locates the semantically inconsistent areas, and obtains a rough positioning feature map; then, according to the global-local feature analysis module (GLFA), relevant features are extracted from both global and local perspectives to analyze the deep semantic feature information of the image in detail, accurately capture the semantic inconsistency between the damaged information and the known information, and thus effectively improve the positioning accuracy of the damaged area.
[0029] It should be noted that S412 is specifically: S4121: The coarse positioning feature map is input into the global local feature analysis module, and the coarse positioning feature map is segmented into channels to obtain a first segmentation feature and a second segmentation feature; S4122: the global feature extraction layer performs global feature extraction processing on the first segmentation feature to obtain a global feature, and the local feature extraction layer performs local feature extraction processing on the second segmentation feature to obtain a local feature; S4123: performing element-wise addition of global features and local features to obtain global local features, and inputting the global local features into the dynamic fusion attention layer for processing to obtain a global local weight matrix, and solving the complementary matrix of the global local weight matrix; S4124: performing element-wise multiplication of the global-local weight matrix and the global feature to obtain a global attention feature, and performing element-wise multiplication of the complementary matrix and the local feature to obtain a local attention feature; S4125: perform element-wise addition of the global attention features, local attention features, global features and local features to obtain the global local attention features, and perform channel connection and mapping processing on the global local attention features, the coarse positioning feature map and the target damaged image to obtain the damaged area positioning image.
[0030] Specifically, the global local feature analysis module is expressed by the following formula: , in, represents the global local attention feature, Represents global features, Represents local features, represents the global-local weight matrix, represents the complementary matrix of the global-local weight matrix, represents element-wise addition, Represents element-wise multiplication.
[0031] Furthermore, the global-local feature analysis module takes the output features (coarse positioning feature map) of the last annular residual module in the damaged area positioning subnetwork as input and outputs the final damaged area positioning image. First, the coarse positioning feature map is subjected to channel segmentation to obtain the first segmentation feature and the second segmentation feature; these two segmentation features are subjected to global feature extraction and local feature extraction operations respectively to generate corresponding global features and local features. Subsequently, the global features and local features are deeply analyzed through the designed dynamically fused attention layer (Dynamically Fused Attention, DFA), and the analysis results are fused to capture the inconsistency between deep semantics from two different perspectives, global and local, and output the analysis result features (global-local attention features); finally, the target damaged image, coarse positioning feature map and global-local attention features are combined to map the damaged area positioning image.
[0032] It should be noted that the global feature extraction layer in S4122 performs global feature extraction processing on the first segmentation feature to obtain a global feature, specifically: S41221: Mapping the first segmentation feature into a key matrix, a query matrix, and a value matrix, wherein the key matrix is obtained by mapping through group convolution, and the value matrix is obtained by mapping through convolution; S41222: performing channel connection on the key matrix and the query matrix to obtain a connection matrix, and performing two consecutive convolution processes on the connection matrix to obtain a connection attention weight; S41223: Multiply the connection attention weights and the value matrix element-wise to obtain the attention value matrix, and add the attention value matrix and the key matrix element-wise to obtain the global features.
[0033] Furthermore, in order to focus on important features and ignore noise, and enhance the understanding of global information; this embodiment maps the first segmentation feature into a key matrix, a query matrix and a value matrix, the key matrix is obtained by mapping through a 3×3 group of convolutions, and the value matrix is obtained by mapping through a 1×1 convolution; the key matrix and the query matrix are channel-connected to obtain a connection matrix, and the connection matrix is convolved twice in a row to obtain a connection attention weight; the connection attention weight is element-wise multiplied with the value matrix to obtain an attention value matrix, and the attention value matrix is element-wise added to the key matrix to obtain a global feature. The goal of global feature extraction is to extract information such as the overall semantic structure in the first segmentation feature. This information can provide the necessary global context support for determining whether the current pixel belongs to a damaged area, and helps to avoid problems such as misjudgment and wrong judgment caused by relying solely on local information.
[0034] Specifically, the global feature extraction layer is expressed by the following formula: , in, Represents global features, represents the key matrix, represents the value matrix, represents the query matrix, Indicates channel connection, represents 1×1 convolution processing, Represents 3×3 groups of convolution processing, represents the first segmentation feature, represents element-wise addition, Represents element-wise multiplication.
[0035] It should be noted that the local feature extraction layer in S4122 performs local feature extraction processing on the second segmentation feature to obtain local features, specifically: S41224: performing multiple stacked convolution processes on the second segmentation feature to obtain a deep feature, and performing convolution processes and activation processes on the deep feature in sequence to obtain a deep activation feature; S41225: performing element-wise multiplication on the depth activation feature and the second segmentation feature to obtain the depth activation segmentation feature, and performing element-wise addition on the depth activation segmentation feature and the second segmentation feature to obtain the local feature; Among them, the process of single stacked convolution processing is: convolution processing is performed on the features, and depth-separable convolution processing is performed on the features after convolution processing, and then element-wise addition is performed on the features after convolution processing and the features processed by depth-separable convolution processing.
[0036] Further, the second segmentation feature is subjected to multiple stacked convolution processes to obtain a depth feature. The process of a single stacked convolution process is: perform 1×1 convolution on the feature, perform 3×3 depth-separable convolution on the convolution-processed feature, and then perform element-wise addition on the feature after 1×1 convolution and the feature after 3×3 depth-separable convolution. In this embodiment, the second segmentation feature is subjected to the stacked convolution process three times to obtain a depth feature. After that, the depth feature is subjected to 1×1 convolution, and the depth feature after 1×1 convolution is activated by a sigmoid function to obtain a depth activation feature. The depth activation feature and the second segmentation feature are element-wise multiplied to obtain a depth activation segmentation feature, and the depth activation segmentation feature and the second segmentation feature are element-wise added to obtain a local feature. Local feature extraction focuses on the local contextual area information in the second segmentation feature, especially those subtle differences that may represent damaged areas, which helps to make up for the perception limitations of the global feature extraction branch, enhance the discrimination of complex damage patterns, and effectively improve the network's ability to locate damaged areas.
[0037] Specifically, the local feature extraction layer is expressed by the following formula: , in, Represents local features, represents the second segmentation feature, represents 1×1 convolution processing, represents stacked convolution processing, represents sigmoid activation, represents element-wise multiplication, Represents element-wise addition.
[0038] It should be noted that in S4123, the global and local features are input into the dynamic fusion attention layer for processing to obtain the global and local weight matrix, which is specifically: S41231: performing global average pooling, convolution processing, activation processing, and convolution processing on the global and local features in sequence to obtain a first weight matrix; S41232: performing global average pooling and global maximum pooling on the global and local features respectively to obtain average pooling features and maximum pooling features; S41233: performing channel connection on the average pooling feature and the maximum pooling feature to obtain a connection pooling feature, and performing convolution processing on the connection pooling feature to obtain a second weight matrix; S41234: performing element-wise addition of the first weight matrix and the second weight matrix to obtain an added weight matrix, and performing channel connection of the added weight matrix and the global feature to obtain a concatenated weight matrix; S41235: Perform channel scrambling, group convolution and activation processing on the concatenated weight matrix in sequence to obtain a global local weight matrix.
[0039] Furthermore, the global-local features are input into the dynamic fusion attention layer, and the dynamic fusion attention layer models the contextual correlation of the global-local features in parallel at the spatial and channel levels through channel attention and spatial attention, respectively, to obtain the corresponding first weight matrix and second weight matrix; channel modeling is to obtain the first weight matrix by performing global average pooling, 1×1 convolution, ReLu activation, and 1×1 convolution on the global local features in sequence; and spatial modeling is to perform global average pooling and global maximum pooling on the global local features respectively to obtain average pooling features and maximum pooling features; the average pooling features and the maximum pooling features are channel-connected to obtain connected pooling features, and 7×7 convolution is performed on the connected pooling features. The first weight matrix and the second weight matrix are processed to obtain the second weight matrix; then, the first weight matrix and the second weight matrix are added element-wise to obtain the added weight matrix, and the added weight matrix is channel-connected with the global feature to obtain the spliced weight matrix; finally, the spliced weight matrix is sequentially subjected to channel scrambling, 7×7 group convolution processing and sigmoid function activation processing to obtain the global local weight matrix; the global local weight matrix comprehensively considers the correlation information of both space and channel, which helps to achieve more detailed and comprehensive feature expression; the dynamic fusion attention mechanism DFA is used to simultaneously calculate the spatial correlation and channel correlation between features to dynamically adjust the attention to different regions, better highlight the key information and suppress irrelevant features, and achieve effective fusion.
[0040] It should be noted that S420 is specifically: S421: multiple convolution modules process the target damaged image step by step according to the damaged area positioning image to obtain multiple convolution features of different scales; S422: multiple residual modules process the convolution features output by the last convolution module step by step, and output the residual convolution features through the last residual module; S423: The adaptive feature aggregation module processes the multiple convolution features respectively through multiple re-encoding layers to obtain multiple enhanced coding features, and performs channel connection on the multiple enhanced coding features to obtain adaptive fusion features; S424: passing the adaptive fusion feature to all transposed convolution modules through skip connections, and processing the residual convolution feature and the adaptive fusion feature through the first transposed convolution module to obtain a transposed feature; S425: all transposed convolution modules except the first transposed convolution module process the transposed features step by step according to the adaptive fusion features to obtain target features; S426: The transposed convolution output module processes the target features to obtain a target repaired image.
[0041] Furthermore, the target damaged image is processed step by step according to the damaged area positioning image through four convolution modules with different numbers of channels, so that the four convolution modules generate four convolution features of different scales respectively; the convolution features generated by the last convolution module are input into four residual modules for residual processing step by step, and the residual convolution features are output through the last residual module; the adaptive feature aggregation module includes four re-encoding layers, and the number of channels of the four re-encoding layers corresponds to the number of channels of the four convolution layers. According to the principle of the same number of channels, the four re-encoding layers process the four convolution features one by one to obtain four enhanced coding features, and multiple enhanced coding features are connected by channels to obtain adaptive fusion features; The adaptive fusion feature is jump-connected to all transposed convolution modules. Before the jump connection enters the transposed convolution module, the adaptive fusion feature is adjusted according to the number of channels and size of the transposed convolution module to be jump-connected to ensure that the adaptive fusion feature can enter the transposed convolution module. The residual convolution feature and the adaptive fusion feature are processed by the first transposed convolution module to obtain the transposed feature; the second transposed convolution module processes the transposed feature and the adaptive fusion feature, and the third transposed convolution module processes the output of the second transposed convolution module and the adaptive fusion feature to obtain the target feature; finally, the target feature is processed by the transposed convolution output module to obtain the target repaired image.
[0042] It should be noted that the re-encoding layer in S423 processes the convolutional features to obtain enhanced coding features; specifically: S4231: performing horizontal average pooling and vertical average pooling on the convolutional features to obtain horizontal pooling features and vertical pooling features respectively, and performing channel connection on the horizontal pooling features and the vertical pooling features to obtain a connection average pooling feature; S4232: performing convolution processing, batch normalization processing, and nonlinear activation processing on the connection average pooling feature in sequence to obtain an activated average pooling feature; S4233: Splitting the activated average pooling feature along the horizontal axis and the vertical axis to obtain a horizontal axis activated pooling feature and a vertical axis activated pooling feature, and performing convolution processing and activation processing on the horizontal axis activated pooling feature and the vertical axis activated pooling feature in sequence to obtain a horizontal axis weight and a vertical axis weight; S4234: Perform element-wise multiplication on the horizontal axis weight, the vertical axis weight and the convolution feature to obtain the weighted convolution feature, perform multiple convolution processes with different expansion rates on the weighted convolution feature to obtain multiple enhanced coding sub-features with different receptive fields, and perform channel connection on the multiple different enhanced coding sub-features to obtain the enhanced coding feature.
[0043] Furthermore, for a single re-encoding layer, first, the re-encoding layer uses two pooling kernels (H, 1) and (1,W) The convolutional features are globally averaged pooled along the horizontal axis (x) and the vertical axis (y) respectively to obtain their global information in the two coordinate directions (horizontal pooling features and vertical pooling features), and the horizontal pooling features and the vertical pooling features are channel-connected to obtain the connected average pooling features, and the connected average pooling features are successively convolved, batch normalized and nonlinearly activated to obtain the activated average pooling features; then, the activated average pooling features are split along the horizontal and vertical axes to obtain the horizontal axis activation pooling features and the vertical axis activation pooling features, and the horizontal axis activation pooling features and the vertical axis activation pooling features are successively convolved and activated to obtain the horizontal axis weight and the vertical axis weight; the horizontal axis weight, the vertical axis weight and the convolutional features are element-wise multiplied to obtain the weighted convolutional features. In view of the fact that the objects and structures in the to-be-repaired content of the target damaged image are often multi-scale and diverse, in order to make the image restoration method more complex, Good performance can also be obtained in complex scenes. By performing convolution processing with 4 different expansion rates on the weighted convolution features, the 4 convolution processing with different expansion rates are 3×3 convolution processing with an expansion rate of 1, 3×3 convolution processing with an expansion rate of 5, 3×3 convolution processing with an expansion rate of 7, and 3×3 convolution processing with an expansion rate of 9, 4 enhanced coding sub-features with different receptive fields are obtained, and then the 4 enhanced coding sub-features with different receptive fields are channel-connected to obtain enhanced coding features; there are 4 re-encoding layers in the adaptive feature aggregation module, and then 4 enhanced coding features are obtained. The adaptive feature aggregation module then connects the 4 enhanced coding features through channels to obtain adaptive fusion features; by re-encoding the convolution features, the effective information therein is enhanced and the interference of irrelevant information is weakened, thereby effectively alleviating the negative impact of the damaged area positioning error on the subsequent image restoration process, and improving the robustness of the high-robustness restoration sub-network.
[0044] Specifically, the recoding layer is expressed by the following formula: , , , , in, Indicates The enhanced encoding features output by the re-encoding layer, Indicates the step length is The maximum pooling of represents the dimensionality reduction operation achieved by 1×1 convolution, Indicates channel connection, Indicates The weighted convolution features in the re-encoding layer, represents a 3×3 convolution with a dilation rate of 1, represents a 3×3 convolution with a dilation rate of 5, represents a 3×3 convolution with a dilation rate of 7, represents a 3×3 convolution with a dilation rate of 9, Indicates the input The size of the convolutional features of the re-encoding layer, represents the size of the convolutional features input to the 4th re-encoding layer, Indicates the input The convolutional features of the re-encoding layers, represents the horizontal axis weight, represents the vertical axis weight, represents element-wise multiplication, represents the activation of average pooling features, represents the horizontal axis pooling feature, represents the vertical axis pooling feature, represents 1×1 convolution, represents nonlinear activation, Represents the horizontal axis activation pooling feature, Indicates the vertical axis activation pooling feature, Represents sigmoid activation.
[0045] According to a blind image restoration method of this embodiment, it is composed of two sub-networks: a damaged region localization sub-network (DRLN) and a high robustness restoration sub-network (HRN). DRLN captures the semantic inconsistency between the damaged region and the known region by analyzing the deep semantic information of the damaged image to locate the damaged region; the global-local feature analysis (GLFA) module analyzes the semantic information of the entire damaged image in detail from two different levels, global and local, respectively, so as to effectively improve the positioning accuracy of the damaged region. When the damage information is not a single constant value, or even similar to the known information, the accurate damaged region can still be located. HRN is guided by the damaged region obtained by positioning to complete the subsequent image restoration task. In order to weaken the negative impact of the damaged region positioning error on blind image restoration and enhance the restoration robustness, this embodiment proposes an adaptive feature aggregation module (AFA). In this module, firstly, multiple convolutional features (outputs of the convolutional module) are enhanced by the proposed multiple recoding layers. Subsequently, the AFA module passes the enhanced feature channel connection aggregation to the decoding stage (transposed convolution module and transposed convolution output module), providing important supporting information, thereby effectively improving the image restoration performance and generating high-quality restored images; when the damage information is not a single constant value and the size, position, and shape of the damaged area are uncertain, high-quality restored images are generated with high robustness.
[0046] It should be noted that the blind image restoration method provided in the embodiment of the present application can be executed by a blind image restoration system, or a control module in the blind image restoration system for executing and loading a blind image restoration method. In the embodiment of the present application, a blind image restoration method provided in the embodiment of the present application is described by taking a blind image restoration system executing and loading a blind image restoration method as an example.
[0047] A blind image restoration system, comprising: Acquisition module: acquires image data set, preprocesses the image data set, and obtains damaged image data set; Construction module: Construct a highly robust blind image restoration network model, which includes a damaged area positioning sub-network and a highly robust restoration sub-network; Training and testing module: train the highly robust blind image restoration network model based on part of the data in the damaged image dataset, and test the trained highly robust blind image restoration network model based on another part of the data; Prediction module: The target damaged image is input into the tested high robustness blind image restoration network model for processing to obtain the target restored image of the target damaged image.
[0048] A blind image restoration system in an embodiment of the present application may be a device, or a component, integrated circuit, or chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a PDA, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), etc., which is not specifically limited in the embodiment of the present application.
[0049] A blind image restoration system in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0050] The blind image restoration system provided in the embodiment of the present application can achieve Figures 1 to 7 In order to avoid repetition, each process of implementing a blind image restoration method in the method embodiment will not be described here.
[0051] According to a blind image restoration system of this embodiment, an image data set is firstly acquired through an acquisition module, and the image data set is preprocessed to obtain a damaged image data set; then a high-robustness blind image restoration network model is constructed through a construction module, and the high-robustness blind image restoration network model includes a damaged region positioning subnetwork and a high-robustness restoration subnetwork; The training and testing module trains the highly robust blind image restoration network model according to part of the data in the damaged image dataset, and tests the trained robust blind image restoration network model according to another part of the data; finally, the prediction module inputs the target damaged image into the tested highly robust blind image restoration network model for processing to obtain the target restoration image of the target damaged image; through the cooperation between various modules, the automatic restoration of the damaged image is realized.
[0052] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, each process of the above-mentioned blind image restoration method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0053] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned blind image restoration method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0054] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0055] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0056] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0057] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A blind image restoration method, characterized in that: The method comprises: Acquire an image data set, and preprocess the image data set to obtain a damaged image data set; Constructing a highly robust blind image restoration network model, wherein the highly robust blind image restoration network model includes a damaged region positioning subnetwork and a highly robust restoration subnetwork; Training the highly robust blind image restoration network model according to part of the data in the damaged image dataset, and testing the trained highly robust blind image restoration network model according to another part of the data; The target damaged image is input into the tested high robustness blind image restoration network model for processing to obtain a target restoration image of the target damaged image.
2. A blind image restoration method according to claim 1, characterized in that: The target damaged image is input into the tested high robustness blind image restoration network model for processing to obtain a target restoration image of the target damaged image; specifically: The target damaged image is input into the highly robust blind image restoration network model, and the damaged area positioning subnetwork performs positioning processing on the target damaged image to obtain a damaged area positioning image; The highly robust repair subnetwork repairs the target damaged image according to the damaged area positioning image to obtain the target repaired image.
3. A blind image restoration method according to claim 2, characterized in that: The damaged area positioning sub-network includes multiple annular residual modules and a global local feature analysis module. The damaged area positioning sub-network performs positioning processing on the target damaged image to obtain a damaged area positioning image; specifically: The target damaged image is input into the damaged area positioning sub-network, and the multiple annular residual modules process the target damaged image to obtain a rough positioning feature map; The global-local feature analysis module processes the rough positioning feature map and the target damaged image to obtain the damaged area positioning image.
4. A blind image restoration method according to claim 3, characterized in that: The global-local feature analysis module includes a global feature extraction layer, a local feature extraction layer and a dynamic fusion attention layer; the global-local feature analysis module processes the rough positioning feature map and the target damaged image to obtain the damaged area positioning image; specifically: The coarse positioning feature map is input into the global local feature analysis module, and channel segmentation is performed on the coarse positioning feature map to obtain a first segmentation feature and a second segmentation feature; The global feature extraction layer performs global feature extraction processing on the first segmentation feature to obtain a global feature, and the local feature extraction layer performs local feature extraction processing on the second segmentation feature to obtain a local feature; Performing element-wise addition on the global features and the local features to obtain global local features, and inputting the global local features into the dynamic fusion attention layer for processing to obtain a global local weight matrix, and solving a complementary matrix of the global local weight matrix; Performing element-wise multiplication processing on the global-local weight matrix and the global feature to obtain a global attention feature, and performing element-wise multiplication processing on the complementary matrix and the local feature to obtain a local attention feature; The global attention features, local attention features, global features and local features are added element-wise to obtain global local attention features, and the global local attention features, the coarse positioning feature map and the target damaged image are channel-connected and mapped to obtain the damaged area positioning image.
5. A blind image restoration method according to claim 4, characterized in that: The global feature extraction layer performs global feature extraction processing on the first segmentation feature to obtain a global feature, specifically: Mapping the first segmentation features into a key matrix, a query matrix and a value matrix, wherein the key matrix is mapped by group convolution, and the value matrix is mapped by convolution; Performing channel connection on the key matrix and the query matrix to obtain a connection matrix, and performing two consecutive convolution processes on the connection matrix to obtain a connection attention weight; The connection attention weight is element-wise multiplied by the value matrix to obtain an attention value matrix, and the attention value matrix is element-wise added to the key matrix to obtain the global feature.
6. A blind image restoration method according to claim 4, characterized in that: The local feature extraction layer performs local feature extraction processing on the second segmentation feature to obtain a local feature, specifically: Performing multiple stacked convolution processes on the second segmentation feature to obtain a deep feature, and performing convolution processes and activation processes on the deep feature in sequence to obtain a deep activation feature; Performing element-wise multiplication on the depth activation feature and the second segmentation feature to obtain a depth activation segmentation feature, and performing element-wise addition on the depth activation segmentation feature and the second segmentation feature to obtain the local feature; Among them, the process of single stacked convolution processing is: convolution processing is performed on the features, and depth-separable convolution processing is performed on the features after the convolution processing, and then element-wise addition is performed on the features after the convolution processing and the features processed by the depth-separable convolution processing.
7. A blind image restoration method according to claim 4, characterized in that: The global and local features are input into the dynamic fusion attention layer for processing to obtain a global and local weight matrix, which is specifically: Performing global average pooling, convolution processing, activation processing, and convolution processing on the global and local features in sequence to obtain a first weight matrix; Performing global average pooling and global maximum pooling on the global and local features respectively to obtain average pooling features and maximum pooling features; Perform channel connection on the average pooling feature and the maximum pooling feature to obtain a connection pooling feature, and perform convolution processing on the connection pooling feature to obtain a second weight matrix; Performing element-wise addition of the first weight matrix and the second weight matrix to obtain an added weight matrix, and performing channel connection of the added weight matrix and the global feature to obtain a concatenated weight matrix; The concatenated weight matrix is sequentially subjected to channel scrambling, group convolution and activation processing to obtain the global-local weight matrix.
8. A blind image restoration method according to claim 2, characterized in that: The highly robust restoration subnetwork includes a plurality of different convolution modules, a plurality of different transposed convolution modules, a plurality of identical residual modules, a transposed convolution output module and an adaptive feature fusion module; the number of the convolution modules is one more than the number of the transposed convolution modules; the highly robust restoration subnetwork performs restoration processing on the target damaged image according to the damaged area positioning image to obtain the target restored image; specifically: The plurality of convolution modules process the target damaged image step by step according to the damaged area positioning image to obtain a plurality of convolution features of different scales; The plurality of residual modules process the convolution features output by the last convolution module step by step, and output the residual convolution features through the last residual module; The adaptive feature aggregation module processes the multiple convolutional features respectively through multiple re-encoding layers to obtain multiple enhanced coding features, and performs channel connection on the multiple enhanced coding features to obtain adaptive fusion features; The adaptive fusion feature is transferred to all transposed convolution modules through skip connections, and the residual convolution feature and the adaptive fusion feature are processed by the first transposed convolution module to obtain the transposed feature; All transposed convolution modules except the first transposed convolution module process the transposed features step by step according to the adaptive fusion features to obtain target features; The transposed convolution output module processes the target features to obtain the target repaired image.
9. A blind image restoration method according to claim 8, characterized in that: The re-encoding layer processes the convolutional features to obtain the enhanced encoding features; Specifically: Performing horizontal average pooling and vertical average pooling on the convolutional features to obtain horizontal pooling features and vertical pooling features, respectively, and performing channel connection on the horizontal pooling features and the vertical pooling features to obtain a connection average pooling feature; The connected average pooling feature is sequentially subjected to convolution processing, batch normalization processing, and nonlinear activation processing to obtain an activated average pooling feature; The activated average pooling feature is split along the horizontal axis and the vertical axis to obtain a horizontal axis activated pooling feature and a vertical axis activated pooling feature, and the horizontal axis activated pooling feature and the vertical axis activated pooling feature are sequentially subjected to convolution processing and activation processing to obtain a horizontal axis weight and a vertical axis weight; The horizontal axis weight, the vertical axis weight and the convolution feature are element-wise multiplied to obtain a weighted convolution feature, the weighted convolution feature is subjected to a plurality of convolution processes with different dilation rates to obtain a plurality of enhanced coding sub-features with different receptive fields, and a plurality of different enhanced coding sub-features are channel-connected to obtain the enhanced coding feature.
10. A blind image restoration system capable of implementing a blind image restoration method according to any one of claims 1 to 9, characterized in that: The system comprises: Acquisition module: acquires an image data set, preprocesses the image data set, and obtains a damaged image data set; Construction module: constructing a highly robust blind image restoration network model, wherein the highly robust blind image restoration network model includes a damaged region positioning subnetwork and a highly robust restoration subnetwork; Training and testing module: training the highly robust blind image restoration network model according to part of the data in the damaged image data set, and testing the trained highly robust blind image restoration network model according to another part of the data; Prediction module: input the target damaged image into the tested high robustness blind image restoration network model for processing to obtain a target restoration image of the target damaged image.
Citation Information
Patent Citations
Image blind restoration method based on semantic inconsistency detection
CN114897738A
Image restoration method, image restoration device, electronic equipment and storage medium
CN118628406A
Image restoration method based on multi-scale residual module and feature fusion
CN119599914A
Method for image motion deblurring, apparatus, electronic device and medium therefor
US20240404025A1