Image restoration method and system based on structured texture reconstruction from global and local perspectives

By adopting a structured texture reconstruction method with global and local perspectives in image repair, combined with spatial adaptive normalization and anti-normalization strategies, the problem of information loss during convolutional downsampling is solved, and high-quality image repair effect is achieved.

CN119323715BActive Publication Date: 2025-05-13HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411413492.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-10
Publication Date
2025-05-13
Estimated Expiration
2044-10-10

AI Technical Summary

Technical Problem

The prior art ignores the information of the structure and texture feature maps during convolution downsampling, resulting in information loss and non-ideal sampling output.

Method used

The structured texture reconstruction method based on global and local perspectives is adopted. By extracting texture feature maps and global structural feature maps, combining spatial adaptive normalization and anti-normalization strategies, the global and local normalized texture feature maps are reconstructed, and image repair is performed through feature map fusion and balance modules.

Benefits of technology

It reduces the information loss of structural feature maps, improves the reconstruction effect of texture feature maps, and ensures the quality and consistency of image repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119323715B_ABST
    Figure CN119323715B_ABST
Patent Text Reader

Abstract

The present invention discloses an image restoration method and system based on structured texture reconstruction from global and local perspectives, and relates to the field of image processing. The texture feature map and global structure feature map of an image are extracted, and the residual local structure feature map is extracted from the texture feature map; the global texture feature map and the local texture feature map are reconstructed through the global structure feature map by using spatial adaptive normalization and anti-normalization strategies; the global texture feature map and the local texture feature map are output into two streams by element addition according to the normalization strategy, and are prepared for the next layer to be reconstructed through the global structure feature map and the local structure feature map; the reconstructed global structure feature map is enhanced twice to keep the reconstructed global structure feature map balanced with the local structure feature map; the reconstructed global and local structure feature maps are divided into two groups, and feature balancing is performed in a cross-layer balancing module to complete the image restoration. The present invention has advantages in both low-resolution and high-resolution images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and more particularly to an image restoration method and system based on structured texture reconstruction from global and local perspectives. Background Art

[0002] Image restoration has made substantial progress due to the encoder and decoder pipeline, which benefits from the convolutional downsampling of convolutional neural networks (CNNs), drawing masked areas from known region semantics within the encoder, coupled with the upsampling process of the decoder for the final restoration output. Recent studies intuitively identify high-frequency structures and low-frequency textures extracted from the encoder through convolutional neural networks (CNNs), and then use them for ideal upsampling restoration. However, existing techniques inevitably ignore the information loss of structural and texture feature maps during the convolutional downsampling process, and are therefore susceptible to non-ideal sampling outputs. How to solve the above problems urgently requires continued research by those skilled in the art. Summary of the invention

[0003] In view of this, the present invention provides an image restoration method and system based on structured texture reconstruction from global and local perspectives.

[0004] In order to achieve the above object, the present invention adopts the following technical solution:

[0005] An image restoration method based on structured texture reconstruction from global and local perspectives comprises the following steps:

[0006] Extracting the texture feature map and the global structure feature map of the image to be repaired, and extracting the residual local structure feature map from the texture feature map;

[0007] Adopting spatial adaptive normalization and denormalization strategies, the global texture feature map with global structure feature map is reconstructed into global normalization, and the local texture feature map with local residual structure feature map is reconstructed into local normalization.

[0008] The global texture feature map and the local texture feature map are fused together by element-wise addition and convolved to be downsampled to the next layer. The output is split into two streams according to the texture normalization strategy and reconstructed by the global structure feature map and the local structure feature map in preparation for the next layer.

[0009] The reconstructed global structure feature map is enhanced twice to keep the reconstructed global structure feature map balanced with the local structure feature map, and is upsampled from the decoder at the same time; the reconstructed global and local structure feature maps are divided into two groups, and feature balancing is performed in the cross-layer balancing module to complete the image restoration.

[0010] Optionally, a globally normalized global texture feature map is reconstructed through a global structural feature map, and a locally normalized local texture feature map is reconstructed through a local residual structural feature map, as follows:

[0011] Global normalization from c T The statistical average of the pixels on the entire feature map is calculated by the channel and variance for:

[0012]

[0013] Among them, H k Represents the height of the k-th layer feature map; W k Indicates the width of the k-th layer feature map; Represents the kth layer C of the convolutional neural network T Texture feature map of the channel; Indicates that at the kth layer C T The statistical average of pixels on the entire texture feature map calculated by the channel;

[0014] Local normalization calculates the statistical average of pixels in different channels at x and y positions and variance for:

[0015]

[0016] Optionally, it is characterized by a spatial adaptive normalization and anti-normalization strategy, the calculation formula is as follows:

[0017]

[0018] In the formula, Represents the kth layer C of the convolutional neural network T(S) Texture or structural feature map of the channel; is the texture or structural feature map reconstructed by the k-th layer of the convolutional neural network. Upsample(·) Upsample to match γ is the weight, β is the bias, Indicates that at the kth layer C T(s) The statistical average of pixels on the entire texture feature map calculated by the channel, Indicates that at the kth layer C T(s) The variance of pixels on the entire texture feature map calculated by the channel.

[0019] Optionally, convolution of high-frequency structure and low-frequency texture feature maps is also included during the convolution downsampling process.

[0020] Optionally, the overall loss function of the image restoration neural network is:

[0021]

[0022] Among them, λ r , adv , and is a balancing hyperparameter, represents the reconstruction loss, represents auxiliary loss, Stands for Against Loss.

[0023] Optionally, it also includes using an edge-preserving smoothing method to remove structural information in the image to be repaired while retaining texture information.

[0024] Optionally, the texture normalization strategy includes two types: reconstructing a globally normalized global texture feature map through a global structure feature map; and reconstructing a locally normalized local texture feature map through a local residual structure feature map.

[0025] An image restoration system based on structured texture reconstruction from global and local perspectives, comprising:

[0026] Feature extraction module: used to extract the texture feature map and global structure feature map of the image to be repaired, and extract the residual local structure feature map from the texture feature map;

[0027] Feature normalization module: used to reconstruct a globally normalized global texture feature map through a global structure feature map and a locally normalized local texture feature map through a local residual structure feature map by using a spatially adaptive normalization and anti-normalization strategy;

[0028] Feature map fusion module: used to fuse the global texture feature map and the local texture feature map together through element addition, and perform convolution to downsample to the next layer, divide the output into two streams according to two texture normalization strategies, and prepare for the next layer to reconstruct through the global structure feature map and the local structure feature map;

[0029] Feature balancing module: used to enhance the reconstructed global structure feature map twice to balance the reconstructed global structure feature map with the local structure feature map, while upsampling from the decoder; divide the reconstructed global and local structure feature maps into two groups, and perform feature balancing in the cross-layer balancing module to complete image restoration.

[0030] It can be seen from the above technical solutions that, compared with the prior art, the present invention provides an image restoration method and system based on structured texture reconstruction from global and local perspectives, which has the following beneficial effects:

[0031] 1. The statistical information of the structural feature map is merged into the texture feature maps of different layers, which reduces the information loss of the structural feature map. At the same time, it is found that the structural feature map can better reconstruct the texture feature map, rather than making it difficult to reconstruct the texture feature map, because the structure is always sparse and can be easily destroyed by the reconstruction of dense texture feature maps.

[0032] 2. Combined with the statistical information of the structural feature map, denormalization is performed on a global and local basis to reconstruct the texture feature map. The structural feature map can reconstruct the local texture normalization better than the global feature map because the structural feature map tends to be more globally sparsity, thereby better supplementing the information of the local dense texture feature map through local normalization. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0034] Figure 1 It is a schematic diagram of the overall structure of the present invention. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0036] The embodiment of the present invention discloses an image restoration method based on structured texture reconstruction from global and local perspectives, comprising the following steps:

[0037] Extracting the texture feature map and the global structure feature map of the image to be repaired, and extracting the residual local structure feature map from the texture feature map;

[0038] Adopting spatial adaptive normalization and denormalization strategies, the global texture feature map with global structure feature map is reconstructed into global normalization, and the local texture feature map with local residual structure feature map is reconstructed into local normalization.

[0039] The global texture feature map and the local texture feature map are fused together by element-wise addition and convolved to be downsampled to the next layer. The output is split into two streams according to two texture normalization strategies and reconstructed by the global structure feature map and the local structure feature map in preparation for the next layer.

[0040] The reconstructed global structure feature map is enhanced twice to keep the reconstructed global structure feature map balanced with the local structure feature map, and is upsampled from the decoder at the same time; the reconstructed global and local structure feature maps are divided into two groups, and feature balancing is performed in the cross-layer balancing module to complete the image restoration.

[0041] Example 1

[0042] like Figure 1 As shown, the present invention does not simply fuse texture and structure feature maps during the encoder convolution downsampling process to suppress the sparse structure feature map, but adopts a spatial adaptive normalization and denormalization strategy (Formula 6) to reconstruct the texture feature map from the structure feature map after each layer of convolution downsampling. The statistical information of the structure feature map is merged into the texture feature maps of different layers, which reduces the information loss of the structure feature map. At the same time, it is found that the structure feature map can better reconstruct the texture feature map, rather than the opposite. It is difficult to reconstruct the texture feature map, because the structure is always sparse and is easily destroyed by the reconstruction of dense texture feature maps. On the contrary, this sparsity can promote the reconstruction of dense textures. Therefore, the existing techniques for guiding each other between structure and texture feature maps cannot present ideal restoration outputs;

[0043] Based on this, the present invention proposes to reconstruct a texture feature map through a structural feature map, and normalize the texture feature map through global and local normalization of the texture feature map; given a texture feature map, the global and local normalization are defined as: Definition 1: Global normalization calculates the statistical mean and variance of pixels on the entire texture feature map from any channel (see Formula 7 for details); while local normalization calculates the statistical mean and variance of pixels of different channels at a given specific position (see Formula 8 for details).

[0044] On this basis, global normalization can highlight the global statistical properties of the entire feature map, which is called the global texture feature map; local normalization can highlight the local statistical properties of each position, which is called the local texture feature map. Then, combined with the statistical information of the structural feature map, denormalization is performed on a global and local basis to reconstruct the texture feature map. At the same time, it is observed that the structural feature map can better reconstruct the local texture normalization than the global feature map, because the structural feature map is more inclined to global sparsity, thereby better supplementing the information of the local dense texture feature map through local normalization;

[0045] The above observations further prompted the present invention to reconstruct the normalization of the global texture feature map. The present invention proposes to extract the residual local structure from the texture feature map instead of the above global structure feature map. In summary, the present invention discusses 4 variant modules, and has a deep understanding of the global and local texture feature map normalization and structural denormalization. Interestingly, the present invention finds that the best effect is to reconstruct the texture feature map from the structural feature map under global normalization and reconstruct the texture feature map from the local residual structure feature map under local normalization. The two texture feature maps reconstructed from the global and local structures are fused together by element addition, and then convolved to downsample to the next layer. The output is divided into two streams according to the two texture normalization strategies, and is prepared for reconstruction by the global and local structure feature maps for the next layer;

[0046] In the early stage of the entire convolution process inside the encoder, the information from the global structure feature map decreases more slowly than the information from the local structure feature map, and in the later stage, it exceeds the information from the local structure feature map. Therefore, the present invention proposes: enhancing the reconstructed global structure feature map twice to keep it balanced with the local structure feature map, and upsampling from the decoder at the same time. The reconstructed global and local structure feature maps are divided into two groups, and feature balancing is performed in the cross-layer balancing module. The overall pipeline is as follows Figure 1 shown.

[0047] For covered The input image (0 represents the covered part, 1 represents the uncovered part), image restoration will convert the covered image I m =I gt ⊙M is transformed into the complete image I out The core of the pipeline of the present invention is to convolve the high-frequency structure and low-frequency texture feature maps during the convolution downsampling process to mitigate their feature map losses, especially the texture feature map losses.

[0048] The specific steps of partial convolution of the structural feature map are as follows:

[0049] Given an input image I gt The present invention first uses the canny edge detector to construct I gt The edge graph of s , grayscale corresponding Likewise, I m The edge graph is: The corresponding grayscale is: On this basis, the present invention uses a partial convolutional layer to extract structural feature maps from uncovered / known areas. M is used as the input of the first partial convolutional layer, denoted as X S , where the corresponding mask is M. Given XS The current sliding window Xs and its cover Ms on the corresponding area, represents the scaling factor for adjusting the known area scale, so the partial convolution operation at each position can be expressed as

[0050]

[0051] In the formula, x's is the output feature map A pixel, W, b is the weight matrix and bias vector under convolution filtering; T r is the transpose operation; ⊙ represents element-by-element multiplication. After each partial convolution operation, the mask is updated as follows:

[0052]

[0053] where m's is x' s Corresponding mask value. On this basis, the present invention cascades N partial convolution layers together to repair the covered part and update the feature map at the same time. After the partial convolution layer, the feature map is processed by the normalization layer and the activation function. In addition, with the convolution downsampling, the structural information gradually decreases. In order to reduce this structural loss, the present invention adds more structural feature channels so that the network has the ability to encode more structural information. Alternatively, a self-attention mechanism is proposed in the prior art to reduce the loss of feature maps during convolution downsampling. However, unlike dense texture information, the structure is always sparse and not easy to reconstruct as the input of the attention mechanism. Based on the above, the present invention can obtain the structural feature map of a known area, that is, the feature map of different partial convolution layers, denoted as Where N is the number of layers.

[0054] Some convolutions and transformers are used for texture feature maps. The specific steps are as follows:

[0055] First, an edge-preserving smoothing method is used to remove the structural information in the input image, while retaining the texture information. gt The edge-preserving smoothed image is denoted as The corresponding cover image is given by In order to extract the texture feature map with global and local patterns well, in addition to the above local pattern based In addition to the local convolution of , we also deploy the Visual Transformer (ViT) and specialize it by the transformer block strategy to extract the global correlation between patches from the texture feature map after passing through a total of N convolutional downsampling stages corresponding to N layers. First, to facilitate the transformer, we use a partial convolution head composed of two partial convolutional blocks to perform downsampling as the first stage to obtain a 1 / 2 size feature map for the logo. Its input is I m, and M are connected in series, where the first convolutional layer is used to change the input dimension and the second one is used as a downsampling layer to reduce the resolution.

[0056] exist Based on this, the transformer body is further used to process the mark by establishing a remote communication, which includes the N-1 level of the transformer block adjusted after the first level. For the lth block of the tth stage (t=2,…,N), the output feature map Depend on:

[0057]

[0058] In the formula, is the input feature map of the lth block at the tth stage; FC(·) is the fully connected layer, MLP(·) is the multi-layer perceptron. MCA(·) represents the multi-head contextual self-attention operation, and the calculation formula is:

[0059]

[0060] Where Q t,l ,K t,l ,V t,l are query, key, and value matrices respectively, d t,l is the embedding dimension. τ (a large positive integer, set to 100 in the experiment of the present invention) is used to adjust the attention area. is the mask of the lth block of the tth segment, expressed as:

[0061]

[0062] where i is the index of the pixel in each patch. Based on this, the mask The refinement of follows a rule, that is, as long as there is at least one valid flag before, all flags in the window will be updated to be valid after the operation; if all flags in the window are invalid, they will still be invalid after the operation. On this basis, the present invention can obtain the feature map of texture information from the known area, that is, the feature map of different stages, denoted as

[0063] Given the structural and texture feature maps of each layer, this paper revisits and deploys spatially adaptive normalization and denormalization strategies to merge statistical information on feature maps across different layers, as follows:

[0064]

[0065] In the formula, Denotes the kth layer C of the convolutional neural network (CNNs) T(S) Texture (structural) feature map of the channel; is the texture (structure) feature map reconstructed by the k-th layer structure (texture) feature map of convolutional neural networks (CNNs), Upsample(·) Upsample to match In this case, through statistics The average and variance One is reconstructed by normalizing the other. The present invention alternately reconstructs the two cases, and shows the feature maps reconstructed on the three layers of masked images during the convolution process in Table 1, wherein the present invention finds that the texture feature map reconstructed by the structural feature map in Table 1(a) shows consistently superior performance than the texture feature map reconstructed by the structural feature map in Table 1(b), especially for the unmasked areas, across three layers. This is because the structural feature map is always sparse and can be easily reconstructed and decomposed by dense texture feature maps. On the contrary, this sparsity can well restore the lost structural information in the dense texture feature map, which means that the structural feature map should guide the reconstruction of the texture feature map, not the other way around.

[0066] First, the texture feature map is normalized, and then the structure feature map is denormalized and reconstructed. The present invention implements two normalization strategies for the texture feature map, namely global normalization and local normalization, which correspond to the global and local texture feature maps respectively. Given a texture feature map, global normalization is performed from the cth T The statistical average of the pixels on the entire feature map is calculated by the channel and variance for:

[0067]

[0068]

[0069] Therefore, global normalization can highlight the global statistical information of the entire feature map, which is called the global texture feature map; local normalization calculates the statistical average of the pixels of different channels at the x and y positions. and variance for:

[0070]

[0071]

[0072] Therefore, local normalization can highlight the local statistical information of each location, which is called the local texture feature map.

[0073] Based on the above, the reconstructed texture feature maps on different layers of different scales in the encoder are used as the input of the cross-layer balancing module proposed in the present invention, in particular: in the early stage, the texture feature map based on the global structure feature map is reconstructed; in the later stage, the texture feature map based on the local residual structure feature map is reconstructed; in the final stage, the texture feature map of the last layer of the convolutional neural network (CNNs) is reconstructed, where each stream contains five partial convolutions with the same kernel size, but the kernel sizes between different streams are different. Afterwards, the reconstructed texture feature maps from the three streams are connected to obtain output feature maps of the same size. Finally, the present invention fuses the feature maps from the local residual stream and the global structure stream, expressed as Then, the present invention adopts a feature equalization method to balance the local residual and global structural feature maps in the decoder from different stages of convolutional neural networks (CNNs). In the channel domain, the feature map Equalization is performed by:

[0074]

[0075] Where AvgPool c (·) represents the global average pooling of each channel, W g is a learnable linear projection matrix; σ(·) is the sigmoid function. Next, in the spatial domain, assume that x i yes The eigenvector at position i in , x j is the neighboring feature channel around i at position j, and the equalization is achieved in the following way:

[0076]

[0077]

[0078]

[0079] in and is the feature vector after spatial similarity and range similarity measurement respectively. Gaussian(·) represents Gaussian function, C' is the normalization factor. s and v are the corresponding neighboring regions. By formula (9)-formula (12), Corrected to produce consistent local residual and global structural feature maps. Multiple loss functions are introduced during the training process, including reconstruction loss to improve image quality, auxiliary loss to enhance structural information, and adversarial loss to ensure content consistency.

[0080] Reconstruction loss. The present invention uses reconstruction loss to measure the restoration result. outCompared with the real image I gt The pixel-by-pixel difference between , expressed as:

[0081]

[0082] where ||·||1 means Norm.

[0083] Auxiliary loss. In order to reconstruct texture through structure, the present invention designs an auxiliary loss function to facilitate the extraction of structural information. Specifically, the present invention introduces a decoder To repair the edge map Based on this, the present invention measures With the unmasked image I H The pixel-by-pixel difference between , expressed as:

[0084]

[0085] In order to enhance the perception of image quality, the present invention uses adversarial loss to To distinguish the restored fake image I out With the real image, this can be expressed as

[0086]

[0087] The present invention processes the texture feature map of the cross-layer balancing module to make it have the same size as the input reconstructed texture feature atlas, so as to perform upsampling through a bottom-up strategy to generate the final image restoration result. Specifically, the present invention uses the output of the last layer of the encoder as the input of the decoder and splices it with the texture feature map from the reconstructed texture feature atlas in the channel dimension. The upsampling process is achieved by convolution fusion. This iterative process will continue until the desired resolution is reached.

[0088] Overall loss. In summary, the present invention defines overall loss as:

[0089]

[0090] Among them, λ r , adv , and are balancing hyperparameters, and in the experiments of the present invention, these parameters are empirically set to 1, 0.1, 1, and 1, respectively. The entire process is trained by minimizing equation (16).

[0091] Finally, this embodiment also discloses an image restoration system based on structured texture reconstruction from global and local perspectives, including:

[0092] Feature extraction module: used to extract the texture feature map and global structure feature map of the image to be repaired, and extract the residual local structure feature map from the texture feature map;

[0093] Feature normalization module: used to reconstruct a globally normalized global texture feature map through a global structure feature map and a locally normalized local texture feature map through a local residual structure feature map by using a spatially adaptive normalization and anti-normalization strategy;

[0094] Feature map fusion module: used to fuse the global texture feature map and the local texture feature map together through element addition, and perform convolution to downsample to the next layer, divide the output into two streams according to two texture normalization strategies, and prepare for the next layer to reconstruct through the global structure feature map and the local structure feature map;

[0095] Feature balancing module: used to enhance the reconstructed global structure feature map twice to balance the reconstructed global structure feature map with the local structure feature map, while upsampling from the decoder; divide the reconstructed global and local structure feature maps into two groups, and perform feature balancing in the cross-layer balancing module to complete image restoration.

[0096] Example 2

[0097] The proposed method is verified on three typical datasets with different characteristics, including: Paris Street View (PSV), a collection of Paris street view images, including 14,900 training images and 100 verification images; CelebA, a face dataset containing 30,000 aligned face images, divided into 28,000 training images and 2,000 verification images; Places2 contains more than 1.8 million natural images from different scenes. The network is trained using 256×256 pixel images and irregular masks [8]. Since the number of channels of the texture feature map is related to the dimension of the label vector, a lower vector dimension will affect the representation ability of the label, while a higher dimension will lead to the computational burden of the self-attention mechanism. Taking these factors into consideration, the present invention proposes a multi-scale texture feature map, in which the number of feature maps of all layers is set to 180. The present invention adopts the Adam optimizer, where β1=0.5 and β2=0.99, and the learning rates of the generator and discriminator are set to 5×10 -4 and 10 -4 We train the proposed method for a total of 150 epochs on PSV and CelebA, and 20 epochs on Places2. In particular, we use 6 stages to extract multi-scale features, and the number of early and late stages in the cross-layer balancing module is set to 3. All experiments are implemented using the pytorch framework and run on 4 NVIDIA 2080TI GPUs.

[0098] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0099] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An image restoration method based on structured texture reconstruction from global and local perspectives, characterized in that: The following steps are involved: Extracting the texture feature map and the global structure feature map of the image to be repaired, and extracting the residual local structure feature map from the texture feature map; Adopting spatial adaptive normalization and denormalization strategies, the global texture feature map with global structure feature map is reconstructed into global normalization, and the local texture feature map with local residual structure feature map is reconstructed into local normalization. Spatial adaptive normalization and anti-normalization strategy, the calculation formula is as follows: In the formula, Represents the kth layer C of the convolutional neural network T(S) Texture or structural feature map of the channel; is the texture or structural feature map reconstructed by the k-th layer of the convolutional neural network. Upsample(·) Upsample to match γ is the weight, β is the bias, Indicates that at the kth layer C T(s) The statistical average of pixels on the entire texture feature map calculated by the channel, Indicates that at the kth layer C T(s) The variance of pixels on the entire texture feature map calculated by the channel; The global texture feature map and the local texture feature map are fused together by element-wise addition and convolved to be downsampled to the next layer. The output is split into two streams according to the texture normalization strategy and prepared for the next layer to be reconstructed by the global structure feature map and the local structure feature map. The reconstructed global structure feature map is enhanced twice to keep the reconstructed global structure feature map balanced with the local structure feature map, and is upsampled from the decoder at the same time; the reconstructed global and local structure feature maps are divided into two groups, and feature balancing is performed in the cross-layer balancing module to complete the image restoration.

2. The image restoration method based on structured texture reconstruction from global and local perspectives according to claim 1, characterized in that: The global normalized global texture feature map is reconstructed through the global structure feature map, and the local normalized local texture feature map is reconstructed through the local residual structure feature map, as follows: Global normalization from c T The statistical average of the pixels on the entire feature map is calculated by the channel and variance for: Among them, H k Represents the height of the k-th layer feature map; W k Indicates the width of the k-th layer feature map; Represents the kth layer C of the convolutional neural network T Texture feature map of the channel; Indicates that at the kth layer C T The statistical average of pixels on the entire texture feature map calculated by the channel; Local normalization calculates the statistical average of pixels in different channels at x and y positions and variance for: .

3. The image restoration method based on structured texture reconstruction from global and local perspectives according to claim 1, characterized in that: It also includes convolving high-frequency structure and low-frequency texture feature maps during convolution downsampling.

4. The image restoration method based on global and local perspective structured texture reconstruction according to claim 1, characterized in that: The overall loss function of the image restoration neural network is: Among them, λ r , adv , and is a balancing hyperparameter, represents the reconstruction loss, represents auxiliary loss, Stands for Against Loss.

5. The image restoration method based on structured texture reconstruction from global and local perspectives according to claim 1, characterized in that: It also includes using an edge-preserving smoothing method to remove structural information in the image to be repaired while retaining texture information.

6. The image restoration method based on global and local perspective structured texture reconstruction according to claim 1, characterized in that: There are two texture normalization strategies: reconstructing a globally normalized global texture feature map through a global structure feature map; and reconstructing a locally normalized local texture feature map through a local residual structure feature map.

7. An image restoration system based on structured texture reconstruction from global and local perspectives, characterized in that: include: Feature extraction module: used to extract the texture feature map and global structure feature map of the image to be repaired, and extract the residual local structure feature map from the texture feature map; Feature normalization module: used to reconstruct a globally normalized global texture feature map through a global structure feature map and a locally normalized local texture feature map through a local residual structure feature map by using a spatially adaptive normalization and anti-normalization strategy; Spatial adaptive normalization and anti-normalization strategy, the calculation formula is as follows: In the formula, Represents the kth layer C of the convolutional neural network T(S) Texture or structural feature map of the channel; is the texture or structural feature map reconstructed by the k-th layer of the convolutional neural network. Upsample(·) Upsample to match γ is the weight, β is the bias, Indicates that at the kth layer C T(s) The statistical average of pixels on the entire texture feature map calculated by the channel, Indicates that at the kth layer C T(s) The variance of pixels on the entire texture feature map calculated by the channel; Feature map fusion module: used to fuse the global texture feature map and the local texture feature map together through element addition, and perform convolution to downsample to the next layer, divide the output into two streams according to two texture normalization strategies, and prepare for the next layer to reconstruct through the global structure feature map and the local structure feature map; Feature balancing module: used to enhance the reconstructed global structure feature map twice to balance the reconstructed global structure feature map with the local structure feature map, while upsampling from the decoder; divide the reconstructed global and local structure feature maps into two groups, and perform feature balancing in the cross-layer balancing module to complete image restoration.

Citation Information

Patent Citations

  • Image restoration method based on global texture and structure

    CN115035170A

  • Image reconstruction device and method, electronic equipment and readable storage medium

    CN117274095A