Image restoration method and system based on dense multi-scale fusion
By adopting a dense multi-scale fusion method in image repair, combined with hollow convolution and self-attention mechanism, the problem of large-area irregular missing image repair of complex backgrounds and textures is solved, and high-quality repair results are achieved, meeting the requirements of global and local semantic consistency and detail clarity.
Patent Information
- Application Number
- CN202111528555.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-12-14
AI Technical Summary
The prior art is difficult to achieve effective repair in large-area irregular missing images of complex backgrounds or fine textures, especially in maintaining consistency of global semantic structures and texture details.
The image repair method based on dense multi-scale fusion is adopted to construct the structure repair network and the detail repair network, and the global structure and detail texture repair of the image are processed respectively. This method combines dense multi-scale hollow convolution and self-attention mechanisms to enhance the network's receptive field and generation ability, and performs final repair through a dual-spectrum normalized discriminator network.
The large-area missing image repair effect on complex backgrounds and textures is significantly improved. The generated repair images have smooth boundaries, clear details, global and local semantic consistency, and are better than the comparison algorithm in terms of visual effects, peak signal-to-noise ratio, structural similarity and average error.
Smart Images

Figure CN114155171B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image restoration, and in particular to an image restoration method and system based on dense multi-scale fusion. Background Art
[0002] Image restoration is an image processing technique that reconstructs the damaged part of a damaged image by using the remaining area information of the damaged image. The goal is to ensure that the filled restoration area has texture and structural consistency with the remaining area and meets visual authenticity. This task is a research hotspot in the field of image processing and has a wide range of applications, such as cultural relic restoration, object removal, image editing, etc.
[0003] Traditional image restoration methods are mainly divided into methods based on geometric diffusion, texture matching and average sub-ellipticization. The geometric diffusion method is to reconstruct the geometric information of the defective content by using the edge information of the defective area using partial differential equations and variational methods. This work is to propagate the remaining area information to the area to be repaired using the diffusion equation, or to establish a priori data model by studying the geometric information of the image model, and use the variational idea to smoothly transfer the known information to the defective area. This type of method works well when repairing small-area damaged images, but because it is impossible to propagate and reconstruct texture detail information, when the texture background is complex or the defect area is large, the repair result has the disadvantages of inconsistent texture and blurred content in the reconstructed area.
[0004] Therefore, in order to solve the above problems, some scholars later proposed to reconstruct and repair the texture details of the image to be repaired based on the texture matching algorithm. The first idea of this type of technology is to decompose the image model into two parts: structure and texture, and use the above-mentioned variational method to complete the structural information, while in the texture part, the texture details are filled in by using texture synthesis technology. Another idea is to use a similar block texture matching algorithm to find the texture block that is most similar to a pixel point in the area to be repaired, and then copy the texture block to the corresponding defective area, and use texture similarity search iteratively to achieve image repair. This type of algorithm can generate reasonable repair results when the background is simple textured. However, when repairing images with rich texture details or large areas of defects, it is impossible to generate content that is not in the remaining area, resulting in a sharp drop in repair performance and limitations.
[0005] In order to solve the problem of image restoration when large areas of images are damaged, some scholars have proposed the average sub-ellipticization algorithm, which is a reasonable combination of the sub-Riemann sub-elliptic diffusion algorithm and the special local average technology. The restoration is completed in four steps: preprocessing, main diffusion, advanced average and weak smoothing. This has a good effect on the restoration of large-area damaged images, but it requires that the damaged points of the damaged image are well distributed, which greatly limits the scope of application of the algorithm.
[0006] To address the shortcomings of traditional restoration methods, deep neural networks are used for image restoration. The context encoder (CE) restoration algorithm is the earliest to use deep learning to restore images. Combining the encoder-decoder network and the generative adversarial network (GAN), it first learns image features and generates a prediction map corresponding to the area to be repaired in the image, and then determines whether the prediction map comes from the training set and the prediction set. When the generated prediction map is consistent with the real image, the network model parameters reach the optimal state, but the algorithm has poor restoration effect on large-area irregular missing images. Later, some scholars added superimposed dilated convolutions and global and local discriminators to CE, which improved the global semantic consistency of the restoration results and optimized local details, solving the limitation of poor restoration effect of large-area missing images. However, the dilated convolution layer used in the algorithm will lose fine texture information and has a limited information receptive field, so it will generate artifacts and unreasonable structures when processing complex texture images. The convolution algorithm based on contextual attention uses the known convolution filter characteristics to generate patches. The network introduces a spatial propagation layer to enhance the spatial consistency of the restoration results and increase the network's receptive field, which has a good restoration effect on complex texture incomplete images. However, when the unknown missing image is not closely related to the neighboring area, the repair result of the algorithm drops sharply. The multi-discriminator image repair algorithm based on the hybrid dilated convolutional network uses a hybrid dilated convolution kernel to solve the problem of losing key information caused by dilated convolution sparsity. Although this method has a good repair effect on large-area regular incomplete images, it has a poor repair effect when the area of interest is completely missing. In addition, the algorithm has a poor repair effect on large-area incomplete images in irregular areas.
[0007] Based on the above, there is an urgent need to propose a new image restoration method that can take into account both the global semantic structure and texture details of the restoration results when repairing large-area irregular missing parts of complex backgrounds or fine textures. Summary of the invention
[0008] The purpose of the present invention is to provide an image restoration method and system based on dense multi-scale fusion, which improves the restoration effect when repairing large-area missing of complex background and texture.
[0009] To achieve the above object, the present invention provides the following solutions:
[0010] An image restoration method based on dense multi-scale fusion, the restoration method comprising:
[0011] Building a structural repair network;
[0012] Inputting the image to be repaired into the structure repair network to obtain the image after structural repair;
[0013] Construct a detail restoration network;
[0014] Inputting the structurally repaired image into the detail repair network to obtain a detail-repaired image;
[0015] Get real images;
[0016] Using the real image to train a dual-spectrum normalized discriminator network;
[0017] The image after detail restoration is input into a trained dual-spectrum normalization discriminator to obtain a final restored image.
[0018] Optionally, the structure repair network includes: a first encoding module and a first decoding module;
[0019] The first encoding module includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a 16th dense multi-scale dilated convolutional fusion layer; the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the 16th dense multi-scale dilated convolutional fusion layer are connected in sequence;
[0020] The first decoding module includes: a fifth convolutional layer, a first deconvolutional layer, a sixth convolutional layer, a first upsampling layer and a seventh convolutional layer; the fifth convolutional layer, the first deconvolutional layer, the sixth convolutional layer, the first upsampling layer and the seventh convolutional layer are connected in sequence; the fifth convolutional layer is also connected to the first-sixteenth dense multi-scale hole convolution fusion layer.
[0021] Optionally, the detail restoration network specifically includes: a second encoding module and a second decoding module;
[0022] The second encoding module includes two layers, the first layer includes: an eighth convolution layer, a ninth convolution layer, a tenth convolution layer, an eleventh convolution layer, a self-attention mechanism layer, a twelfth convolution layer and a thirteenth convolution layer; the eighth convolution layer, the ninth convolution layer, the tenth convolution layer, the eleventh convolution layer, the self-attention mechanism layer, the twelfth convolution layer and the thirteenth convolution layer are connected in sequence;
[0023] The second layer includes: a fourteenth convolutional layer, a fifteenth convolutional layer, a sixteenth convolutional layer, a seventeenth convolutional layer and a twenty-sixth dense multi-scale dilated convolutional fusion layer, wherein the fourteenth convolutional layer, the fifteenth convolutional layer, the sixteenth convolutional layer, the seventeenth convolutional layer and the twenty-sixth dense multi-scale dilated convolutional fusion layer are connected in sequence;
[0024] The second decoding module includes: a first network connection layer, a second deconvolution layer, an eighteenth convolution layer, an upsampling layer, a nineteenth convolution layer, and a twentieth convolution layer; the first network connection layer, the second deconvolution layer, the eighteenth convolution layer, the upsampling layer, the nineteenth convolution layer, and the twentieth convolution layer are connected in sequence;
[0025] The first network connection layer is connected to the thirteenth convolutional layer and the twenty-sixth dense multi-scale hole convolution fusion layer respectively.
[0026] Optionally, the trained dual-spectrum normalized discriminator network includes:
[0027] Global branch identification layer, local branch identification layer, second network connection layer, third fully connected layer and sigmoid layer.
[0028] Optionally, the global branch identification layer includes: a twenty-first convolutional layer, a twenty-second convolutional layer, a twenty-third convolutional layer, a twenty-fourth convolutional layer, a twenty-fifth convolutional layer, a twenty-sixth convolutional layer and a first fully connected layer; the twenty-first convolutional layer, the twenty-second convolutional layer, the twenty-third convolutional layer, the twenty-fourth convolutional layer, the twenty-fifth convolutional layer, the twenty-sixth convolutional layer and the first fully connected layer are connected in sequence.
[0029] Optionally, the local branch identification layer includes: a twenty-seventh convolutional layer, a twenty-eighth convolutional layer, a twenty-ninth convolutional layer, a thirtieth convolutional layer, a thirty-first convolutional layer and a second fully connected layer; the twenty-seventh convolutional layer, the twenty-eighth convolutional layer, the twenty-ninth convolutional layer, the thirtieth convolutional layer, the thirty-first convolutional layer and the second fully connected layer are connected in sequence.
[0030] Optionally, the number of channels of the first convolutional layer is 64, the number of channels of the second convolutional layer is 128, the number of channels of the third convolutional layer is 128, the number of channels of the fourth convolutional layer is 256, the number of channels of the sixteenth dense multi-scale hole convolutional fusion layer is 256, the number of channels of the fifth convolutional layer is 256, the number of channels of the first deconvolutional layer is 128, the number of channels of the sixth convolutional layer is 128, the number of channels of the first upsampling layer is 64, and the number of channels of the seventh convolutional layer is 3.
[0031] Optionally, the number of channels of the eighth convolutional layer is 64, the number of channels of the ninth convolutional layer is 128, the number of channels of the tenth convolutional layer is 128, the number of channels of the eleventh convolutional layer is 256, the number of channels of the self-attention mechanism layer is 256, the number of channels of the twelfth convolutional layer is 256, the number of channels of the thirteenth convolutional layer is 256, the number of channels of the fourteenth convolutional layer is 64, the number of channels of the fifteenth convolutional layer is 128, the number of channels of the sixteenth convolutional layer is 128, the number of channels of the seventeenth convolutional layer is 256, the number of channels of the twenty-sixth dense multi-scale void convolution fusion layer is 256, the number of channels of the first network connection layer is 512, the number of channels of the second deconvolutional layer is 256, the number of channels of the eighteenth convolutional layer is 128, the number of channels of the upsampling layer is 64, the number of channels of the nineteenth convolutional layer is 64, and the number of channels of the twentieth convolutional layer is 3.
[0032] Optionally, the number of channels of the twenty-first convolutional layer is 64, the number of channels of the twenty-second convolutional layer is 128, the number of channels of the twenty-third convolutional layer is 256, the number of channels of the twenty-fourth convolutional layer is 512, the number of channels of the twenty-fifth convolutional layer is 512, the number of channels of the twenty-sixth convolutional layer is 512, the number of channels of the first fully connected layer is 512, the number of channels of the twenty-seventh convolutional layer is 64, the number of channels of the twenty-eighth convolutional layer is 128, the number of channels of the twenty-ninth convolutional layer is 256, the number of channels of the thirtieth convolutional layer is 512, the number of channels of the thirty-first convolutional layer is 512, the number of channels of the second fully connected layer is 512, the number of channels of the second network connection layer is 1024, and the number of channels of the third fully connected layer is 1024.
[0033] Based on the above method in the present invention, the present invention further provides an image restoration system based on dense multi-scale fusion, the restoration system comprising:
[0034] Repair network building module, used to build a structural repair network;
[0035] A structural repair module, used for inputting the image to be repaired into the structural repair network to obtain the structurally repaired image;
[0036] A detail restoration network construction module is used to construct a detail restoration network;
[0037] A detail restoration module, used for inputting the structurally restored image into the detail restoration network to obtain a detail restored image;
[0038] A real image acquisition module, used for acquiring a real image;
[0039] A training module, used for training a dual-spectrum normalization discriminator network using the real image;
[0040] The final image restoration module is used to input the image after detail restoration into the trained dual-spectrum normalization discriminator to obtain the final restored image.
[0041] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0042] The above-mentioned method and system in the present invention input the image to be repaired into the structure repair module, and perform global structure repair on the image by using a generation network based on dense multi-scale fusion hole convolution; then input the structure repair result into the detail repair module, and perform detail texture repair on the image by using a dense multi-scale hole convolution network and a self-attention mechanism convolution network parallel to it. It can repair large-area defects and complex texture images, generate fine textures and enhance the global and local semantic consistency of the image, perform spectral normalization processing on the discriminator module, stabilize the discriminator training, and improve the generation ability of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0044] Figure 1 This is a flow chart of an image restoration method based on dense multi-scale fusion according to an embodiment of the present invention;
[0045] Figure 2 This is a flowchart of an image restoration method based on dense multi-scale fusion according to an embodiment of the present invention;
[0046] Figure 3 This is a structural diagram of a sixteen-layer dense multi-scale dilated convolution fusion module according to an embodiment of the present invention;
[0047] Figure 4 Schematic diagram of the restoration results of various algorithms in the embodiments of the present invention on the CelebAHQ dataset;
[0048] Figure 5 Schematic diagram of restoration results of various algorithms in the Paris_StreetView dataset according to the embodiments of the present invention;
[0049] Figure 6 The figure is a schematic diagram of the structure of an image restoration system based on dense multi-scale fusion according to an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0051] The purpose of the present invention is to provide an image restoration method and system based on dense multi-scale fusion, which improves the restoration effect when repairing large-area missing of complex background and texture.
[0052] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0053] Figure 1 This is a flow chart of an image restoration method based on dense multi-scale fusion according to an embodiment of the present invention. Figure 2This is a flowchart of an image restoration method based on dense multi-scale fusion according to an embodiment of the present invention. Figure 1 and Figure 2 , the method of the present invention comprises:
[0054] Step 101: Construct a structure repair network.
[0055] Specifically, the structure repair network includes: a first encoding module and a first decoding module;
[0056] The first encoding module includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a 16th dense multi-scale dilated convolutional fusion layer; the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the 16th dense multi-scale dilated convolutional fusion layer are connected in sequence;
[0057] The first decoding module includes: a fifth convolutional layer, a first deconvolutional layer, a sixth convolutional layer, a first upsampling layer and a seventh convolutional layer; the fifth convolutional layer, the first deconvolutional layer, the sixth convolutional layer, the first upsampling layer and the seventh convolutional layer are connected in sequence; the fifth convolutional layer is also connected to the first-sixteenth dense multi-scale hole convolution fusion layer.
[0058] The number of channels of the first convolutional layer is 64, the number of channels of the second convolutional layer is 128, the number of channels of the third convolutional layer is 128, the number of channels of the fourth convolutional layer is 256, the number of channels of the sixteenth dense multi-scale hole convolutional fusion layer is 256, the number of channels of the fifth convolutional layer is 256, the number of channels of the first deconvolutional layer is 128, the number of channels of the sixth convolutional layer is 128, the number of channels of the first upsampling layer is 64, and the number of channels of the seventh convolutional layer is 3.
[0059] Most existing methods use dilated convolution to increase the information receptive field. Although a larger receptive field is guaranteed without increasing the number of learnable weights, the kernel of dilated convolution is sparse, and many pixels will be skipped during iterative calculation, resulting in fence phenomenon and poor detail restoration. In order to expand the receptive field and solve the problem of sparse dilated convolution, it is proposed to use dense multi-scale fused dilated convolution blocks to increase the receptive field layer by layer, and replace the large convolution kernel with a 3×3 convolution kernel. Figure 3As shown in the figure, the first column is a convolution with a kernel size of 3, the second column is a dilated convolution with a kernel size of 3, the dilation rates are 1, 2, 4, and 8 from top to bottom, the third column is an element-level feature addition, the fourth column is a convolution with a kernel size of 3, the fifth column is a Concat layer based on a specified feature axis, the sixth column is a convolution with a kernel size of 1, and the seventh column is an element-level feature addition. The number of channels of the first layer of the dense multi-scale dilated convolution block convolution input is reduced to 64 to reduce parameters, and then sent to four branches using dilated convolutions with different dilation rates, denoted as x i (i=1,2,3,4). In the fourth column, the first and second convolution layers and the Concat layer all use instance normalization and ReLU activation function, and the number of output channels is 64. The third convolution layer only uses instance normalization, and the number of output channels is 256. i There is a corresponding convolution, using K i (·) indicates that dense multi-scale features are obtained from the combination of sparse multi-scale features through accumulation. i K i The output of (·), the combined part is expressed as:
[0060]
[0061] Finally, 1×1 convolution is used to merge the cascade features. In summary, the dense multi-scale fusion block densely connects the dilated convolutions with different dilation rates, and transmits the output feature map to the next layer, so that each layer has the initial input information, ensuring the maximum information transmission, greatly enhancing the receptive field of the dilated convolution and reducing the sparsity of the dilated convolution.
[0062] Step 102: Input the image to be repaired into the structure repair network to obtain the image after structural repair.
[0063] Among them, the structure restoration network obtains the global preferential total energy of the image and repairs the structure of the missing area.
[0064] Step 103: Construct a detail restoration network.
[0065] The detail restoration network specifically includes: a second encoding module and a second decoding module;
[0066] The second encoding module includes two layers, the first layer includes: an eighth convolution layer, a ninth convolution layer, a tenth convolution layer, an eleventh convolution layer, a self-attention mechanism layer, a twelfth convolution layer and a thirteenth convolution layer; the eighth convolution layer, the ninth convolution layer, the tenth convolution layer, the eleventh convolution layer, the self-attention mechanism layer, the twelfth convolution layer and the thirteenth convolution layer are connected in sequence;
[0067] The second layer includes: a fourteenth convolutional layer, a fifteenth convolutional layer, a sixteenth convolutional layer, a seventeenth convolutional layer and a twenty-sixth dense multi-scale dilated convolutional fusion layer, wherein the fourteenth convolutional layer, the fifteenth convolutional layer, the sixteenth convolutional layer, the seventeenth convolutional layer and the twenty-sixth dense multi-scale dilated convolutional fusion layer are connected in sequence;
[0068] The second decoding module includes: a first network connection layer, a second deconvolution layer, an eighteenth convolution layer, an upsampling layer, a nineteenth convolution layer, and a twentieth convolution layer; the first network connection layer, the second deconvolution layer, the eighteenth convolution layer, the upsampling layer, the nineteenth convolution layer, and the twentieth convolution layer are connected in sequence;
[0069] The first network connection layer is connected to the thirteenth convolutional layer and the twenty-sixth dense multi-scale hole convolution fusion layer respectively.
[0070] The number of channels of the eighth convolutional layer is 64, the number of channels of the ninth convolutional layer is 128, the number of channels of the tenth convolutional layer is 128, the number of channels of the eleventh convolutional layer is 256, the number of channels of the self-attention mechanism layer is 256, the number of channels of the twelfth convolutional layer is 256, the number of channels of the thirteenth convolutional layer is 256, the number of channels of the fourteenth convolutional layer is 64, the number of channels of the fifteenth convolutional layer is 128, the number of channels of the sixteenth convolutional layer is 128, the number of channels of the seventeenth convolutional layer is 256, the number of channels of the twenty-sixth dense multi-scale void convolution fusion layer is 256, the number of channels of the first network connection layer is 512, the number of channels of the second deconvolutional layer is 256, the number of channels of the eighteenth convolutional layer is 128, the number of channels of the upsampling layer is 64, the number of channels of the nineteenth convolutional layer is 64, and the number of channels of the twentieth convolutional layer is 3.
[0071] The self-attention mechanism module is introduced below.
[0072] The self-attention layer is used to extract the internal correlation between data and features, obtain global context information to obtain a larger receptive field, and enhance the lack of network semantic information. The self-attention module first converts the image features of the previous layer into Convert to two feature spaces f and g to calculate attention. These two feature spaces are represented as: f(x) = W f x,g(x)=W g x; then calculate: Where: s ij =f(x i ) T g(x i ), β j,i Indicates the degree to which the model pays attention to the i-th layer when generating the j-th layer;
[0073] C is the number of channels, N is the number of feature positions contained in the previous layer, and the output of the attention layer is: Among them, h(x i ) = W h x i , v(x i ) = W v x i ;
[0074] In the above formula, is the learning weight matrix of 1×1 convolution, The module multiplies the output of the attention layer by the scale parameter and adds it to the input image features. Therefore, the output is: y i = γo i + x i ; where γ is a learnable scalar initialized to 0. γ specifies the local information that the network learns first, and then gradually transfers more weights to the learning of non-local information.
[0075] Step 104: Input the picture with the repaired structure into the detail repair network, which includes a dense multi-scale dilated convolution network layer and a self-attention mechanism convolution network layer parallel to it. Connect this parallel convolution layer to the decoder and the transposed convolution network to obtain the picture after detail repair, generate fine textures and enhance global and local semantic consistency.
[0076] Step 105: Obtain the real image.
[0077] Step 106: Train the double-spectrum normalized discriminator network with the real image.
[0078] Step 107: Input the picture after detail repair into the trained double-spectrum normalized discriminator to obtain the final repaired image.
[0079] Input the detail repair result into the double-spectrum normalized discriminator network, and continuously feedback to improve the repair ability of the entire generation network, and output the repaired image. Use spectrum normalization to replace batch normalization (BatchNormalization, BN) in the global-local discriminator network, solve the dependence of batch normalization on the Batchsize size, and stabilize the discriminator training. The double-spectrum normalized discriminator takes the repair result and the original image as the network input. The global discriminator consists of 6 convolutional layers with a convolutional kernel size of 5×5 and a stride of 2. The local discriminator consists of 5 convolutional layers with a convolutional kernel size of 5×5 and a stride of 2. The discriminator all uses the activation function Leaky RelU, fuses the information of the global and local discriminators, passes through the fully connected layer and the Sigmoid activation function to output the result, and uses the GAN loss to measure the gap between the model-repaired image and the original image.
[0080] The trained double-spectrum normalized discriminator network includes:
[0081] Global branch identification layer, local branch identification layer, second network connection layer, third fully connected layer and sigmoid layer.
[0082] The global branch identification layer includes: a twenty-first convolutional layer, a twenty-second convolutional layer, a twenty-third convolutional layer, a twenty-fourth convolutional layer, a twenty-fifth convolutional layer, a twenty-sixth convolutional layer and a first fully connected layer; the twenty-first convolutional layer, the twenty-second convolutional layer, the twenty-third convolutional layer, the twenty-fourth convolutional layer, the twenty-fifth convolutional layer, the twenty-sixth convolutional layer and the first fully connected layer are connected in sequence.
[0083] The local branch identification layer includes: a twenty-seventh convolutional layer, a twenty-eighth convolutional layer, a twenty-ninth convolutional layer, a thirtieth convolutional layer, a thirty-first convolutional layer and a second fully connected layer; the twenty-seventh convolutional layer, the twenty-eighth convolutional layer, the twenty-ninth convolutional layer, the thirtieth convolutional layer, the thirty-first convolutional layer and the second fully connected layer are connected in sequence.
[0084] The number of channels of the 21st convolutional layer is 64, the number of channels of the 22nd convolutional layer is 128, the number of channels of the 23rd convolutional layer is 256, the number of channels of the 24th convolutional layer is 512, the number of channels of the 25th convolutional layer is 512, the number of channels of the 26th convolutional layer is 512, the number of channels of the first fully connected layer is 512, the number of channels of the 27th convolutional layer is 64, the number of channels of the 28th convolutional layer is 128, the number of channels of the 29th convolutional layer is 256, the number of channels of the 30th convolutional layer is 512, the number of channels of the 31st convolutional layer is 512, the number of channels of the second fully connected layer is 512, the number of channels of the second network connection layer is 1024, and the number of channels of the third fully connected layer is 1024.
[0085] The following is a detailed introduction to the dual spectrum normalization discriminator module:
[0086] The discriminator of WGAN uses WassersteinDistance training, which can eliminate the convergence problem that occurs in traditional GAN training and make the training process stable. However, the parameter matrix of the discriminator in WGAN needs to satisfy the Lipschitz constraint. Therefore, WGAN directly restricts the elements in the parameter matrix and prevents them from being greater than a certain value. Although this method can make the parameter matrix of the discriminator satisfy the Lipschitz constraint, it destroys the proportional relationship between the structure of the entire parameter matrix and the parameters while truncating the top. In response to the above problem, it is proposed to use a method that satisfies the Lipschitz condition without destroying the matrix structure - spectral normalization. The discriminator is regarded as a multi-layer network, and the input and output relationship of its nth layer is expressed as:
[0087]
[0088] Where a n (·) is the nonlinear activation function of the network layer, using the ReLU activation function; W l is the network parameter matrix, b l is the bias of the network. For the convenience of derivation, b is omitted l , then the above formula can be written as:
[0089]
[0090] Where D n It is a diagonal matrix, which is used to represent the effect of ReLU. When its input is negative, the diagonal element is 0, otherwise it is 1. Therefore, the input-output relationship of a multi-layer neural network (assuming it is N layers) can be expressed as: f(x) = D N W N …D 1 W 1 X
[0091] The Lipschitz constraint places requirements on the gradient of f(x):
[0092]
[0093] Where W represents the spectral norm of the matrix W, which is defined as:
[0094]
[0095] σ(W) is the maximum singular value of the matrix W. For a diagonal matrix D, σ(D) = max(d 1 ,…,d n ), which is the largest element on the diagonal element. Thus, It can be expressed as:
[0096]
[0097] Because the spectral norm of the ReLU diagonal matrix is at most 1, normalization is performed to satisfy the Lipschitz constraint:
[0098]
[0099] From the above formula, we can see that the constraint of Lipschitz=1 can be satisfied by simply dividing the network parameters of each layer of the network by the spectral norm of the parameter matrix of that layer.
[0100] The dual discriminator information is fused and output through the fully connected layer and the Sigmoid activation function. The Sigmoid activation function is used to encode the nonlinear expression, thereby capturing the nonlinear factors of the data and the role of feature selection.
[0101] The present invention uses the internationally recognized CelebAHQ and Paris StreetView datasets to train the inpainting model. The images in both datasets contain large posture changes, complex backgrounds, and fine textures. The CelebAHQ dataset has 25,000 training sets and 5,000 test sets, which are composed of face images. The Paris StreetView dataset has 14,900 training sets and 100 test sets, which are composed of city street scenes. The proposed algorithm is compared with the Image Inpainting with Learnable Bidirectional Attention Maps (LBAM), the Pluralistic Image Completion (PIC), and the Region Normalization for Image Inpainting (RN) to verify the effectiveness of the proposed algorithm.
[0102] In order to subjectively compare the present invention with other algorithms, experiments were conducted on irregular masks in the above two data sets. Figure 4 As shown, Figure 4 Part (a) in the figure represents the original image. Figure 4 Part (b) in the figure is a defective image with a random mask added, and part 4 (c) in the figure represents the restoration result of the multivariate image restoration algorithm. Figure 4 Part (d) in the figure represents the repair result of the learnable bidirectional attention map repair algorithm, and part (e) in the figure represents the repair result of the regional normalization repair algorithm. Figure 4 Part (f) in the figure represents the restoration result of the restoration method of the present invention. The restoration result of the PIC algorithm has a chaotic texture and poor effect. The structure generated by the LBAM algorithm is relatively complete, but there are artifacts and color differences, and the restoration effect is poor. The structure restored by the RN algorithm is complete, but there are watermarks and distortions, and the restoration effect is average. The restoration result of the method of the present invention not only has a reasonable overall structure, but also has high detail clarity, good granularity, and a good restoration effect. Figure 5 As shown, Figure 5 Part (a) in the figure represents the original image. Figure 5 Part (b) is a defective image with a large irregular mask added. The loss area has rich texture. Figure 5 Part (c) in the figure represents the restoration result of the multivariate image restoration algorithm. Figure 5 Part (d) in the figure represents the repair result of the learnable bidirectional attention map repair algorithm. Figure 5Part (e) in the figure represents the restoration result of the regional normalization restoration algorithm. The restoration result of the PIC algorithm has a strong sense of smearing and no texture details, and the effect is poor. The structure generated by the LBAM algorithm is complete, but there are distortions and blurs in some areas, and the effect is average. The restoration result of the RN algorithm has artifacts and the effect is poor. The restoration result of the method of the present invention ensures the integrity and rationality of the overall structure, restores the rich texture details of the incomplete area, and has a good restoration effect.
[0103] In order to objectively evaluate the performance of the proposed algorithm and the comparison algorithm, the Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM) and L1 loss (MAE) indicators are selected for comparison under the same number of iterations and training sets. As can be seen from the table, the PSNR, SSIM and MAE evaluation indicators of the proposed algorithm are better than those of the comparison algorithm.
[0104] Table 1 Quantitative comparison on the dataset CelebAHQ\Paris_streetview
[0105] Repair Algorithm Peak signal-to-noise ratio (↑) Structural similarity (↑) L1 loss (↓) PIC 18.46\18.34 0.721\0.703 0.0393\0.0445 LBAM 25.25\24.68 0.882\0.821 0.0239\0.0343 RN 22.25\21.76 0.838\0.786 0.0368\0.0402 The present invention 29.80\28.69 0.926\0.837 0.0177\0.0294
[0106] Figure 6 The present invention is a schematic diagram of the structure of an image restoration system based on dense multi-scale fusion according to an embodiment of the present invention, wherein the restoration system comprises:
[0107] A repair network construction module 201 is used to construct a structural repair network;
[0108] The structural repair module 202 is used to input the image to be repaired into the structural repair network to obtain the structurally repaired image;
[0109] A detail restoration network construction module 203 is used to construct a detail restoration network;
[0110] A detail restoration module 204 is used to input the structurally restored image into the detail restoration network to obtain a detail restored image;
[0111] A real image acquisition module 205, used to acquire a real image;
[0112] A training module 206, configured to train a bi-spectral normalized discriminator network using the real image;
[0113] The final image restoration module 207 is used to input the image after detail restoration into the trained bi-spectral normalization discriminator to obtain a final restored image.
[0114] The present invention discloses an image restoration algorithm based on dense multi-scale fusion dilated convolution. First, the damaged image is input into a global structure generation network containing dense multi-scale fusion dilated convolution blocks. Then, the output result of the structure generation network is input into a detail generation network, which contains a layer of dense multi-scale fusion dilated convolution blocks and a parallel convolution network with a self-attention mechanism for capturing global context information. Finally, the restoration result is subjected to an improved dual discriminator to enhance the global and local content consistency and detail features of the restored image. The proposed algorithm is trained and tested on an internationally recognized dataset, and the experimental results show that the proposed algorithm can restore large-area missing images, and the restoration results have smooth boundaries and clear details, meeting visual coherence and authenticity. In terms of the restored visual effect, peak signal-to-noise ratio, structural similarity and average error, it is superior to the three mainstream algorithms compared.
[0115] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0116] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. An image restoration method based on dense multi-scale fusion, It is characterized in that The repair method comprises: Constructing a structure repair network; the structure repair network includes: a first encoding module and a first decoding module; the first encoding module includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer and a 16th dense multi-scale hole convolutional fusion layer; the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer and the 16th dense multi-scale hole convolutional fusion layer are connected in sequence; the first decoding module includes: a fifth convolutional layer, a first deconvolutional layer, a sixth convolutional layer, a first upsampling layer and a seventh convolutional layer; the fifth convolutional layer, the first deconvolutional layer, the sixth convolutional layer, the first upsampling layer and the seventh convolutional layer are connected in sequence; the fifth convolutional layer is also connected to the 16th dense multi-scale hole convolutional fusion layer; Inputting the image to be repaired into the structure repair network to obtain the image after structural repair; Construct a detail restoration network; the detail restoration network includes: a second encoding module and a second decoding module; the second encoding module includes two layers, the first layer includes: an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a self-attention mechanism layer, a twelfth convolutional layer and a thirteenth convolutional layer; the eighth convolutional layer, the ninth convolutional layer, the tenth convolutional layer, the eleventh convolutional layer, the self-attention mechanism layer, the twelfth convolutional layer and the thirteenth convolutional layer are connected in sequence; the second layer includes: a fourteenth convolutional layer, a fifteenth convolutional layer, a sixteenth convolutional layer, a seventeenth convolutional layer and a twenty-sixth dense multi-scale void convolution layer The fourteenth convolution layer, the fifteenth convolution layer, the sixteenth convolution layer, the seventeenth convolution layer and the twenty-sixth dense multi-scale hole convolution fusion layer are connected in sequence; the second decoding module includes: a first network connection layer, a second deconvolution layer, an eighteenth convolution layer, an upsampling layer, a nineteenth convolution layer and a twentieth convolution layer; the first network connection layer, the second deconvolution layer, the eighteenth convolution layer, the upsampling layer, the nineteenth convolution layer and the twentieth convolution layer are connected in sequence; the first network connection layer is respectively connected to the thirteenth convolution layer and the twenty-sixth dense multi-scale hole convolution fusion layer; Inputting the structurally repaired image into the detail repair network to obtain a detail-repaired image; Get real images; Using the real image to train a dual-spectrum normalized discriminator network; Inputting the detailed restored image into a trained bispectral normalized discriminator to obtain a final restored image; The trained dual-spectrum normalized discriminator network includes: a global branch discriminator layer, a local branch discriminator layer, a second network connection layer, a third fully connected layer and a sigmoid layer.
2. The image restoration method based on dense multi-scale fusion according to claim 1, It is characterized in that The global branch discriminant layer includes: the twenty-first convolutional layer, the twenty-second convolutional layer, the twenty-third convolutional layer, the twenty-fourth convolutional layer, the twenty-fifth convolutional layer, the twenty-sixth convolutional layer, and the first fully connected layer; the twenty-first convolutional layer, the twenty-second convolutional layer, the twenty-third convolutional layer, the twenty-fourth convolutional layer, the twenty-fifth convolutional layer, the twenty-sixth convolutional layer, and the first fully connected layer are connected in sequence.
3. The image inpainting method based on dense multi-scale fusion according to claim 2, wherein, the local branch discriminant layer includes: the twenty-seventh convolutional layer, the twenty-eighth convolutional layer, the twenty-ninth convolutional layer, the thirtieth convolutional layer, the thirty-first convolutional layer, and the second fully connected layer; the twenty-seventh convolutional layer, the twenty-eighth convolutional layer, the twenty-ninth convolutional layer, the thirtieth convolutional layer, the thirty-first convolutional layer, and the second fully connected layer are connected in sequence.
4. The image inpainting method based on dense multi-scale fusion according to claim 1, wherein, the number of channels of the first convolutional layer is 64, the number of channels of the second convolutional layer is 128, the number of channels of the third convolutional layer is 128, the number of channels of the fourth convolutional layer is 256, the number of channels of the sixteenth dense multi-scale dilated convolution fusion layer is 256, the number of channels of the fifth convolutional layer is 256, the number of channels of the first deconvolutional layer is 128, the number of channels of the sixth convolutional layer is 128, the number of channels of the first upsampling layer is 64, and the number of channels of the seventh convolutional layer is 3.
5. The image inpainting method based on dense multi-scale fusion according to claim 1, wherein, the number of channels of the eighth convolutional layer is 64, the number of channels of the ninth convolutional layer is 128, the number of channels of the tenth convolutional layer is 128, the number of channels of the eleventh convolutional layer is 256, the number of channels of the self-attention mechanism layer is 256, the number of channels of the twelfth convolutional layer is 256, the number of channels of the thirteenth convolutional layer is 256, the number of channels of the fourteenth convolutional layer is 64, the number of channels of the fifteenth convolutional layer is 128, the number of channels of the sixteenth convolutional layer is 128, the number of channels of the seventeenth convolutional layer is 256, the number of channels of the twenty-sixth dense multi-scale dilated convolution fusion layer is 256, the number of channels of the first network connection layer is 512, the number of channels of the second deconvolutional layer is 256, the number of channels of the eighteenth convolutional layer is 128, the number of channels of the upsampling layer is 64, the number of channels of the nineteenth convolutional layer is 64, and the number of channels of the twentieth convolutional layer is 3.
6. The image inpainting method based on dense multi-scale fusion according to claim 3, wherein, The number of channels of the 21st convolutional layer is 64, the number of channels of the 22nd convolutional layer is 128, the number of channels of the 23rd convolutional layer is 256, the number of channels of the 24th convolutional layer is 512, the number of channels of the 25th convolutional layer is 512, the number of channels of the 26th convolutional layer is 512, the number of channels of the first fully connected layer is 512, the number of channels of the 27th convolutional layer is 64, the number of channels of the 28th convolutional layer is 128, the number of channels of the 29th convolutional layer is 256, the number of channels of the 30th convolutional layer is 512, the number of channels of the 31st convolutional layer is 512, the number of channels of the second fully connected layer is 512, the number of channels of the second network connection layer is 1024, and the number of channels of the third fully connected layer is 1024.
7. An image restoration system based on dense multi-scale fusion, It is characterized in that The restoration system is applied to the image restoration method based on dense multi-scale fusion according to any one of claims 1 to 6, and the image restoration system based on dense multi-scale fusion comprises: Repair network building module, used to build a structural repair network; A structural repair module, used for inputting the image to be repaired into the structural repair network to obtain the structurally repaired image; A detail restoration network construction module is used to construct a detail restoration network; A detail restoration module, used for inputting the structurally restored image into the detail restoration network to obtain a detail restored image; A real image acquisition module, used for acquiring a real image; A training module, used for training a dual-spectrum normalization discriminator network using the real image; The final image restoration module is used to input the image after detail restoration into the trained dual-spectrum normalization discriminator to obtain the final restored image.
Citation Information
Patent Citations
Face image restoration method based on dense expansion convolution self-coding adversarial network
CN110689499A
Face image semantic restoration method based on multi-scale feature fusion
CN113112411A