Image restoration system and method based on generative adversarial network

CN120339130APending Publication Date: 2025-07-18SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510420287.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

Smart Images

  • Figure CN120339130A_ABST
    Figure CN120339130A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image restoration, in particular to an image restoration system and method based on a generative adversarial network, and the system comprises an image collection module, a preprocessing module, a restoration module and an output module. The image acquisition module is responsible for receiving external image information; the preprocessing module is used for cutting and cutting according to requirements and preparing a to-be-restored image; the restoration module uses a pre-trained generative adversarial network deep learning model to restore the missing part of the image by using an image patch technology through alternate training of a generator and a discriminator; and the output module displays the repaired complete image on a screen. According to the system, through framework innovation, algorithm optimization and module cooperation, the existing technical problems are effectively solved, a high-quality repaired image can be generated, large-area and irregular defects can be processed, and high consistency with an original image is kept. The breakthroughs bring new opportunities to the field of image restoration and are expected to play an important role in the fields of digital content restoration, ancient book restoration, medical image processing and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image restoration, and in particular to an image restoration system and method based on a generative adversarial network. Background Art

[0002] With the rapid development of digital image technology, image restoration has become an important research direction in the field of computer vision and image processing. Traditional image restoration methods mainly rely on manually designed features and rules. Although they can achieve good results in some simple scenes, they are often unable to handle complex and large-area image defects. In recent years, the rise of deep learning technology has brought new opportunities and challenges to image restoration.

[0003] At present, deep learning-based image restoration methods can be roughly divided into two categories: context encoder-based methods and generative adversarial network-based methods. Context encoder-based methods fill in missing areas by learning the context information of the image, but such methods often have difficulty in generating highly realistic details. Although the generative adversarial network-based methods perform well in generating realistic details, they still have shortcomings in maintaining the consistency of the overall structure of the image and dealing with large-area defects.

[0004] Existing image restoration methods based on generative adversarial networks usually adopt a single generator structure, which has limitations when dealing with image features of different scales. In addition, existing methods often have difficulty balancing the authenticity of the generated image and the consistency with the original image during training, resulting in restoration results that are either too blurry or inconsistent with the surrounding environment. Another common problem is that existing methods often have difficulty generating reasonable semantic content and structural information when dealing with large, irregularly shaped defective areas. Summary of the invention

[0005] In view of the above problems, the present invention proposes a new image restoration system and method based on generative adversarial networks. The present invention aims to solve the problems existing in the prior art, such as low image restoration quality, difficulty in handling large-area defects, and inconsistency between the restoration result and the original image, and to provide a technical solution that can generate high-quality, well-structured, and detailed restoration images.

[0006] The present invention proposes an image restoration system based on a generative adversarial network, the system comprising:

[0007] An image acquisition module, used to receive captured images or picture information from an external device;

[0008] An image preprocessing module, which is in communication with the image acquisition module and is used to crop and cut the received captured image or picture information according to user requirements to obtain an image to be repaired;

[0009] A repair module, communicatively connected to the image preprocessing module, for processing the input image to be repaired through a pre-trained deep learning model to obtain a patched complete image;

[0010] An output module, communicatively connected to the repair module, for displaying the repaired complete image on the screen;

[0011] Wherein, the deep learning model is a generative adversarial network, and the generative adversarial network realizes deep learning by putting real data and data generated by a generator into a discriminator for discrimination, and continuously and alternately training the generator and the discriminator, and the generator repairs the missing area in the image through image patches.

[0012] Preferably, the generative adversarial network includes a GAN framework structure and a deep generative neural network architecture;

[0013] The GAN framework structure includes a generator and a discriminator. The input of the generator is an image patch, and the output is the content of the missing area in the image; the input of the discriminator is an entire image, and the output is the probability of judging authenticity;

[0014] The discriminator has multiple convolutional layers, each convolutional layer is followed by a Leaky ReLU layer, and finally a Sigmoid layer is used as the output layer;

[0015] The generator is a U-Net architecture, the input is the missing area of the image, and the output is the filled content;

[0016] The training objective of the GAN framework structure is:

[0017] L = min[max(D(G(E(I)))-D(I)) + λ1||G(E(I)) - I|| + λ2||E(G(E(I)))-E(I)||]

[0018] Wherein, L is the value of the objective function, D represents the discriminator, G represents the generator, E represents the encoder, I represents the input image, and λ1 and λ2 are two constants, respectively representing the weights of the corresponding loss functions;

[0019] During the training process, the generator and the discriminator are alternately trained until the objective function converges.

[0020] Preferably, the deep generative neural network architecture includes:

[0021] An encoder, composed of convolutional operation layers, for gradually reducing the size of the feature map, and the number of channels of the feature map continuously increases to capture the content information in the image;

[0022] The decoder, composed of convolutional operation layers, is used to gradually increase the size of the feature map while continuously reducing the number of channels of the feature map, and outputs the repaired image.

[0023] Preferably, the generative adversarial network processes the image to be repaired through a pre-trained deep learning model to obtain the repaired complete image, specifically including:

[0024] Obtain input features, where the input features include: damaged image features F ∈ R H×W×C ; The damaged image features refer to the image feature information F ∈ R of the damaged image extracted by the convolutional encoder H×W×C , where H, W, and C respectively represent the height, width, and number of channels of the input image;

[0025] The process of generating the repaired image, which includes:

[0026] The image feature processing stage, that is, sending the input features into the encoder for calculation to output the obtained encoded vector, and the encoded vector is used as the initial encoded information in the stage of generating the repaired image. The encoder includes: multiple convolutional layers and a pooling layer, and the obtained encoded vector is input into the decoder, and the decoder includes: multiple transposed convolutional layers;

[0027] The image completion stage, obtaining the repaired image according to the encoded vector, that is, calculating the pixel value at each position, and the repaired image includes: real image and fake image.

[0028] Preferably, in the image completion stage, calculating the pixel value at each position specifically includes:

[0029] Through the cross-connection structure of the encoder-decoder, the pixel value at each position in the image comes from the same feature map, where C1, C2,..., C n represents the feature information of different scales extracted during the image feature processing process.

[0030] Preferably, the convolutional kernel of the transposed convolutional layer keeps the size of the output feature map unchanged through zero padding.

[0031] Preferably, the generator includes:

[0032] The generator encoding module is used to compress and process the image information through convolutional operations and extract image features to obtain feature information;

[0033] The generator decoding module is used to decode the feature information obtained by the compression process and output the repaired image;

[0034] The generator skip connection module is used to map high-dimensional image information to low-dimensional during the encoding and decoding processes of the generator.

[0035] Preferably, the generator encoding module includes:

[0036] An image compression sub-module, configured to compress the input image, reduce the image size while increasing the number of channels, and obtain low-resolution image feature information;

[0037] An image feature extraction sub-module, configured to extract the low-resolution image feature information to obtain image feature information.

[0038] Preferably, the generator decoding module includes:

[0039] An image feature mapping sub-module, configured to perform image feature mapping in the decoder, map high-dimensional image information to low-dimensional, and obtain a feature map;

[0040] An image feature fusion sub-module, configured to perform fusion processing on the feature map to obtain a first feature map;

[0041] A feature map transformation sub-module, configured to transform the first feature map to obtain a second feature map;

[0042] An image completion sub-module, configured to obtain a generated image according to the second feature map.

[0043] A method for an image inpainting system based on a generative adversarial network includes:

[0044] Processing the input image to be inpainted through a pre-trained deep learning model to obtain a patched complete image;

[0045] The deep learning model is a generative adversarial network, and the generative adversarial network realizes deep learning by putting real data and data generated by the generator into the discriminator for discrimination, and continuously and alternately training the generator and the discriminator. The generator repairs the missing area in the image through image patches;

[0046] The generative adversarial network includes a GAN framework structure and a deep generative neural network architecture; the GAN framework structure includes a generator and a discriminator. The input of the generator is an image patch, and the output is the content of the missing area in the image; the input of the discriminator is an entire image, and the output is the probability of judging authenticity;

[0047] The GAN framework structure is the model structure of a deep learning generative adversarial network; the discriminator has multiple convolutional layers, and a Leaky ReLU layer is adopted after each convolutional layer, and finally a Sigmoid layer is used as the output layer; the generator is a U-Net architecture, the input is the missing area of the image, and the output is the filled content;

[0048] The training objective of the GAN framework structure is:

[0049] L = min[max(D(G(E(I)))) - D(I)) + λ1||G(E(I)) - I|| + λ2||E(G(E(I))) - E(I)||]

[0050] Among them, L is the value of the objective function, D represents the discriminator, G represents the generator, E represents the encoder, I represents the input image, and λ1 and λ2 are two constants, representing the weights of the corresponding loss functions respectively;

[0051] During the training process, the generator and the discriminator are alternately trained until the objective function converges;

[0052] The process of generating the inpainting image, the process of generating the inpainting image includes:

[0053] The image feature processing stage, the input features are sent into the encoder for calculation, and the obtained encoded vector is output. The encoded vector is used as the initial encoded information in the inpainting image generation stage. The encoder includes: multiple convolutional layers and one pooling layer. The obtained encoded vector is input into the decoder. The decoder includes: multiple deconvolutional layers;

[0054] The image completion stage, the repaired image is obtained according to the encoded vector, that is, the pixel value of each position is calculated. The repaired image includes: real images and fake images;

[0055] The image completion stage calculates the pixel value of each position, specifically including: through the cross-connection structure of the encoder-decoder, the pixel value of each position in the image comes from the same feature map, where c1, c2,..., c n represents the feature information of different scales extracted during the image feature processing process;

[0056] The image completion stage includes:

[0057] The generator encoding module is used to compress and process the image information through convolution operations and extract image features to obtain feature information;

[0058] The generator decoding module is used to decode the feature information obtained by the compression process and output the repaired image;

[0059] The generator skip connection module is used to map high-dimensional image information to low-dimensional during the encoding and decoding processes of the generator.

[0060] Through innovative system architecture design and algorithm optimization, the present invention has achieved significant technological breakthroughs and improvements in multiple aspects. First of all, the system architecture of multi-module collaborative work proposed by the present invention effectively integrates key links such as image acquisition, preprocessing, restoration, and output, forming a complete and efficient image restoration process. This architecture design not only improves the overall performance of the system but also enhances its flexibility and adaptability in practical applications.

[0061] In the core restoration module, the present invention adopts an innovative generative adversarial network structure. By introducing a deep generative neural network architecture, the system can better capture the multi-scale features of images, thus performing excellently when dealing with defect areas of different sizes and shapes. In particular, the generator of the present invention adopts an encoder-decoder structure and combines skip connection technology. This design can not only effectively extract the high-level semantic information of images but also retain the detailed texture, thus generating more realistic and natural restoration results.

[0062] Another important innovation point is the multi-objective optimization strategy proposed by the present invention. By cleverly designing the loss function, the system can take into account the authenticity of the generated image, the consistency with the original image, and the semantic rationality during the training process. This multi-objective optimization not only improves the quality of the restored image but also enhances the robustness of the system when dealing with different types of images.

[0063] In addition, the present invention has also made innovative improvements in the image preprocessing and postprocessing links. By introducing intelligent cropping and cutting algorithms, the system can more accurately locate and process defect areas. In the postprocessing stage, advanced image enhancement techniques are adopted to further improve the visual quality of the restored image.

[0064] From a microscopic perspective, the present invention has carried out careful design and optimization on each component of the network structure. For example, the Leaky ReLU activation function and batch normalization technology are adopted in the discriminator. These improvements in details greatly improve the training stability and generation ability of the network. At the same time, by introducing advanced technologies such as residual learning and attention mechanism, the system can better process the complex structure and texture information in images.

[0065] In summary, through the innovative design of the system architecture, the optimization and improvement of algorithms, and the collaborative effect among various modules, the present invention effectively solves many problems existing in the prior art. It can not only generate high-quality restored images but also has the ability to process large-area and irregular defects, while ensuring a high degree of consistency between the restoration result and the original image. These technological breakthroughs bring new possibilities to the field of image restoration and are expected to play an important role in multiple fields such as digital content restoration, ancient book restoration, and medical image processing.

[0066] There are several obvious innovation points and advantages:

[0067] 1. Selection and Optimization of Deep Learning Models:

[0068] Application of Generative Adversarial Networks (GANs): Use the GAN framework for image inpainting, and continuously optimize the inpainting effect through the alternating training of the generator and discriminator. This method can generate more realistic and natural image completion results.

[0069] Generator Design of U-Net Architecture: Adopt the U-Net architecture as the generator. This architecture is particularly suitable for tasks that require pixel-level accuracy, such as image segmentation or inpainting. The characteristic of U-Net is its skip connections, which can effectively retain the detailed information in the input image and improve the quality of the output image.

[0070] 2. Design of Objective Function:

[0071] Use of Comprehensive Loss Function: The proposed loss function not only includes the difference between the image generated by the generator and the real image (measured by the $L_1$ norm), but also considers the accuracy of the encoder in extracting the features of the original image. This comprehensive consideration helps to improve the quality of the final image inpainting, ensuring that the generated image is both close to the real data distribution and faithful to the structural information of the original input image.

[0072] Flexibly Adjustable Weight Coefficients: By adjusting the two constants λ1 and λ2, the importance of different parts of the loss can be flexibly balanced according to the requirements of specific application scenarios, thus achieving a better inpainting effect.

[0073] 3. Efficient Image Processing Pipeline:

[0074] Image Preprocessing Module: Allows cropping and cutting of images according to user needs, improving the flexibility and applicability of the system.

[0075] Encoder-Decoder Structure: The system adopts an encoder to gradually reduce the size of the feature map while increasing the number of channels to capture more content information, and the decoder does the opposite, gradually restoring the image resolution. This method effectively improves the efficiency and quality of image inpainting.

[0076] 4. Cross-Layer Connection Mechanism:

[0077] Skip Connections: The skip connections used in the U-Net architecture help to solve the common problem of gradient vanishing in deep networks, and enable the network to obtain deeper semantic information while maintaining high-resolution details, which is particularly important for image inpainting in complex backgrounds.

[0078] 5. Adaptability and Scalability in Practical Applications:

[0079] = Support for multiple convolution operations: Whether it is ordinary convolution in the encoder or transposed convolution (deconvolution) in the decoder, it provides the system with powerful feature extraction and reconstruction capabilities. In particular, the zero-padding technique in the transposed convolution layer ensures the consistency of the output feature map size, enhancing the stability and reliability of the system.

[0080] These innovative points work together, making the GAN-based image inpainting system not only able to efficiently complete the automatic image inpainting task, but also perform excellently in dealing with complex and variable image damage situations, with high practical value and broad application prospects. Brief Description of the Drawings

[0081] Figure 1 is the overall system architecture diagram of the present invention;

[0082] Figure 2 is the internal structure diagram of the inpainting module of the present invention;

[0083] Figure 3 is the structure diagram of the generative adversarial network of the present invention;

[0084] Figure 4 is the structure diagram of the display generator of the present invention; Detailed Description of the Invention

[0085] Please refer to the attached Figures 1-4 , the present invention provides an image inpainting system and method based on a generative adversarial network. The system includes an image acquisition module 1, an image preprocessing module 2, an inpainting module 3, and an output module 4. These modules work together to achieve high-quality image inpainting functions.

[0086] The image acquisition module 1 is used to receive the captured image or picture information from an external device. In a preferred embodiment of the present invention, this module can support multiple image formats, such as JPEG, PNG, BMP, etc., to adapt to different application scenarios.

[0087] The image preprocessing module 2 is communicatively connected to the image acquisition module 1. The main function of this module is to crop and cut the received captured image or picture information according to user requirements to obtain the image to be inpainted. Preferably, this module can also perform operations such as image enhancement and noise removal to improve the subsequent inpainting effect.

[0088] The inpainting module 3 is the core component of the present invention and is communicatively connected to the image preprocessing module 2. This module uses a pre-trained deep learning model to process the input image to be inpainted and obtains the complete inpainted image. The deep learning model adopted by the present invention is a generative adversarial network (GAN) 5, and this network structure performs excellently in image generation and inpainting tasks.

[0089] The working principle of the generative adversarial network 5 is to put the real data and the data generated by the generator into the discriminator for discrimination, and continuously train the generator and the discriminator alternately to achieve deep learning. In the present invention, the generator repairs the missing regions in the image through image patches. This method can effectively learn the high-level semantic features and texture information of the image, so as to achieve high-quality image repair.

[0090] The output module 4 is communicatively connected to the repair module 3 and is responsible for displaying the repaired complete image on the screen. Preferably, this module can also provide functions such as image saving and exporting to meet the different needs of users.

[0091] Furthermore, the generative adversarial network 5 of the present invention includes a GAN framework structure 51 and a deep generative neural network architecture 52. The GAN framework structure 51 includes a generator 511 and a discriminator 512, and these two components play against each other to continuously improve their respective performances.

[0092] The input of the generator 511 is an image patch, and the output is the content of the missing region in the image. In an embodiment of the present invention, the generator 511 adopts a U-Net architecture, and this architecture performs excellently in retaining image details. The input of the U-Net is the missing region of the image, and the output is the filled content.

[0093] The input of the discriminator 512 is an entire image, and the output is the probability of judging authenticity. The discriminator 512 is composed of multiple convolutional layers, and a Leaky ReLU layer is adopted after each convolutional layer, and finally a Sigmoid layer is used as the output layer. This structure can effectively distinguish between real images and generated images, thus prompting the generator to continuously improve its generation ability.

[0094] The training objective of the GAN framework structure 51 of the present invention can be expressed by the following mathematical formula:

[0095]

[0096] Among them, L is the value of the objective function, D represents the discriminator, G represents the generator, E represents the encoder, I represents the input image, and λ1 and λ2 are two constants, representing the weights of the corresponding loss functions respectively. In practical applications, the values of λ1 and λ2 can be adjusted according to specific image inpainting tasks. Generally, the value range of λ1 is from 0.1 to 1, while the value range of λ2 is from 0.01 to 0.1. The adjustment of these parameters is crucial for optimizing the model performance, because they directly affect the quality of the generated image and the stability during the model training process. By reasonably setting the values of λ1 and λ2, the model can achieve the best performance in different application scenarios. For example, in tasks that require high-precision detail inpainting, a higher value of λ1 may be preferred to emphasize pixel-level accuracy; while in scenarios that pay more attention to the overall structural consistency, an appropriate increase in the value of λ2 may be needed to enhance the feature consistency. This flexibility enables the image inpainting system based on the generative adversarial network to adapt to various complex application requirements.

[0097] During the training process, the present invention adopts an alternating training strategy, that is, alternately training the generator and the discriminator until the objective function converges. This training method can effectively balance the performance of the generator and the discriminator and avoid problems such as mode collapse.

[0098] The deep generative neural network architecture 52 includes an encoder 521 and a decoder 522. The encoder 521 is composed of multiple convolutional operation layers, which are used to gradually reduce the size of the feature map while increasing the number of channels of the feature map, so as to capture the content information in the image. In a preferred embodiment of the present invention, the encoder 521 may include 3 to 5 convolutional layers, the convolutional kernel size of each layer is 3x3 or 5x5, and the stride is 2, so that the size of the feature map can be effectively reduced while retaining the image information.

[0099] The decoder 522 is also composed of multiple convolutional operation layers, but its function is to gradually increase the size of the feature map while reducing the number of channels of the feature map, and finally output the inpainted image. In order to better restore the image details, the decoder 522 can adopt transposed convolution (also known as deconvolution) operations, or use a combination of upsampling and convolution.

[0100] This structural design of the present invention enables the system to effectively learn the multi-scale features of the image, thereby achieving high-quality image inpainting. Through the collaborative work of the encoder 521 and the decoder 522, the system can first compress the image into a highly abstract feature representation, and then gradually restore it to a complete image, and the missing areas in the image can be effectively filled during this process.

[0101] The generative adversarial network 5 of the present invention processes the image to be repaired through a pre-trained deep learning model to obtain a complete repaired image. Specifically, this process includes two main stages: obtaining input features and generating a repaired image.

[0102] In the stage of obtaining input features, the system of the present invention processes the damaged image feature F∈R H×W×C . Here, the damaged image feature refers to the image feature information of the damaged image extracted by the convolutional encoder. Among them, H, W, and C respectively represent the height, width, and number of channels of the input image. Preferably, in an embodiment of the present invention, H and W can be 256×256 or 512×512 pixels, depending on specific application requirements and computing resources. C is usually 3 (corresponding to the three RGB channels), but may be more in some special applications.

[0103] The process of generating a repaired image is the core part of the present invention, including an image feature processing stage and an image completion stage. In the image feature processing stage, the system sends the input features into the encoder for calculation, and the output encoded vector is used as the initial encoded information in the stage of generating a repaired image. The encoder of the present invention includes multiple convolutional layers and a pooling layer. Preferably, 3 to 5 convolutional layers can be used, and each convolutional layer is followed by a batch normalization layer and a ReLU activation function to improve the effect of feature extraction. The final pooling layer can choose max pooling or average pooling to further reduce the feature dimension.

[0104] The obtained encoded vector is then input into the decoder. The decoder of the present invention includes multiple transposed convolutional layers. In a preferred embodiment, the decoder can contain the same number of transposed convolutional layers as the encoder to ensure that the output image has the same size as the input image. Each transposed convolutional layer can also be followed by a batch normalization layer and a ReLU activation function.

[0105] In the image completion stage, the system obtains the repaired image according to the encoded vector. This process is actually calculating the pixel value at each position. The repaired image of the present invention includes a real image and a fake image. The real image refers to the undamaged part of the original input image, while the fake image refers to the part generated by the generator to fill the missing area.

[0106] Furthermore, in the image completion stage of the present invention, when calculating the pixel value at each position, a cross-connection structure of the encoder-decoder is adopted. The characteristic of this structure is that the pixel

[0107] value at each position in the image comes from the same feature map. Among them, C1, C2, …, C nIt represents the feature information of different scales extracted by the image feature processing process. This design can effectively retain the multi-scale information of the image, thereby improving the repair quality.

[0108] In a preferred embodiment of the present invention, a skip connection can be adopted to implement this cross-connection structure. Specifically, the output of each layer in the encoder can be directly connected to the corresponding layer in the decoder. The advantage of doing this is that the low-level feature information can be reused during the decoding process, which helps to generate more detailed and accurate image details.

[0109] It should be noted that the convolution kernel of the transposed convolution layer of the present invention keeps the size of the output feature map unchanged through zero padding. This design ensures that the spatial dimension of the feature map remains consistent throughout the encoding-decoding process, facilitating subsequent image reconstruction.

[0110] The generator 6 of the present invention includes a generator encoding module 61, a generator decoding module 62, and a generator skip connection module 63. These three modules work together to jointly complete the image repair task.

[0111] The main function of the generator encoding module 61 is to compress the image information through convolution operations and extract image features to obtain feature information. In an embodiment of the present invention, this module may include multiple convolutional layers, and each convolutional layer is followed by a batch normalization layer and a LeakyReLU activation function. This combination can effectively extract the hierarchical features of the image while avoiding the problem of gradient disappearance. Preferably, a convolutional operation with stride = 2 can be used to implement downsampling, which can reduce the size of the feature map while retaining more spatial information.

[0112] The generator decoding module 62 is responsible for decoding the feature information obtained by the compression process and outputting the repaired image. This module usually includes multiple transposed convolution layers (or a combination of upsampling layers and convolutional layers). After each transposed convolution operation, a batch normalization layer and a ReLU activation function can be added to enhance the non-linear expression ability of the network. Preferably, a Tanh activation function can be used in the last layer to normalize the output value to the range of [-1, 1], which helps to generate a more stable and realistic image.

[0113] During the encoding and decoding processes of the generator, the generator skip connection module 63 maps high-dimensional image information to low-dimensional. The design of this module is inspired by the U-Net structure. By adding direct connections between the encoder and the decoder, the decoder can access the low-level feature information in the encoder. In a preferred embodiment of the present invention, skip connections can be added between each pair of corresponding encoding layers and decoding layers. Such connections can employ simple concatenation operations or weighted summation operations. In this way, the system can better retain the detailed information of the image and improve the restoration quality.

[0114] This generator structure design of the present invention fully considers the characteristics of the image restoration task, can effectively capture the multi-scale features of the image, and retain the detailed information during the reconstruction process. Through the collaborative work of the encoding module, the decoding module, and the skip connection module, the system can generate high-quality restored images and effectively fill the missing areas in the original image.

[0115] The generator encoding module 61 of the present invention further includes an image compression sub-module 611 and an image feature extraction sub-module 612. These two sub-modules work together to achieve effective processing and feature extraction of the input image.

[0116] The main function of the image compression sub-module 611 is to compress the input image, increase the number of channels while reducing the image size, so as to obtain low-resolution image feature information. In a preferred embodiment of the present invention, this sub-module can adopt a multi-layer convolutional network structure. For example, 3 to 5 convolutional layers can be used, and the stride of each convolutional layer is set to 2, so that the image size can be halved in each layer. At the same time, by gradually increasing the number of convolutional kernels layer by layer, the number of channels of the feature map can be increased. Preferably, it can start from 64 channels and double each layer, and finally reach 512 channels. This design can effectively compress the spatial information and extract rich features at the same time.

[0117] The image feature extraction sub-module 612 is responsible for further extracting the low-resolution image feature information to obtain more abstract and high-level image feature information. In an embodiment of the present invention, this sub-module can adopt a residual block structure. The use of residual blocks can effectively alleviate the problem of gradient disappearance in the training of deep networks and can extract more rich feature information at the same time. Preferably, 3 to 5 residual blocks can be used, and each residual block contains two 3x3 convolutional layers, with a batch normalization layer and a ReLU activation function added in the middle. This design can further refine the image features while keeping the size of the feature map unchanged.

[0118] The generator decoding module 62 of the present invention includes an image feature mapping sub-module 621, an image feature fusion sub-module 622, a feature map transformation sub-module 623, and an image completion sub-module 624. The design of these sub-modules aims to gradually transform abstract feature information into specific image content.

[0119] The main function of the image feature mapping sub-module 621 is to perform image feature mapping in the decoder, mapping high-dimensional image information to low-dimensional to obtain a feature map. In a preferred embodiment of the present invention, this sub-module can adopt a deconvolution (transpose convolution) operation. For example, 3 to 5 deconvolution layers can be used, with a stride of 2 for each layer, which can gradually increase the spatial size of the feature map. At the same time, by reducing the number of convolution kernels layer by layer, the number of channels of the feature map can be reduced. This design can effectively transform abstract features into specific spatial information step by step.

[0120] The image feature fusion sub-module 622 is responsible for fusing the feature map to obtain a first feature map. In an embodiment of the present invention, this sub-module can introduce low-level feature information from the encoder using skip connections. Preferably, an attention mechanism can be adopted to dynamically determine which features should be focused on. For example, a combination of channel attention and spatial attention can be used, which can achieve selective enhancement of features in both the channel dimension and the spatial dimension.

[0121] The role of the feature map transformation sub-module 623 is to transform the first feature map to obtain a second feature map. In a preferred embodiment of the present invention, this sub-module can adopt the method of multi-scale convolution. Specifically, convolution kernels of different sizes (such as 3x3, 5x5, and 7x7) can be used to process the feature map in parallel, and then the results are fused. This design can capture image features of different scales, which helps to generate more detailed and realistic image content.

[0122] The image completion sub-module 624 is responsible for obtaining the generated image based on the second feature map. In an embodiment of the present invention, this sub-module can include one or more convolution layers, and the number of convolution kernels in the last layer is the same as the number of image channels (usually 3, corresponding to the three RGB channels). Preferably, the Tanh activation function can be used in the last layer to normalize the output value to the range of [-1, 1]. This can ensure that the pixel values of the generated image are within a reasonable range, which helps to generate a more natural-looking image visually.

[0123] The present invention also provides a method for an image inpainting system based on a generative adversarial network. The core idea of this method is to use a pre-trained deep learning model, especially a generative adversarial network, to process the input image to be inpainted, so as to obtain a complete inpainted image.

[0124] In a preferred embodiment of the present invention, the method first uses the generator in the generative adversarial network to preliminarily inpaint the input image. Through an encoder-decoder structure, the generator encodes the input damaged image into a latent feature vector, and then decodes it to generate the inpainted image. In this process, the generator tries to fill in the missing areas in the image and generate content that looks reasonable and coherent.

[0125] The generated image is fed into the discriminator for evaluation. The task of the discriminator is to distinguish the generated image from the real complete image. Through this adversarial training method, the generator is forced to continuously improve its generation ability to produce more realistic and natural inpainting results.

[0126] The method of the present invention also includes an important training objective, which can be expressed as the following mathematical formula:

[0127]

[0128] In this formula, L represents the overall loss function, G, D, and E respectively represent the generator, discriminator, and encoder, and I is the input image. λ1 and λ2 are weight coefficients used to balance the importance of different loss terms.

[0129] Specifically, the term D(G(E(I)))-D(I) represents the adversarial loss, which prompts the generator to generate images that can deceive the discriminator. The term ||G(E(I))-I|| is the pixel-level reconstruction loss, ensuring that the generated image is similar to the original image at the pixel level.

[0130] The term ||E(G(E(I)))-E(I)|| is the feature-level reconstruction loss, which encourages the generated image to be similar to the original image in the high-level feature space. In practical applications, the values of λ1 and λ2 need to be adjusted according to specific tasks. Generally, λ1 can be set between 0.1 and 1, while λ2 can be set between 0.01 and 0.1. The selection of these values will affect the quality and style of the inpainting results, and the optimal values need to be determined through experiments.

[0131] The method of the present invention optimizes the above objective function by alternately training the generator and the discriminator. In each round of training, first fix the discriminator parameters and update the generator parameters to minimize the loss function; then fix the generator parameters and update the discriminator parameters to maximize the discrimination ability. This alternating training method can effectively balance the performance of the generator and the discriminator and avoid problems such as mode collapse during the training process.

[0132] Through this method, the system of the present invention can learn the high-level semantic features and texture information of the image, thereby achieving high-quality image restoration. This method can not only effectively fill in the missing areas in the image, but also maintain the consistency and natural transition between the restored area and the surrounding content, providing a powerful and flexible solution for various image restoration applications.

[0133] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image inpainting system based on a generative adversarial network, characterized in that, The repair system includes: An image acquisition module, which is used to receive captured images or picture information from external devices; An image preprocessing module, which is communicatively connected to the image acquisition module and is used to crop and cut the received captured images or picture information according to user requirements to obtain the image to be repaired; A repair module, which is communicatively connected to the image preprocessing module and is used to process the input image to be repaired through a pre-trained deep learning model to obtain a patched complete image; An output module, which is communicatively connected to the repair module and is used to display the repaired complete image on the screen; Among them, the deep learning model is a generative adversarial network. The generative adversarial network realizes deep learning by putting real data and the data generated by the generator into the discriminator for discrimination, and continuously and alternately training the generator and the discriminator. The generator repairs the missing areas in the image through image patches.

2. The image inpainting system based on a generative adversarial network according to claim 1, characterized in that, The generative adversarial network includes a GAN framework structure and a deep generative neural network architecture; The GAN framework structure includes a generator and a discriminator. The input of the generator is an image patch, and the output is the content of the missing area in the image; the input of the discriminator is an entire image, and the output is the probability of judging authenticity; The discriminator has multiple convolutional layers, and a Leaky ReLU layer is adopted after each convolutional layer, and finally a Sigmoid layer is used as the output layer; The generator is a U-Net architecture, the input is the missing area of the image, and the output is the filled content; The training objective of the GAN framework structure is: L = min[max(D(G(E(I)))-D(I))+λ1||G(E(I))-I||+λ2||E(G(E(I)))-E(I)||], where L is the value of the objective function, D represents the discriminator, G represents the generator, E represents the encoder, I represents the input image, and λ1 and λ2 are two constants, respectively representing the weights of the corresponding loss functions; During the training process, the generator and the discriminator are alternately trained until the objective function converges.

3. The image inpainting system based on a generative adversarial network according to claim 2, wherein The deep generative neural network architecture includes: An encoder, which is composed of convolutional operation layers and is used to gradually reduce the size of the feature map, and the number of channels of the feature map continuously increases to capture the content information in the image; A decoder, which is composed of convolutional operation layers and is used to gradually increase the size of the feature map, and the number of channels of the feature map continuously decreases to output the repaired image.

4. The image inpainting system based on a generative adversarial network according to claim 2, wherein The generative adversarial network processes the image to be repaired through a pre-trained deep learning model to obtain a patched complete image, specifically including: Obtain input features, where the input features include: damaged image feature F ∈ R H×W×C ; The damaged image feature refers to the image feature information F ∈ R of the damaged image extracted by the convolutional encoder H×W×C , where H, W, and C respectively represent the height, width, and number of channels of the input image; The process of generating a patched image, where the process of generating a patched image includes: The image feature processing stage, that is, sending the input features into the encoder for calculation to output the obtained encoded vector. The encoded vector is used as the initialization encoded information in the stage of generating a patched image. The encoder includes: multiple convolutional layers and a pooling layer, and inputting the obtained encoded vector into the decoder. The decoder includes: multiple deconvolutional layers; The image completion stage, obtaining the repaired image according to the encoded vector, that is, calculating the pixel value at each position. The repaired image includes: a real image and a fake image.

5. The image inpainting system based on a generative adversarial network according to claim 4, wherein In the image completion stage, the pixel value at each position is calculated, specifically including: Through the cross-connection structure of the encoder-decoder, the pixel values at each position in the image come from the same feature map, where C1, C2, ..., C n represent the feature information of different scales extracted during the image feature processing process.

6. The image inpainting system based on a generative adversarial network according to claim 4, wherein The convolution kernel of the deconvolution layer keeps the size of the output feature map unchanged through zero padding.

7. The image inpainting system based on the generative adversarial network according to claim 1, characterized in that, The generator includes: A generator encoding module for compressing image information through convolution operations, extracting image features, and obtaining feature information. A generator decoding module for decoding the feature information obtained from the compression process and outputting the repaired image. A generator skip connection module for mapping high-dimensional image information to low-dimensional during the encoding and decoding processes of the generator.

8. The image inpainting system based on a generative adversarial network according to claim 7, wherein The generator encoding module includes: An image compression sub-module for compressing the input image, reducing the image size while increasing the number of channels to obtain low-resolution image feature information. An image feature extraction sub-module for extracting the low-resolution image feature information to obtain image feature information.

9. The image inpainting system based on a generative adversarial network according to claim 7, characterized in that The generator decoding module includes: An image feature mapping sub-module for performing image feature mapping in the decoder, mapping high-dimensional image information to low-dimensional to obtain a feature map. An image feature fusion sub-module for fusing the feature map to obtain a first feature map. A feature map transformation sub-module for transforming the first feature map to obtain a second feature map. An image completion sub-module for obtaining the generated image based on the second feature map.

10. Method of an image inpainting system based on a generative adversarial network, based on the system according to any one of claims 1-9, characterized in that, Including: Processing the input image to be repaired through a pre-trained deep learning model to obtain the patched complete image. The deep learning model is a generative adversarial network. The generative adversarial network realizes deep learning by putting real data and the data generated by the generator into the discriminator for discrimination, and continuously and alternately training the generator and the discriminator. The generator repairs the missing area in the image through image patches. The generative adversarial network includes a GAN framework structure and a deep generative neural network architecture; the GAN framework structure includes a generator and a discriminator. The input of the generator is an image patch, and the output is the content of the missing area in the image. The input of the discriminator is an entire image, and the output is the probability of judging authenticity. The GAN framework structure is the model structure of the deep learning generative adversarial network; the discriminator has multiple convolutional layers, and each convolutional layer is followed by a Leaky ReLU layer, and finally a Sigmoid layer is used as the output layer; the generator is a U-Net architecture, the input is the missing area of the image, and the output is the filled content. The training objective of the GAN framework structure is: L = min[max(D(G(E(I)))-D(I)) + λ1||G(E(I)) - I|| + λ2||E(G(E(I)))-E(I)||] Where L is the value of the objective function, D represents the discriminator, G represents the generator, E represents the encoder, I represents the input image, and λ1 and λ2 are two constants, respectively representing the weights of the corresponding loss functions. During the training process, the generator and the discriminator are alternately trained until the objective function converges. The process of generating the patched image, the process of generating the patched image includes: In the image feature processing stage, the input features are sent into the encoder for calculation, and the resulting encoded vectors are output. The encoded vectors are used as the initial encoded information in the patched image generation stage. The encoder includes: a plurality of convolutional layers and a pooling layer. The obtained encoded vectors are input into the decoder, and the decoder includes: a plurality of deconvolutional layers; In the image completion stage, the restored image is obtained according to the encoded vectors, that is, the pixel value at each position is calculated. The restored image includes: a real image and a fake image; In the image completion stage, calculating the pixel value at each position specifically includes: through the cross-connection structure of the encoder-decoder, the pixel value at each position in the image comes from the same feature map, where c1, c2,..., cn represent the feature information of different scales extracted during the image feature processing process; The image completion stage includes: A generator encoding module for compressing the image information through convolution operations and extracting image features to obtain feature information; A generator decoding module for decoding the feature information obtained by the compression process and outputting the restored image; A generator skip connection module for mapping high-dimensional image information to low-dimensional during the encoding and decoding processes of the generator.