Image inpainting method and system based on deep learning interactive network

By using an interactive deep learning-based network to generate a target mask image and combining it with an image inpainting model, the problem of low detection accuracy in image inpainting is solved, and an efficient inpainting effect that conforms to human visual perception is achieved.

CN116051392BActive Publication Date: 2026-04-21JINAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JINAN UNIVERSITY
Filing Date
2022-09-16
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing image restoration methods suffer from poor controllability and low accuracy in detecting damaged areas, resulting in unsatisfactory restoration effects and a high failure rate.

Method used

An interactive deep learning-based network is used to obtain the damaged areas of the target image, generate a target mask image using the interactive segmentation network, and input it together with the target image into a pre-trained image inpainting model for inpainting.

Benefits of technology

It achieves accurate detection and repair of damaged areas, and the repair results conform to human visual perception, improving the flexibility and controllability of image restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116051392B_ABST
    Figure CN116051392B_ABST
Patent Text Reader

Abstract

This application relates to an image inpainting method and system based on a deep learning-based interactive network. The method includes: acquiring a target image to be inpainted, the target image having multiple damaged regions; using a deep learning-based interactive segmentation network to process the target image after selecting the damaged regions using a preset method, generating a target mask image corresponding to the target image, the target mask image being used to represent the damaged regions to be inpainted in the target image; jointly inputting the target mask image and the target image into a preset image inpainting model, and outputting the inpainted image. This application solves the problem in related technologies where only considering the inpainting network's inpainting effect during image inpainting easily leads to image inpainting failure and poor image inpainting results. It achieves the beneficial effects of obtaining accurate damaged regions, obtaining images that conform to human visual perception, and having good flexibility and controllability in image inpainting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital image processing technology, and in particular to image inpainting methods and systems based on interactive networks of deep learning. Background Technology

[0002] For ancient murals, the unfavorable wall conditions often lead to various forms of damage, such as fading, discoloration, and peeling. Traditional manual restoration is an irreversible process, risking further irreversible damage. Digital image processing, however, eliminates the need for direct manipulation of the original artwork and allows for adjustments to the restoration based on artistic needs, making non-destructive restoration possible and increasing its flexibility.

[0003] In related technologies, mural image restoration mostly refers to the texture context near the damaged area in the image, filling it based on the surrounding texture; or copying and pasting a template from a template library with the current area. The restoration results of these methods often lack natural transitions and are difficult to conform to human visual perception. Meanwhile, with the maturity of deep learning neural networks, image restoration using neural network-based models can synthesize a smoother and more natural restoration area based on the surrounding texture and structure of a given area.

[0004] In related technologies, conventional image restoration methods all require obtaining the damaged area for virtual restoration. However, the fully automatic damaged area detection in related technologies suffers from poor controllability and low accuracy. As a result, even if the trained image restoration model has a high restoration effect, the image restoration will still fail due to the inaccuracy of the damaged area.

[0005] Currently, no effective solution has been proposed to address the problem that image restoration techniques that only consider the restoration effect of image restoration networks are prone to failure and poor image restoration results. Summary of the Invention

[0006] This application provides an image restoration method, system, and storage medium based on an interactive deep learning network, which at least solves the problem in related technologies that, when only considering the restoration effect of the image restoration network, image restoration is prone to failure and poor image restoration results.

[0007] In a first aspect, embodiments of this application provide an image inpainting method based on an interactive deep learning network, comprising: acquiring a target image to be inpainted, wherein the target image has multiple damaged regions; processing the target image after selecting the damaged regions using a preset method using an interactive deep learning-based segmentation network to generate a target mask image corresponding to the target image, wherein the target mask image is used to characterize the damaged regions to be inpainted in the target image; jointly inputting the target mask image and the target image into a preset image inpainting model, and outputting an image after inpainting, wherein the image inpainting model is trained based on preset sample images, damaged sample images constructed based on the sample images, and mask images corresponding to the damaged sample images.

[0008] Secondly, embodiments of this application provide an image inpainting system based on an interactive deep learning network, comprising:

[0009] The acquisition module is used to acquire the target image to be repaired, wherein the target image has multiple damaged areas;

[0010] The segmentation module is used to process the target image after the damaged area is selected by a preset method using an interactive segmentation network based on deep learning, and generate a target mask map corresponding to the target image, wherein the target mask map is used to characterize the damaged area of ​​the target image to be repaired;

[0011] The repair module is used to jointly input the target mask image and the target image into a preset image repair model and output the repaired image. The image repair model is trained based on preset sample images, damaged sample images constructed based on the sample images, and the mask images corresponding to the damaged sample images.

[0012] Thirdly, embodiments of this application provide an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the image inpainting method based on a deep learning-based interactive network as described in the first aspect.

[0013] Fourthly, embodiments of this application provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the image inpainting method based on a deep learning interactive network as described in the first aspect above.

[0014] Compared to related technologies, the image inpainting method, system, and storage medium based on deep learning interactive networks provided in this application obtain a target image to be repaired, wherein the target image has multiple damaged areas; using a deep learning-based interactive segmentation network, the target image after selecting the damaged areas in a preset manner is processed to generate a target mask image corresponding to the target image, wherein the target mask image is used to represent the damaged areas of the target image to be repaired; the target mask image and the target image are jointly input into a preset image inpainting model, and the repaired image is output. The image inpainting model is trained based on preset sample images, damaged sample images constructed based on the sample images, and the mask images corresponding to the damaged sample images. This solves the problem in related technologies that only consider the repair effect of the image inpainting network when inpainting images, which easily leads to image inpainting failure and poor image inpainting effect. It achieves the beneficial effects of obtaining accurate damaged areas, repairing images that conform to human visual perception, and having good flexibility and controllability in image inpainting.

[0015] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0017] Figure 1 This is a hardware structure block diagram of a terminal for an image restoration method based on a deep learning-based interactive network according to an embodiment of this application;

[0018] Figure 2 This is a flowchart of an image inpainting method based on an interactive deep learning network according to an embodiment of this application;

[0019] Figure 3 This is a grayscale image of the target image to be repaired in this embodiment of the application after binarization.

[0020] Figure 4 This is the target mask image output by the interactive segmentation network in the embodiments of this application;

[0021] Figure 5 It is the target mask image generated by fusing the target image and the target mask image in the embodiments of this application;

[0022] Figure 6 This is a schematic diagram of the image after the target image has been repaired according to an embodiment of this application;

[0023] Figure 7 This is a flowchart of an image inpainting method based on a deep learning-based interactive network according to a preferred embodiment of this application;

[0024] Figure 8 This is a structural block diagram of a deep learning-based interactive image restoration system according to an embodiment of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application. Furthermore, it is understood that although the efforts made in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, modifications to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.

[0026] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0027] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application means two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The terms “first,” “second,” “third,” etc., used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.

[0028] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. Taking running on a terminal as an example, Figure 1 This is a hardware structure block diagram of the terminal for the image inpainting method based on a deep learning-based interactive network, according to an embodiment of this application. Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. Optionally, the terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0029] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the image restoration method based on a deep learning interactive network in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0030] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0031] This embodiment provides an image inpainting method based on an interactive deep learning network running on the aforementioned terminal. Figure 2 This is a flowchart of an image inpainting method based on a deep learning-based interactive network according to an embodiment of this application, as shown below. Figure 2 As shown, the process includes the following steps:

[0032] Step S201: Obtain the target image to be repaired, wherein the target image has multiple damaged areas.

[0033] In this embodiment, the target image includes, but is not limited to, images of cultural heritage obtained through specific devices (e.g., photocopies of murals). The actual cultural heritage image corresponding to the target image is damaged. By restoring and repairing this image, the protection of the cultural heritage is achieved. In this embodiment, the specific method of acquiring the target image to be restored is not limited. Before executing the image restoration method of this embodiment, the corresponding target image to be restored has already been prepared. In this embodiment, the target image to be restored can be obtained by directly reading it from the corresponding stored related cultural heritage images.

[0034] Step S202: Using a deep learning-based interactive segmentation network, the target image after the damaged area is selected in a preset way is processed to generate a target mask map corresponding to the target image. The target mask map is used to represent the damaged area of ​​the target image to be repaired.

[0035] In this embodiment, the deep learning-based interactive segmentation network is a pre-trained segmentation network; specifically, the interactive segmentation network is an interactive segmentation network generated by using the DeepLabV3+ semantic segmentation algorithm with the ResNet as the backbone network and refining f-BRS through backpropagation in multiple feature directions.

[0036] In this embodiment, the damaged areas in the target image can be distinguished by visual inspection. Therefore, the image restoration operator can select the damaged areas using a preset method, such as by clicking on the seed point (the pixel point corresponding to the target image) of the boundary of the visually identified damaged area.

[0037] In this embodiment, after the user selects the damaged area of ​​a given target image by interactively clicking on seed points, the target mask image (precise mask image) of the target image is obtained through an interactive segmentation network based on deep learning. This provides the precise damaged area for subsequent image restoration using an image restoration model, so that the restored image conforms to the results of human visual perception.

[0038] In this embodiment, by providing subjective information interactively (the user's subjective selection of the clicked seed point), the interactive segmentation network outputs accurate damaged areas. Thus, when the inpainting network completes the inpainting, the corresponding target image to be repaired actually completes the repair of the damaged areas, rather than merely verifying the repair effect of the image inpainting model.

[0039] Figure 3 This is a grayscale image of the target image to be repaired in this embodiment of the application after binarization. Figure 4 This is the target mask image output by the interactive segmentation network in the embodiments of this application. Figure 5 This is a target mask image generated by fusing the target image and the target mask image in the embodiments of this application. In some optional embodiments, after obtaining the target image, the user (image restoration operator) selects the damaged area (see reference) by clicking on seed points in the corresponding restoration system. Figure 4 and Figure 5 The white area in the middle is then processed by an interactive segmentation network to binarize the grayscale image (see reference). Figure 3 The target mask image is obtained by segmenting the image as shown in the figure. Figure 4 (As shown).

[0040] Step S203: Input the target mask image and the target image into the preset image inpainting model, and output the inpainted image. The image inpainting model is trained based on preset sample images, damaged sample images constructed based on the sample images, and the mask images corresponding to the damaged sample images.

[0041] In this embodiment, the image restoration model is a pre-trained neural network, that is, a trained neural network model.

[0042] In this embodiment, the image inpainting model includes, but is not limited to, the adversarial edge model EdgeConnect, which includes an edge generator and an image completion network; in this embodiment, after obtaining the target mask image output by the interactive segmentation network (reference... Figure 4 (As shown) After that, it will be based on the target mask and the target image (see reference). Figure 3 As shown in the image, the data input is the target edge map generated by the edge generator. The target mask image is then fused with the target image to generate a target mask image. Next, an image completion network performs image completion processing based on the target edge map and the target mask image, ultimately outputting the repaired image (see reference). Figure 6 (The image shown is the repaired version).

[0043] Through steps S201 to S203 described above, a target image to be repaired is obtained, wherein the target image has multiple damaged areas; an interactive segmentation network based on deep learning is used to process the target image after the damaged areas are selected in a preset manner, generating a target mask image corresponding to the target image, wherein the target mask image is used to represent the damaged areas of the target image to be repaired; the target mask image and the target image are jointly input into a preset image inpainting model, and the repaired image is output. The image inpainting model is trained based on preset sample images, damaged sample images constructed based on the sample images, and the mask images corresponding to the damaged sample images. This solves the problem in related technologies where only the repair effect of the image inpainting network is considered when repairing images, which easily leads to image repair failure and poor image repair effect. It achieves the beneficial effects of obtaining accurate damaged areas, repairing images that conform to human visual perception, and having good flexibility and controllability in image repair.

[0044] It should be noted that, in the embodiments of this application, the specified area (damaged area) of the input target image can be automatically repaired according to the user's specific needs, and the following functions can be achieved: interactively selecting the area to be repaired in the input target image to obtain the target mask image of the specified area, and automatically repairing the selected specified area.

[0045] It should be further explained that, in this embodiment, the repair area (damaged area) can be automatically selected semi-interactively according to the user's subjective needs, thereby determining the precise damaged area; in this embodiment, a semi-interactive segmentation network is used, and the specified area is selected by the user clicking on seed points. Only a little manual operation is needed to determine the precise damaged area; at the same time, in this embodiment, the image restoration model adopts the feature backpropagation thinning scheme f-BRS and the EdgeConnect restoration model. The EdgeConnect restoration model can make the restoration result more realistic and accurate.

[0046] In some embodiments, an interactive segmentation network based on deep learning is used to process the target image after the damaged region has been selected using a preset method, generating a target mask image corresponding to the target image. This is achieved through the following steps:

[0047] Step 21: Preprocess the target image to obtain the initial image. The preprocessing includes at least one of the following: flip enhancement, cropping, and noise reduction.

[0048] In this embodiment, the target image is cropped to a preset size and then enhanced by horizontal and vertical flipping. Additionally, noise reduction processing can be performed on the enhanced image.

[0049] Step 22: Obtain the seed points clicked by the user in the pixels corresponding to the initial image, and generate a segmentation guidance map based on all the clicked seed points. The seed points are used to represent the boundary points corresponding to the damaged areas.

[0050] In this embodiment, before selecting the boundary points of the damaged area by clicking on the seed points, the initial image can be binarized to obtain a grayscale image. Then, the pixel corresponding to the seed point in the grayscale image is clicked to complete the selection of all seed points and generate a segmentation guide map.

[0051] Step 23: Using the segmentation guide image as a guide, use an interactive segmentation network to segment the initial image to obtain the target mask image.

[0052] In this embodiment, the interactive segmentation network displays the segmentation guide map as the segmentation, thereby using the pixels within the segmentation guide map region as the corresponding pixels in the target mask map (see reference). Figure 4 (The white area in the image), and at the same time, after determining the area corresponding to the target mask, the pixels of other areas (refer to the white area in the image) are... Figure 4 The black areas in the image are masked to obtain the complete target mask image.

[0053] The initial image is obtained by preprocessing the target image in the above steps. The preprocessing includes at least one of the following: flipping enhancement, cropping, and noise reduction. Seed points clicked by the user in the corresponding pixels of the initial image are obtained, and a segmentation guidance map is generated based on all clicked seed points. Seed points are used to represent the boundary points corresponding to the damaged areas. Guided by the segmentation guidance map, the initial image is segmented using an interactive segmentation network to obtain a target mask image. This realizes the use of an interactive segmentation network to segment a target mask image that accurately represents the damaged areas from the target image, thereby improving the repair accuracy and further achieving the repair of an image that conforms to human visual perception.

[0054] The training process of the interactive segmentation network in this application embodiment is described below. The training process of the interactive segmentation network includes the following steps:

[0055] Step 31: Collect the SBD dataset, which contains 8498 training images and 2820 test images; crop the training and test images to a size of 320 pixels × 480 pixels, and use horizontal and vertical flipping as enhancement. Before the cropping operation, the image size will be randomly scaled by a factor of 1.25.

[0056] Step 32: Input the SBD dataset collected in Step 31 into the f-BRS interactive segmentation network and train it. The network uses ResNet as the backbone network.

[0057] Step 33: Construct the f-BRS interactive segmentation network, which consists of a distance map fusion (DMF) module and a standard DeepLabV3+ module. The distance map fusion module adaptively fuses RGB images and distance maps, taking the RGB image and two distance maps (one for positive clicks and one for negative clicks) as input. The DMF block processes the 5-channel input using a 1×1 convolution, then uses LeakyReLU as the activation function and outputs a 3-channel tensor, which can be passed to the backbone network pre-trained on the RGB image. The standard DeepLabV3+ module consists of a pre-trained ResNet backbone network and a DeepLabV3+ decoder, and also includes three feature propagation refinement schemes (f-BRS): f-BRS-A optimizes the scale and bias of features after the pre-trained backbone, f-BRS-B optimizes the scale and bias of features after ASPP, and f-BRS-C optimizes the scale and bias of features after the first separable convblock.

[0058] In some embodiments, step 22, generating a segmentation guidance map based on all clicked seed points, includes the following steps:

[0059] Step 41: Connect all seed points in sequence to generate the split boundary.

[0060] In this embodiment, connecting all the clicked seed points generates the corresponding segmentation boundary (see reference). Figure 4 and Figure 5 (The boundary of the white area in the middle).

[0061] Step 42: Using the segmentation boundary, generate a positive segmentation guide map for the pixels in the initial image that are within the segmentation boundary, and generate a negative segmentation guide map for the pixels in the initial image that are outside the segmentation boundary. The segmentation guide map includes a positive segmentation guide map and a negative segmentation guide map.

[0062] By connecting all seed points sequentially in the above steps, a segmentation boundary is generated. Using the segmentation boundary, a positive segmentation guide map is generated for the pixels in the initial image that are within the segmentation boundary, and a negative segmentation guide map is generated for the pixels in the initial image that are outside the segmentation boundary. The segmentation guide map includes both positive and negative segmentation guide maps, thus realizing the generation of distance maps (segmentation guide maps) for positive and negative clicks.

[0063] In some embodiments, step S203, which involves jointly inputting the target mask image and the target image into a preset image inpainting model and outputting the inpainted image, is achieved through the following steps:

[0064] Step 51: Fuse the target image and the target mask image to generate the target mask image.

[0065] Step 52: Binarize the target mask image to generate a first grayscale image corresponding to the target mask image. Then, perform edge detection in the generated first grayscale image using an edge detector corresponding to the edge generator to obtain the first edge image corresponding to the target mask image.

[0066] In this embodiment, the first grayscale image corresponds to... Figure 5 The image shown; in this embodiment, the edge detector includes an edge detector based on the Canny edge detection algorithm.

[0067] Step 53: Input the target mask image, the first grayscale image, and the first edge image into the edge generator to generate the target edge image corresponding to the target image.

[0068] In this embodiment, the target mask image, the first grayscale image, and the first edge image are used as inputs, and edge detection is performed by an edge generator to obtain the target edge image.

[0069] Step 54: Input the target edge image and the target mask image into the image completion network to complete and repair the target mask image through the image completion network, and output the repaired image.

[0070] The target image and target mask image are fused together in the above steps to generate a target mask image. The target mask image is then binarized to generate a first grayscale image corresponding to the target mask image. In the generated first grayscale image, an edge detector corresponding to the edge generator is used to perform edge detection to obtain a first edge image corresponding to the target mask image. The target mask image, the first grayscale image, and the first edge image are then jointly input into the edge generator to generate a target edge image corresponding to the target image. The target edge image and the target mask image are then fed into an image completion network to complete and repair the target mask image. The repaired image is then output, thus completing the repair of the user-interacted segmented selected region in the target image and achieving an image that conforms to human visual perception.

[0071] Figure 7 This is a flowchart of an image inpainting method based on a deep learning-based interactive network according to a preferred embodiment of this application, with reference to... Figure 7 The process includes:

[0072] Step S701: Load the image, then proceed to step S702.

[0073] Step S702: Determine whether the user selected the damaged area of ​​the image by clicking the seed point. If yes, proceed to step S703; otherwise, repeat step S702.

[0074] In this embodiment, once it is determined that the user selects a damaged area of ​​the image by clicking on a seed point, an interactive segmentation network is then used to segment the target mask image.

[0075] Step S703: Save the target mask image. Then, proceed to step S704.

[0076] Step S704: Repair the damaged areas of the image, and then proceed to step S705.

[0077] Step S705: Display the obtained image restoration results.

[0078] The training process of the image restoration model in the embodiments of this application is also described below. The training of the image restoration model specifically includes the following steps:

[0079] Step 1: Collect the publicly available dataset Places2, which contains 10 million images. The image size is 256×256.

[0080] Step 2: Input the dataset collected in Step 1 into the Edge Connect repair network and train it.

[0081] Step 3: Build the Edge Connect model.

[0082] In this embodiment, the Edge Connect model includes an edge generator and an image inpainting network, wherein,

[0083] In the edge generator, let I gt As ground truth images, their edge maps and grayscale counterparts will be represented by C. gt and I gray This indicates that a masked grayscale image is used in the edge generator. As input, its edge map Using an image mask M as a precondition (1 represents the missing region, 0 represents the background), and ⊙ representing the Hadamard product, the generator predicts the edge map of the masked region: In this embodiment, I is used gray The conditional Cgt and Cpred are used as inputs to the discriminator to predict whether the edge map is real.

[0084] In this embodiment, the training objectives of the edge generator include adversarial loss and feature matching loss, and satisfy the following:

[0085]

[0086] Where λ adv,1 and λ FM It is the regularization parameter.

[0087] Adversarial loss is defined as:

[0088]

[0089] Feature matching loss Defined as:

[0090]

[0091] Where L is the last convolutional layer of the discriminator, and N i D1 is the number of elements in the i-th activation layer. (i) It is the activation in the i-th layer of the discriminator; feature matching loss. It is the activation map in the intermediate layer of the comparison discriminator. In this embodiment, the training process is stabilized by forcing the generator to produce a representation that is similar to the real image.

[0092] In this embodiment, the image completion network uses an incomplete color image. As input, use the composite edge map C compAdjustments are made; the composite edge map is constructed by combining the background region of the true edge with the edges generated in the damaged region of the previous stage, i.e., C. comp =C gt ⊙(1-M)+C pred ⊙M, During the image completion process, the image completion network returns a color image. This fills in the missing areas, which have the same resolution as the input image.

[0093] In this embodiment, color image I pred It is trained on a joint loss, which includes L1 loss and adversarial loss. Perceived loss and style loss, among which,

[0094] The L1 loss is determined by normalizing the mask size;

[0095]

[0096]

[0097] Where, φ i It is the activation map of the i-th layer of the pre-trained network, φ i The activation maps corresponding to layers relu1_1, relu2_1, relu3_1, relu4_1, and relu5_1 of the VGG-19 network pre-trained on the ImageNet dataset. Penalize results that are perceptually dissimilar to the labels by defining a distance metric between activation graphs of a pre-trained network.

[0098] The activation maps of layers relu1_1, relu2_1, relu3_1, relu4_1, and relu5_1 in the VGG-19 network are also used to compute the style loss; given a size C j ×H j ×W j The feature map, style loss is calculated by the following formula:

[0099]

[0100] in, It is formed by the activation map φ j Constructed C j ×C j Gram matrix.

[0101] In summary, the overall loss function for the image restoration model is:

[0102]

[0103] This embodiment also provides an interactive image restoration device based on deep learning, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the terms "module," "unit," "subunit," etc., can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0104] Figure 8 This is a structural block diagram of a deep learning-based interactive image restoration system according to an embodiment of this application, as shown below. Figure 8 As shown, the system includes:

[0105] The acquisition module 81 is used to acquire the target image to be repaired, wherein the target image has multiple damaged areas.

[0106] The segmentation module 82, coupled to the acquisition module 81, is used to process the target image after the damaged area is selected in a preset manner using an interactive segmentation network based on deep learning, and generate a target mask map corresponding to the target image, wherein the target mask map is used to characterize the damaged area of ​​the target image to be repaired.

[0107] The repair module 83, coupled to the segmentation module 83, is used to jointly input the target mask image and the target image into a preset image repair model and output the repaired image. The image repair model is trained based on preset sample images, damaged sample images constructed based on the sample images, and the mask images corresponding to the damaged sample images.

[0108] The deep learning-based interactive image restoration system of this application embodiment acquires a target image to be restored, wherein the target image has multiple damaged areas; it uses a deep learning-based interactive segmentation network to process the target image after selecting the damaged areas using a preset method, generating a target mask map corresponding to the target image, wherein the target mask map is used to represent the damaged areas of the target image to be restored; the target mask map and the target image are jointly input into a preset image restoration model, and the restored image is output. The image restoration model is trained based on preset sample images, damaged sample images constructed based on the sample images, and the corresponding mask maps of the damaged sample images. This solves the problem in related technologies where image restoration only considers the restoration effect of the image restoration network, which easily leads to image restoration failure and poor image restoration effect. It achieves the beneficial effects of obtaining accurate damaged areas, restoring images that conform to human visual perception, and having good flexibility and controllability in image restoration.

[0109] In some embodiments, the segmentation module 82 further includes:

[0110] The preprocessing unit is used to preprocess the target image to obtain an initial image, wherein the preprocessing includes at least one of the following: flip enhancement, cropping, and noise reduction;

[0111] The acquisition unit, coupled to the preprocessing unit, is used to acquire the seed points clicked by the user in the pixels corresponding to the initial image, and to generate a segmentation guidance map based on all the clicked seed points, wherein the seed points are used to characterize the boundary points corresponding to the damaged areas.

[0112] The segmentation unit, coupled to the acquisition unit, is used to segment the initial image using an interactive segmentation network, guided by a segmentation guide map, to obtain the target mask image.

[0113] In some embodiments, the acquisition unit is further configured to connect all seed points sequentially to generate a segmentation boundary; and to generate a positive segmentation guide map for the pixels in the initial image that are within the segmentation boundary, and to generate a negative segmentation guide map for the pixels in the initial image that are outside the segmentation boundary, wherein the segmentation guide map includes a positive segmentation guide map and a negative segmentation guide map.

[0114] In some embodiments, the interactive segmentation network is an interactive segmentation network generated by using the DeepLabV3+ semantic segmentation algorithm with the ResNet as the backbone network and refining f-BRS through backpropagation in multiple feature directions.

[0115] In some embodiments, the image inpainting model includes the adversarial edge model Edge Connect, which includes an edge generator and an image completion network.

[0116] In some embodiments, the repair module 83 further includes:

[0117] The fusion unit is used to fuse the target image and the target mask image to generate the target mask image;

[0118] The processing unit, coupled to the coupling unit, is used to perform binarization processing on the target mask image to generate a first grayscale image corresponding to the target mask image. In the generated first grayscale image, edge detection is performed using an edge detector corresponding to the edge generator to obtain a first edge image corresponding to the target mask image.

[0119] The generation unit, coupled to the processing unit, is used to input the target mask image, the first grayscale image, and the first edge image into the edge generator to generate the target edge image corresponding to the target image.

[0120] The completion unit, coupled to the generation unit, is used to input the target edge image and the target mask image into the image completion network, so that the target mask image can be completed and repaired by the image completion network, and the repaired image can be output.

[0121] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.

[0122] This embodiment also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0123] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0124] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0125] S1, Obtain the target image to be repaired. The target image has multiple damaged areas.

[0126] S2 utilizes a deep learning-based interactive segmentation network to process the target image after selecting the damaged area using a preset method, generating a target mask map corresponding to the target image. The target mask map is used to represent the damaged area of ​​the target image to be repaired.

[0127] S3 inputs the target mask and the target image into a preset image restoration model and outputs the restored image.

[0128] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0129] Furthermore, in conjunction with the image inpainting method based on a deep learning interactive network in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the image inpainting methods based on a deep learning interactive network in the above embodiments.

[0130] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0131] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. An image inpainting method based on an interactive deep learning network, characterized in that, include: Obtain the target image to be repaired, wherein the target image has multiple damaged areas; An interactive segmentation network based on deep learning is used to process the target image after the damaged area is selected in a preset manner, and a target mask map corresponding to the target image is generated. The target mask map is used to represent the damaged area of ​​the target image to be repaired. The target mask image and the target image are jointly input into a preset image restoration model, and the restored image is output. The image restoration model is trained based on preset sample images, damaged sample images constructed from the sample images, and mask images corresponding to the damaged sample images. Specifically, an interactive segmentation network based on deep learning is used to process the target image after selecting the damaged region using a preset method, generating a target mask image corresponding to the target image, including: The target image is preprocessed to obtain an initial image, wherein the preprocessing includes at least one of the following: flip enhancement, cropping, and noise reduction; The seed points clicked by the user in the pixels corresponding to the initial image are obtained, and a segmentation guidance map is generated based on all the clicked seed points. The seed points are used to characterize the boundary points corresponding to the damaged areas. Obtaining the seed point clicked by the user in the corresponding pixel of the initial image includes: The initial image is binarized to obtain a grayscale image, and the pixel corresponding to the seed point in the grayscale image is clicked to complete the selection of all the seed points; A segmentation guidance map is generated based on all the clicked seed points, including: Connect all the seed points in sequence to generate the segmentation boundary; Using the segmentation boundary, a positive segmentation guide map is generated for pixels in the initial image that are within the segmentation boundary, and a negative segmentation guide map is generated for pixels in the initial image that are outside the segmentation boundary, wherein the segmentation guide map includes the positive segmentation guide map and the negative segmentation guide map; Guided by the segmentation guide map, the interactive segmentation network is used to segment the initial image to obtain the target mask image. The interactive segmentation network displays the segmentation guide map as the segmentation guide map, takes the pixels in the segmentation guide map area as the pixels corresponding to the target mask image, and after determining the area corresponding to the target mask image, performs masking processing on the pixels in other areas to obtain the complete target mask image. The image inpainting model includes the adversarial edge model EdgeConnect, which comprises an edge generator and an image completion network. The target mask image and the target image are jointly input into the preset image inpainting model, and the inpainted image is output, including: The target image and the target mask image are fused together to generate a target mask image; The target mask image is binarized to generate a first grayscale image corresponding to the target mask image. In the generated first grayscale image, edge detection is performed using an edge detector corresponding to the edge generator to obtain a first edge image corresponding to the target mask image. The target mask image, the first grayscale image, and the first edge image are jointly input into the edge generator to generate the target edge image corresponding to the target image. The target edge image and the target mask image are fed into the image completion network to complete and repair the target mask image, and the repaired image is output.

2. The image restoration method according to claim 1, characterized in that, The interactive segmentation network is generated by using the DeepLabV3+ semantic segmentation algorithm with ResNet as the backbone network and refining f-BRS through backpropagation in multiple feature directions.

3. The image restoration method according to claim 1, characterized in that, The edge detector includes an edge detector based on the Canny edge detection algorithm.

4. An image inpainting system based on an interactive deep learning network, characterized in that, include: The acquisition module is used to acquire the target image to be repaired, wherein the target image has multiple damaged areas; The segmentation module is used to process the target image after selecting the damaged region using a preset method using an interactive segmentation network based on deep learning, generating a target mask image corresponding to the target image, wherein the target mask image is used to represent the damaged region to be repaired in the target image; the segmentation module is also used to preprocess the target image to obtain an initial image, wherein the preprocessing includes at least one of the following: flip enhancement, cropping, and noise reduction; obtain the seed points clicked by the user in the pixels corresponding to the initial image, and generate a segmentation guidance map based on all the clicked seed points, wherein the seed points are used to represent the boundary points corresponding to the damaged region; guided by the segmentation guidance map, segment the initial image using the interactive segmentation network to obtain the target mask image, wherein the interactive segmentation network... Using the segmentation guide image as the segmentation display, pixels within the segmentation guide image region are used as pixels corresponding to the target mask image. After determining the region corresponding to the target mask image, pixels in other regions are masked to obtain the complete target mask image. The segmentation module is also used to binarize the initial image to obtain a grayscale image, and click the pixel corresponding to the seed point in the grayscale image to complete the selection of all seed points. The segmentation module is also used to connect all the seed points in sequence to generate a segmentation boundary. Using the segmentation boundary, pixels in the initial image within the segmentation boundary are used to generate a positive segmentation guide image, and pixels in the initial image outside the segmentation boundary are used to generate a negative segmentation guide image. The segmentation guide image includes the positive segmentation guide image and the negative segmentation guide image. The repair module is used to jointly input the target mask image and the target image into a preset image repair model, and output a repaired image. The image repair model is trained based on preset sample images, damaged sample images constructed based on the sample images, and the mask images corresponding to the damaged sample images. The image repair model includes an adversarial edge model, EdgeConnect, which includes an edge generator and an image completion network. The repair module is also used to fuse the target image and the target mask image to generate a target mask image; and to repair the target image... The target mask image is binarized to generate a first grayscale image corresponding to the target mask image. Edge detection is then performed on the generated first grayscale image using an edge detector corresponding to the edge generator to obtain a first edge image corresponding to the target mask image. The target mask image, the first grayscale image, and the first edge image are then jointly input into the edge generator to generate a target edge image corresponding to the target image. The target edge image and the target mask image are then fed into the image completion network to complete and repair the target mask image, and the repaired image is output.

5. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the steps of the image inpainting method based on an interactive deep learning network as described in any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements step 2 of the image inpainting method based on an interactive deep learning network as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Road image-based shelter coverage area filling method and device

    CN111476213A

  • User real-time smearing interactive image segmentation method based on deep learning

    CN114037712A