Generating modified digital images using a deep vision-guided patch matching model for image inpainting
By combining deep vision-guided and patch-matching models with generator neural networks or teacher-student network frameworks, the problem of insufficient accuracy and efficiency in high-resolution image restoration by traditional systems is solved, achieving efficient and flexible image restoration results.
Patent Information
- Application Number
- CN202210202058.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-15
- Filing Date
- 2022-03-03
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2042-03-03
Smart Images

Figure CN115082329B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Exemplary embodiments of the present disclosure generally relate to the field of computers, and in particular, to systems, methods, and non-transitory computer-readable media for generating modified digital images using a deep visual-guided patch matching model for image inpainting. BACKGROUND
[0002] In recent years, there has been significant development in software and hardware platforms for digital image inpainting to reconstruct missing or defective regions of digital images. In practice, some digital image editing applications utilize inpainting functionality to remove unwanted objects or distracting elements from digital images and automatically fill in the removed regions of pixels with reliable results. For example, many digital image editing systems can utilize a patch-based approach to borrow example pixels from other portions of the digital image to fill in the defective regions. Other digital image editing systems fill in regions of digital images by implementing a learning-based deep network to learn natural image distribution by training on large datasets. Despite these advances, conventional digital image editing systems still face many obstacles or disadvantages, particularly in terms of accuracy, efficiency, and flexibility. SUMMARY
[0003] One or more embodiments described herein provide benefits and address one or more of the foregoing or other issues in the art by utilizing systems, methods, and non-transitory computer-readable media that accurately, efficiently, and flexibly generate modified digital images using a guided inpainting approach. Specifically, in one or more embodiments, the disclosed systems implement a hybrid guided patch matching model that implements a patch-based deep network approach in a unique digital image processing pipeline. Specifically, the disclosed systems combine the high-quality texture synthesis capabilities of patch-based approaches with the image semantic understanding capabilities of deep network approaches. In some embodiments, the disclosed systems automatically generate guidance maps to assist in identifying replacement pixels for inpainting regions of digital images. For example, the disclosed systems generate guidance maps in the form of structure maps, depth maps, segmentation maps (or other visual guidance). The disclosed systems can generate these guidance maps using an inpainting neural network in conjunction with a visual guidance algorithm or by utilizing a separate visual guidance algorithm such as a generator neural network or a teacher-student network architecture. Furthermore, in some embodiments, the disclosed systems implement a patch matching model to identify replacement pixels for filling in regions of digital images according to these deep visual guidance. By utilizing deep visual guidance and a patch matching model, the disclosed systems can accurately, efficiently, and flexibly generate realistic modified digital images of various resolutions.
[0004] Additional features and advantages of one or more embodiments of the present disclosure are set forth in the description that follows, and in part will be obvious from the description, or can be learned by practice of such an exemplary embodiment. BRIEF DESCRIPTION OF DRAWINGS
[0005] One or more embodiments of the present disclosure are described in additional specificity and detail by reference to the drawings, in which:
[0006] Figure 1 An example system environment in which a guided inpainting system operates is shown in accordance with one or more embodiments;
[0007] Figure 2 An overview of generating modified digital images by utilizing patch matching models and deep visual guidance to inpaint one or more regions of an input digital image is shown in accordance with one or more embodiments;
[0008] Figure 3 An example process of generating modified digital images by utilizing deep visual guidance generated from inpainting digital images is shown in accordance with one or more embodiments;
[0009] Figure 4 An example process of utilizing a generator neural network to generate deep visual guidance from an input digital image is shown in accordance with one or more embodiments;
[0010] Figure 5 An example process of utilizing a teacher-student neural network framework to generate deep visual guidance from an input digital image is shown in accordance with one or more embodiments;
[0011] Figure 6 Utilizing multiple deep visual guidance and patch matching models to identify replacement pixels is shown in accordance with one or more embodiments;
[0012] Figure 7 A comparison of modified digital images generated by a traditional patch matching system and a structure guided inpainting system is shown in accordance with one or more embodiments;
[0013] Figure 8 A comparison of modified digital images generated by a traditional patch matching system and a deep guided inpainting system is shown in accordance with one or more embodiments;
[0014] Figure 9 A comparison of modified digital images generated by a traditional patch matching system and a segmentation guided inpainting system is shown in accordance with one or more embodiments;
[0015] Figure 10 A schematic diagram of a guided inpainting system is shown in accordance with one or more embodiments;
[0016] Figure 11 A flow diagram illustrating a series of actions for generating a modified digital image by identifying replacement pixels with a guided patch matching model is shown in accordance with one or more embodiments; and
[0017] Figure 12 A block diagram illustrating an example computing device in accordance with one or more embodiments is shown. DETAILED DESCRIPTION
[0018] One or more embodiments described herein include a guided inpainting system that accurately, efficiently, and flexibly generates modified digital images with a guided inpainting method. Specifically, in one or more embodiments, the guided inpainting system generates deep visual guidance that informs a patch matching model to identify replacement pixels for inpainting regions of a digital image. To generate the deep visual guidance, the guided inpainting system utilizes a visual guidance algorithm, such as a segmentation image neural network, an image depth neural network, a structure image model, or a combination of two or more of the foregoing. Indeed, in some implementations, the guided inpainting system generates one or more of structure image guidance, image depth guidance, or segmentation image guidance from a digital image. Moreover, the guided inpainting system implements a patch matching model to identify replacement pixels indicated by the deep visual guidance and uses the replacement pixels to inpaint regions of the digital image. By utilizing the deep visual guidance and the patch matching model, the guided inpainting system can accurately, efficiently, and flexibly generate realistic modified digital images of nearly any resolution.
[0019] As described above, in one or more embodiments, the guided inpainting system utilizes deep visual guidance and a patch matching model to inpaint missing, blurred, or otherwise undesirable regions of a digital image. For example, the guided inpainting system utilizes a visual guidance algorithm to generate deep visual guidance for identifying replacement pixels to fill regions of a digital image. In some cases, the guided inpainting system receives a request to edit a digital image that includes missing or undesirable pixels in one or more regions. In some embodiments, based on the request, the guided inpainting system generates an inpainted version of the digital image with a pre-trained inpainting neural network. For example, the guided inpainting system processes the digital image with the inpainting neural network and fills the regions of pixels to be replaced with an initial set of replacement pixels. In certain cases, the inpainted digital image is a lower resolution version of the digital image and the initial set of replacement pixels is a preliminary replacement of the region of the digital image.
[0020] As described above, in certain embodiments, the guided inpainting system generates deep visual guidance to assist in accurately filling regions of a digital image. In certain cases, the guided inpainting system generates the deep visual guidance from a preliminary inpainted digital image. In other cases, the guided inpainting system generates the deep visual guidance directly from a digital image having missing or undesirable regions.
[0021] For example, to generate depth visual guidance from the inpainted digital image, the guided inpainting system utilizes a visual guidance algorithm. More specifically, in one or more embodiments, the guided inpainting system utilizes a visual guidance algorithm, such as one or more of a structure image model, an image depth neural network, or a segmentation image neural network. For example, the guided inpainting system utilizes a structure image model to generate depth visual guidance in the form of structure image guidance that indicates one or more structures within the inpainted digital image. In some embodiments, the guided inpainting system utilizes an image depth neural network to generate depth visual guidance in the form of image depth guidance that indicates different depths within the inpainted digital image. In these or other embodiments, the guided inpainting system utilizes a segmentation image neural network to generate depth visual guidance in the form of segmentation image guidance that indicates different semantic segmentations within the inpainted digital image.
[0022] As noted above, in some embodiments, the guided image inpainting system generates depth visual guidance directly from an initial digital image having missing or undesirable regions. For example, the guided inpainting system utilizes one or more of a generator neural network or a teacher-student neural network framework. In particular, in one or more embodiments, the guided inpainting system utilizes a generator neural network that includes an encoder neural network and a decoder neural network to generate depth visual guidance by predicting structures within the missing or undesirable region(s). For example, the guided inpainting system utilizes the generator neural network to predict structures within the region by processing one or more of the initial digital image, an intermediate digital image that indicates structures outside the region, and / or a binary mask that indicates the region.
[0023] In one or more embodiments that utilize a teacher-student neural network framework, the guided inpainting system generates depth visual guidance by utilizing a student neural network to predict one or more of depth or segmentation within a region of pixels to be replaced. For example, the guided inpainting system utilizes the student neural network to predict depth or segmentation from parameters learned from a teacher neural network. In practice, in some cases, the guided inpainting system generates labels for a complete digital image by utilizing the teacher neural network and learns the parameters by utilizing the student neural network to process an incomplete digital image (e.g., a version of the same digital image processed by the teacher neural network but having one or more missing or undesirable regions of pixels). After training, the guided inpainting system can utilize the student neural network to generate digital visual guidance (e.g., structure image or depth map) for digital images that include digital images having holes or other regions of replacement.
[0024] In some cases, the guided repair system combines multiple types or variations of depth visual guidance in generating the modified digital image. For example, the guided repair system assigns weights to structural image guidance (or structural image model), image depth guidance (or image depth neural network), and segmentation image guidance (or segmentation image neural network) to combine (one or more of the above) with a patch matching model for identifying replacement pixels from the digital image. In one or more embodiments, based on generating the depth visual guidance, the guided repair system identifies replacement pixels for filling or repairing a region of the digital image having missing or undesirable pixels. More specifically, the guided repair system utilizes the depth visual guidance and the patch matching model to identify replacement pixels within the initial digital image (or within another digital image). For example, the guided repair system utilizes the structural image guidance to identify pixels within the digital image having a structure that corresponds to a structure of a region of pixels to be replaced. As another example, the guided repair system utilizes the image depth guidance to identify pixels having a depth that corresponds to a depth of the region. As yet another example, the guided repair system utilizes the segmentation image guidance to identify pixels having a semantic segmentation that corresponds to a semantic segmentation of the region.
[0025] In some embodiments, the guided repair system utilizes a cost function of the patch matching model to identify and select replacement pixels from the digital image. For example, the guided repair system utilizes a cost function that determines a distance between a pixel to be replaced and a potential replacement pixel. In certain cases, to combine one or more variations of the depth visual guidance with the patch matching model in the replacement pixel identification and selection process, the guided repair system modifies the cost function. For example, the guided repair system modifies the cost function based on an initial patch matching cost function and a weighted combination of one or more of the structural image guidance, the image depth guidance, and / or the segmentation image guidance (with their respective weights). Thus, the guided repair system identifies replacement pixels based on the patch matching model and one or more variations of the depth visual guidance.
[0026] As described above, in some embodiments, the guided repair system generates a modified digital image by replacing or repairing a region of the digital image having missing or undesirable pixels. For example, the guided repair system utilizes the patch matching model to generate the modified digital image by replacing the region of the digital image with replacement pixels identified via the depth visual guidance. In particular, the guided repair system can guide the patch model to identify replacement pixels for the region based on segments, structures, and / or depths reflected within the depth visual guidance.
[0027] As described above, conventional digital image editing systems exhibit a number of shortcomings, particularly in terms of accuracy, efficiency, and flexibility. By way of illustration, many conventional systems inaccurately modify digital images. For example, digital image editing systems that utilize conventional patch-based approaches often lack semantic or geometric understanding of the content of the digital image. As a result, these conventional systems often inaccurately select visually discordant or misplaced replacement pixels to fill in regions of the digital image. Moreover, conventional systems that utilize neural network approaches often generate digital images that lack realistic texture detail.
[0028] Furthermore, many conventional digital image editing systems are also inefficient. In particular, to generate accurate digital images, conventional systems often require digital image editing systems to require a large amount of user interaction, time, and corresponding computational resources (e.g., processing power and memory). For example, client devices often provide various user interfaces and interactive tools to correct errors in the regions. To illustrate, conventional systems often require user interaction to identify color-guided regions outside of holes within the digital image. Moreover, correcting errors produced by inpainting algorithms often requires client devices to scale, pan, select, and iteratively apply various algorithms to accurately fill in regions within the digital image.
[0029] Furthermore, conventional digital image editing systems are also inflexible. For example, as described above, conventional systems often cannot accurately render digital images that require semantic awareness across different features of the digital image as well as texture consistency. Moreover, due to the significant memory limitations of many conventional systems that utilize neural network approaches, such systems are often strictly limited to inpainting low resolution digital images. Thus, conventional systems often cannot handle high quality textures in digital images that have a resolution that exceeds 1K.
[0030] As described above, embodiments of the guided inpainting system can provide a number of advantages over conventional digital image editing systems. For example, embodiments of the guided inpainting system can provide higher accuracy than conventional systems. While many conventional systems inaccurately modify digital images by copying and pasting incorrect or misplaced pixels to fill in regions of the digital image, the guided inpainting system utilizes a guided approach that considers the semantics and texture of the digital image. Indeed, certain embodiments of the guided inpainting system utilize deep visual guidance that considers different structures, depths, and / or segmentations of the digital image to accurately select replacement pixels that more consistently fit the missing or undesirable regions.
[0031] In comparison to many traditional digital image editing systems, the guided inpainting system can also improve computational efficiency. In fact, by utilizing the inpainting neural network / visual guidance algorithm and patch matching model to generate the digital image and fill in the region, the guided inpainting system can avoid user interfaces, interactions, tools, and algorithms for identifying color regions within or outside the fill region. Moreover, with a more accurate pixel representation, the guided inpainting system can significantly reduce user interface interactions, time, and corresponding computational resources in correcting the fill pixels. In fact, the guided inpainting system can automatically fill in a region of a digital image with a single user interface, minimal user interactions (e.g., a single click of a button), and significantly reduced time (e.g., in seconds).
[0032] Furthermore, embodiments of the guided inpainting system further improve flexibility in comparison to traditional digital image editing systems. First, the guided inpainting system can generate accurate digital images that reflect semantic and texture features of the digital image. Moreover, some embodiments of the guided inpainting system are suitable for generating high-quality, photo-realistic textures at higher resolutions than many traditional systems. In fact, traditional systems utilizing neural networks are often fixed to generate digital images at only lower resolutions (e.g., less than 1K). In contrast, the guided inpainting system uses deep visual guidance to process low-resolution digital images, and a patch matching model that can generate accurate output digital images at various resolutions (e.g., high resolution or low resolution). Thus, the guided inpainting system can be flexibly adapted to generate modified digital images at almost any resolution, even high resolutions typically used in contemporary digital images.
[0033] In summary, traditional systems utilizing patch-based methods lack semantic understanding, while traditional systems utilizing neural network methods lack realistic textures and cannot operate on high-resolution images. In contrast, the guided inpainting system provides high-quality texture synthesis with image semantic understanding in efficiently, accurately, and flexibly generating digital images at various different resolutions.
[0034] Additional details regarding the guided inpainting system will now be provided with reference to the accompanying drawings. For example, Figure 1 A diagram illustrating an example system environment for implementing the guided inpainting system 102, in accordance with one or more embodiments, is shown. Reference is made to Figure 1 An overview of the guided inpainting system 102 is described. Thereafter, more detailed descriptions of components and processes of the guided inpainting system 102 are provided with reference to subsequent drawings.
[0035] As shown, the environment includes server(s) 104, client devices 108, database 112, and network 114. Each component of the environment communicates via the network 114, and the network 114 is any suitable network over which computing devices communicate. Reference is made to Figure 12The example network will be discussed in more detail.
[0036] As described above, this environment includes client device 108. Client device 108 is a computing device among various computing devices, including smartphones, tablets, smart TVs, desktop computers, laptops, virtual reality devices, augmented reality devices, or reference devices. Figure 12 Another computing device described. Although Figure 1 A single client device 108 is shown, but in some embodiments, the environment includes multiple different client devices, each associated with a different user (e.g., a digital image editor). Client device 108 communicates with servers(s)104 via network 114. For example, client device 108 receives user input from a user interacting with client device 108 (e.g., via client application 110) to edit or modify a digital image, for example, by filling or replacing pixels in one or more regions of the digital image. In some cases, client device 108 receives user input via client application 110 to generate and / or execute digital content operation sequences. Thus, boot repair system 102 on servers(s)104 receives information or instructions to generate modified digital content items using one or more digital content editing operations stored in database 112.
[0037] As shown in the figure, client device 108 includes client application 110. Specifically, client application 110 is a web application, a local application installed on client device 108 (e.g., a mobile application, desktop application, etc.), or a cloud-based application in which all or part of its functionality is executed by server(s) 104. Client application 110 presents or displays information to the user, including a digital image editing interface. In some cases, the user interacts with client application 110 to provide user input to perform operations as described above, such as modifying a digital image by removing objects and / or replacing or filling pixels in one or more areas of the digital image.
[0038] In some embodiments, the boot repair system 102 may be implemented entirely or partially using a client application 110 (or client device 108). For example, the boot repair system 102 includes a web-hosted application that allows the client device 108 to interact with servers 104 to send and receive data such as depth-vision guidance and modified digital images (e.g., generated by servers 104). In some cases, the boot repair system 102 generates and provides modified digital images via depth-vision guidance entirely on the client device 108 (e.g., utilizing local processing capabilities on the client device 108), without having to communicate with servers 104.
[0039] likeFigure 1 As shown, the environment includes (multiple) servers 104. Servers 104 generate, track, store, process, receive, and transmit electronic data, such as digital images, visual guidance algorithms, depth visual guidance, patch matching models, and user interaction instructions. For example, servers 104 receive data from client devices 108 in the form of user interaction instructions to select digital image editing operations (e.g., removing objects or replacing pixels in specific areas). Furthermore, servers 104 transmit data to client devices 108 to provide modified digital images, including repaired areas with pixels replaced via patch matching models and depth visual guidance. In practice, servers 104 communicate with client devices 108 to send and / or receive data via network 114. In some embodiments, servers 104 include distributed servers, comprising multiple server devices distributed across network 114 and located in different physical locations. Servers 104 include digital image servers, content servers, application servers, communication servers, web hosting servers, multi-dimensional servers, or machine learning servers.
[0040] like Figure 1 As shown, the (multiple) servers 104 also include a boot repair system 102 as part of the digital content editing system 106. The digital content editing system 106 communicates with client devices 108 to perform various functions associated with client application 110, such as storing and managing a repository of digital images, modifying digital images, and providing modified digital image items for display. For example, the boot repair system 102 communicates with database 112 to access digital images, patch matching models, and visual boot algorithms for modifying digital images. In fact, as... Figure 1 As further shown, the environment includes a database 112. Specifically, database 112 stores information such as digital images, visual guidance algorithms, deep visual guidance, patch matching models, and / or other types of neural networks.
[0041] although Figure 1 A specific arrangement of the environment is shown, but in some embodiments, the environment has a different component arrangement and / or may have a different number of components or sets of components. For example, in some embodiments, the boot repair system 102 is implemented by client device 108 and / or a third-party device (e.g., entirely or partially located on client device 108 and / or the third-party device). Furthermore, in one or more embodiments, client device 108 bypasses network 114 and communicates directly with boot repair system 102. Additionally, in some embodiments, database 112 is located outside of server(s) 104 (e.g., communicating via network 114), or on server(s) 104 and / or client device 108.
[0042] As described above, in one or more embodiments, the guided repair system 102 generates a modified digital image by filling or replacing pixels in one or more regions. Specifically, the guided repair system 102 utilizes a depth vision-guided and patch-matching model to identify replacement pixels for filling one or more regions. Figure 2 The illustration shows how, according to one or more embodiments, a digital image is modified by using depth vision guidance 210 and patch matching model 218 to replace regions 206 and 208 with missing or otherwise undesirable pixels.
[0043] like Figure 2 As shown, the guided repair system 102 generates a modified digital image 204 from an initial digital image 202. Specifically, the guided repair system 102 uses a patch matching model 218 and a depth visual guide 210 to identify replacement pixels and fill regions 206 and 208 with pixels, thereby producing a visually cohesive output in the form of the modified digital image 204 (e.g., where replacement brick pixels are filled where bricks should be, and where replacement shrub pixels are filled where shrubs should be).
[0044] Depth visual guidance (e.g., depth visual guidance 210) includes guidance that instructs or informs a patch matching model (e.g., patch matching model 218) to identify replacement pixels for filling regions (e.g., regions 206 and / or 208) of a digital image (e.g., the initial digital image 202). In some cases, depth visual guidance 210 includes one or more structures within the initial digital image 202, one or more depths within the initial digital image 202, and / or digital and / or visual representations of one or more semantic segments within the initial digital image 202. For example, depth visual guidance 210 includes a structural image guidance 212 (e.g., a structural image where pixels specify objects or structures and edges or boundaries between objects / structures) indicating one or more structures within the initial digital image 202, an image depth guidance 214 (e.g., a depth map where pixels reflect the distance of objects to an observer or camera capturing the digital image), a segmentation image guidance 216 (e.g., a segmentation image where pixels reflect semantic labels of different parts of the digital image) indicating one or more semantic segments within the initial digital image 202, and / or a combination of two or more of the above. In some embodiments, depth vision guidance 210 includes one or more other types of guidance, such as image color guidance indicating one or more colors within a digital image, image normal guidance indicating one or more digital image normals within a digital image, and / or image edge guidance indicating one or more identified edges within a digital image.
[0045] Accordingly, the patch matching model (e.g., patch matching model 218) includes a model or computer algorithm that searches for and / or identifies replacement pixels from a digital image to fill or repair pixel regions (e.g., regions 206 and / or 208). For example, patch matching model 218 modifies an initial digital image 202 to replace regions 206 and 208, which include misaligned or missing pixels, with pixels from the digital image that are visually more clustered relative to other pixels in the initial digital image 202. In practice, using a cost function, patch matching model 218 utilizes one or more pixel sampling techniques (e.g., random or probabilistic) to identify pixels, sampling them and comparing them with those pixels in and around the pixel regions 206 and / or 208 to be replaced. In some embodiments, patch matching model 218 refers to the work of Connelly Barnes, EliShechtman, Adam Finkelstein, and Dan B. Goldman. PatchMatch: A Randomized Correspondence Algorithm for Structural Image Editing The model described in (PatchMatch: A randomized correspondence algorithm for editing structured images), ACM Trans. Graph. 28(3):24(2009), is incorporated herein by reference in its entirety.
[0046] As described above, to generate deep visual guidance 210, the guidance restoration system 102 utilizes a visual guidance algorithm. The visual guidance algorithm includes a computer model or algorithm for generating deep visual guidance (e.g., deep visual guidance 210). For example, the visual guidance algorithm includes a structured image model that generates structured image guidance 212 from a restored digital image (e.g., a restored version of the initial digital image 202). Alternatively, the visual guidance algorithm includes an image deep neural network that generates image depth guidance 214 from the restored digital image. Furthermore, in some cases, the visual guidance algorithm includes segmentation image guidance 216 that generates segmentation image guidance 216 from the restored digital image. As another example, the visual guidance algorithm generates deep visual guidance directly from the initial digital image 202. In these embodiments, the visual guidance algorithm may include one or more of a generator neural network or a teacher-student neural network framework to generate deep visual guidance directly from the initial digital image 202.
[0047] Neural networks encompass machine learning models that can be trained and / or tuned based on inputs to determine a classification or approximation of an unknown function. For example, a neural network includes a model of interconnected artificial neurons (e.g., organized in layers) that communicate and learn to approximate complex functions, generating outputs (e.g., generated digital images) based on multiple inputs provided to the neural network. In some cases, neural networks refer to algorithms (or sets of algorithms) that implement deep learning techniques to model high-level abstractions in data.
[0048] In one or more embodiments, the repaired digital image includes a digital image generated by a deep repair neural network to initially fill pixel regions (e.g., regions 206 and / or 208) for replacement with an initial set of replacement pixels. Specifically, the guided repair system 102 generates a repaired digital image from the initial digital image 202 by utilizing a pre-trained repair neural network to identify the initial set of replacement pixels to (coarsely) repair regions 206 and / or 208 of the initial digital image 202. In some cases, the repaired digital image has a lower resolution than both the initial digital image 202 and the modified digital image 204.
[0049] As described above, through the various actions described for generating the modified digital image 204, in some embodiments, the guided restoration system 102 utilizes multiple neural networks with different architectures to identify and insert replacement pixels. For example, as described above, the guided restoration system utilizes a visual guidance algorithm to generate deep visual guidance 210. As also mentioned, the guided restoration system 102 utilizes a restoration neural network to generate a restored digital image from the initial digital image 202. In some embodiments, the restoration neural network includes a deep neural network that utilizes iterative confidence feedback and guided upsampling to perform digital image restoration.
[0050] like Figure 2 As shown, the guided restoration system 102 utilizes both depth visual guidance 210 (e.g., one or more of structural image guidance 212, image depth guidance 214, or segmentation image guidance 216) and patch matching model 218 to generate a modified digital image 204. In some embodiments, the guided restoration system 102 generates depth visual guidance 210 from the restored digital image. In other embodiments, the guided restoration system 102 utilizes alternative neural networks, such as generator neural networks or teacher-student neural network frameworks, to generate depth visual guidance 210 directly from the initial digital image 202.
[0051] As described above, in some implementations, the guided restoration system 102 generates a depth visual guide (e.g., depth visual guide 210) from the restored digital image. Specifically, the guided restoration system 102 first generates a restored digital image from an input digital image (e.g., an initial digital image 202) before subsequently generating the depth visual guide (e.g., depth visual guide 210). Figure 3 The illustration shows how, according to one or more embodiments, a modified digital image 314 is generated from an input digital image 302 (which has holes or regions with one or more missing pixels) by generating a depth visual guide 310 from a repaired digital image 306.
[0052] like Figure 3As shown, the guided restoration system 102 generates a restored digital image 306 from an input digital image 302. More specifically, the guided restoration system 102 utilizes a deep restoration neural network 304 to generate the restored digital image 306 from the input digital image 302. For example, the deep restoration neural network 304 generates a low-resolution image of the input digital image 302 and processes or analyzes the low-resolution image to identify an initial set of replacement pixels to fill or restore missing or defective regions 303a and 303b of the input digital image 302. In some embodiments, the deep restoration neural network 304 is pre-trained to identify the initial set of replacement pixels using parameters learned from a dataset of sample digital images. For example, the deep restoration neural network 304 identifies (or receives its indication) the pixel regions 303a and 303b to be replaced and, based on the learned parameters, replaces pixels in regions 303a and 303b with pixels copied from other parts of the input digital image 302. In one or more embodiments, the guided repair system 102 utilizes a deep repair neural network 304, such as those developed by Yu Zeng, Zhe Lin, Jimei Yang, Jianming Zhang, EliShechtman, and Huchuan Lu. High-resolution Inpainting with Iterative Confidence Feedback and Guided Upsampling The neural network described in arXiv:2005.11742 (2020), which is incorporated herein by reference in its entirety, is an example of a neural network with iterative confidence feedback and guided upsampling.
[0053] In fact, such as Figure 3 As shown, the repaired digital image 306 includes a (lower resolution) version of the input digital image 302, wherein regions 303a and 303b are at least initially filled with an initial, coarsely matched set of replacement pixels. Upon closer inspection, the pixels used to replace regions 303a and 303b when generating the repaired digital image 306 still have various defects, such as lower resolution, unwanted regions, and mismatched pixels that make the brick wall appear uneven or curved. To ultimately generate a more accurate modified digital image 314, the guided repair system 102 also generates a depth visual guide 310 from the repaired digital image 306.
[0054] More specifically, the guided restoration system 102 utilizes a visual guidance algorithm 308 to generate depth visual guidance 310 from the restored digital image 306. Specifically, the guided restoration system 102 generates depth visual guidance 310 in the form of structural image guidance 311a, image depth guidance 311b, segmentation image guidance 311c, or a combination of two or more of the above. In some cases, the visual guidance algorithm 308 is pre-trained with parameters learned from a database of sample digital images to generate depth visual guidance in one or another form based on these parameters.
[0055] As described above, in some embodiments, the guided restoration system 102 generates structural image guidance 311a as a depth visual guidance 310. More specifically, the guided restoration system 102 utilizes a structural image model as a visual guidance algorithm 308 to process the restored digital image 306 and generate structural image guidance 311a. For example, the guided restoration system 102 uses the structural image model to identify one or more distinct structures within the restored digital image 306. In some cases, the structural image model identifies structures based on detecting edges within the restored digital image 306.
[0056] Specifically, the guided inpainting system 102 extracts structure from the inpainted digital image 306 using a structural image model based on local change metrics (e.g., inherent change and relative total change). For example, the structural image model identifies meaningful content and texture edges without assuming regularity or symmetry. In some cases, as part of identifying edges or boundaries between different structural components, the structural image model further identifies, preserves, or removes texture (or at least some texture components) from the digital image. Therefore, structural images tend to include large, cohesive regions that identify different structures (e.g., without small regional variations). Because structural images tend to provide large, homogeneous regions indicating common structures, these images can provide an accurate resource for patch matching models in determining which regions to extract from when identifying replacement pixels (since small regional variations are easily diluted or ignored when applying patch matching algorithms).
[0057] Furthermore, the structural image model decomposes the optimization problem to extract the principal structure from the measure of variation. Specifically, the structural image model extracts the principal structure based on parameters learned by comparing predicted structures with ground-based real structure information from a dataset of sample digital images. In some embodiments, the guided restoration system 102 utilizes a structural image model in the form of a computer vision algorithm such as RTV smoothing. For example, the guided restoration system 102 utilizes the work of Li Zu, Qiong Yan, Yang Xia, and Jiaya Jia. Structure Extraction from Texture via Relative Total Variation (Texture structure extraction via relative total variation), the structural image model described in ACM Transactions on Graphics 31(6):1-10 (2012), the entire text of which is incorporated herein by reference. In some embodiments, the guided repair system 102 utilizes different structural image models to identify various structures within a digital image.
[0058] In one or more embodiments, the guided restoration system 102 generates image depth guidance 311b as depth visual guidance 310. Specifically, the guided restoration system 102 utilizes an image deep neural network as a visual guidance algorithm 308 to process the restored digital image 306 and generate image depth guidance 311b. For example, the guided restoration system 102 utilizes an image deep neural network to identify one or more different depths within the restored digital image 306.
[0059] Specifically, the guided restoration system 102 utilizes an image deep neural network to determine monocular depth predictions for different objects or structures within the restored digital image 306. Specifically, the image deep neural network analyzes the restored digital image 306 to generate a single-channel depth map. For example, the image deep neural network generates a depth map based on parameters learned from a dataset of sample digital images, used to extract pseudo-depth data for comparison with true ground depth information. In some cases, the guided restoration system 102 utilizes a pre-trained image deep neural network such as DepthNet. For example, the guided restoration system 102 utilizes the work of Ke Xian, Jianming Zhang, Oliver Wang, Long Mai, Zhe Lin, and Zhiguo Cao. Structure- guided Ranking Loss for Single Image Depth Prediction The image deep neural network described in (Structure-Guided Ranking Loss for Depth Prediction of a Single Image), the full text of which is incorporated herein by reference. As another example, the guided inpainting system 102 utilizes different image deep neural networks, such as the ResNet-50 architecture which utilizes convolutional neural network methods (e.g., including fully convolutional networks and / or residual neural networks).
[0060] In some embodiments, the guided restoration system 102 generates segmented image guidance 311c as a depth visual guidance 310. More specifically, the guided restoration system 102 utilizes a segmented image neural network as a visual guidance algorithm 308 to process the restored digital image 306 and generate segmented image guidance 311c. For example, the guided restoration system 102 utilizes a segmented image neural network to identify one or more distinct semantic segments within the restored digital image 306.
[0061] More specifically, the guided inpainting system utilizes a segmentation image neural network to identify parts of the digit image 306 corresponding to different semantic labels. Specifically, the segmentation image neural network assigns labels to parts, objects, or structures in the digit image 306 based on parameters learned from comparing predicted semantic labels with ground truth semantic labels on a dataset of sample digit images. In some cases, the guided inpainting system 102 utilizes a pre-trained segmentation image neural network such as SegmentNet. For example, the guided inpainting system utilizes the work of Ke Sun, Yang Zhao, Borui Jian, TianhengCheng, Bin Xiao, Dong Liu, Yadong Mu, Xinggang Wang, Wenyu Liu, and Jingdong Wang. High-resolution Representations for Labeling Pixels and Regions (High-resolution representations for labeling pixels and regions), a segmentation image neural network described in arXiv:1904.04515 (2019), the full text of which is incorporated herein by reference. As another example, guided inpainting system 102 utilizes different segmentation image neural networks to label parts or objects of a digital image.
[0062] like Figure 3 As further shown, the guided repair system 102 utilizes a guided patch matching model 312 to generate a modified digital image 314. Specifically, the guided repair system 102 utilizes the guided patch matching model 312, which is notified by the depth vision guidance 310. To generate the modified digital image, the guided repair system 102 uses the guided patch matching model 312 to analyze the input digital image 302 to identify replacement pixels for regions 303a and 303b indicated by the depth vision guidance 310. For example, the guided patch matching model 312 identifies regions 303a and 303b and compares pixels from different parts of the input digital image 202 to identify good candidates for filling regions 303a and 303b.
[0063] Specifically, depth vision guidance 310 informs guidance patch matching model 312 where to identify replacement pixels from input digital image 302. For example, depth vision guidance 310 indicates portions of input digital image 302 with structural, depth, and / or semantic segmentation corresponding to regions 303a and / or 303b. Therefore, guidance patch matching model 312 analyzes pixels in those portions to identify and select replacement pixels for filling regions 303a and 303b. In effect, by filling regions 303a and 303b, guidance repair system 102 thereby generates a semantically consistent modified digital image 314 with visually appropriate replacement pixels for regions 303a and 303b (e.g., matching the resolution of input digital image 302 to a high resolution). In some cases, guidance repair system 102 utilizes one or more specific techniques to determine structural, depth, and / or semantic segmentation from input digital image 302. For example, the boot repair system 102 was implemented by Aaron Hertzmann, Charles E. Jacobs, Nuria Oliver, Brian Curless, and David H. Salesin. Image Analogies (Image analogy), Proceedings of the 28 th The techniques described in the Annual Conference on Computer Graphics and Interactive Techniques, 327-40 (2001), are incorporated herein by reference in their entirety.
[0064] In some embodiments, the guidance patch matching model 312 is orthogonal to the visual guidance algorithm 308 and the deep repair neural network 304. Therefore, the guidance patch matching model can be easily adapted to different types of deep models or other systems that can be used to generate deep visual guidance. Thus, the hybrid approach of the guidance repair system 102, combining patch-based and learning-based methods, enables flexible, plug-and-play applications for future developments in generating deep visual guidance.
[0065] although Figure 3 Three different types of depth visual guidance 310 are illustrated: i) structural image guidance, ii) image depth guidance, and iii) segmentation image guidance, but other depth visual guidance is also possible. For example, in some embodiments, the guidance restoration system 102 generates the depth visual guidance 310 from the digital image normal of the input digital image 302. As another example, the guidance restoration system 102 utilizes an edge filter to detect one or more edges within the input digital image 302 to generate the depth visual guidance 310.
[0066] As another example, guided restoration system 102 generates depth visual guidance from the raw colors of the input digital image 302. However, while predictive image color guidance produces the best results in generating modified digital images, researchers have determined that structure, depth, and segmentation are surprisingly better. In some embodiments, guided restoration system 120 combines color with one or more other types of depth visual guidance to enhance accuracy and performance.
[0067] As described above, in one or more embodiments, the guided restoration system 102 directly generates depth visual guidance based on the input digital image (e.g., without generating a restoration digital image). Specifically, in some cases, the guided restoration system 102 extracts an occlusion structure image and a binary mask from the input digital image, and uses a generator neural network to generate depth visual guidance based on the input digital image, the occlusion structure image, and the binary mask. Figure 4 The diagram illustrates how, according to one or more embodiments, a generator neural network 408 directly generates a depth visual guide 410 from an input digital image 402.
[0068] like Figure 4 As shown, the guided restoration system 102 generates an intermediate digital image 404 (i.e., a structural image) based on the input digital image 402. Specifically, the guided restoration system 102 generates the intermediate digital image 404 as an occluding structural image that indicates the structure of the input digital image 402 in one or more portions outside the pixel regions 403a and 403b to be replaced. To generate the intermediate digital image 404, the guided restoration system 102 extracts different structures from the input digital image 402 using a structural image model. Additionally, the guided restoration system 102 generates a binary mask 406 from the input digital image 402. The binary mask 406 indicates the pixel regions 403a and 403b to be replaced.
[0069] like Figure 4 As further shown, the guided restoration system 102 utilizes a generator neural network 408 to process the input digital image 402, the intermediate digital image 404, and the binary mask 406. Furthermore, the generator neural network 408, comprising an encoder neural network (denoted by "E") and a decoder neural network (denoted by "D"), generates a depth visual guide 410 in the form of a structured image guide. In some cases, the generator neural network 408 outputs a partial structured digital image that indicates only the structure of regions 403a and 403b (without indicating the structure of other parts of the input digital image 402). Additionally, the guided restoration system 102 combines the partial structure output from the generator neural network 408 with the intermediate digital image 404 to generate a depth visual guide 410 that indicates the structure throughout the entire input digital image 402.
[0070] althoughFigure 4 The illustration shows the use of a generator neural network 408 to generate structural image guidance, but alternative embodiments are also possible. For example, in some embodiments, the guidance restoration system 102 utilizes the generator neural network 408 to generate image depth guidance and / or segmentation image guidance. To generate image depth guidance, the guidance restoration system 102 generates an intermediate digital image (e.g., using an image depth neural network) that indicates the depth in portions of the input digital image 402 outside regions 403a and 403b. Additionally, the guidance restoration system 102 inputs the intermediate digital image into the generator neural network 408 to predict the depths of regions 403a and 403b. Thus, the guidance restoration system 102 combines the predicted depths with the intermediate digital image to generate depth visual guidance in the form of image depth guidance.
[0071] To generate segmented image guidance, guidance restoration system 102 generates an intermediate digital image (e.g., using a segmented image neural network) that indicates semantic segmentation of portions of input digital image 402 outside regions 403a and 403b. Furthermore, guidance restoration system 102 inputs the intermediate digital image into generator neural network 408 to predict one or more segments of region 403a and one or more segments of region 403b. Therefore, guidance restoration system 102 combines the predicted segments with the intermediate digital image to generate depth visual guidance in the form of segmented image guidance.
[0072] In one or more embodiments, the guided restoration system 102 trains a generator neural network 408 to predict structures. For example, the guided restoration system 102 learns the parameters of the component encoder neural network and the component decoder neural network that make up the generator neural network 408. Specifically, the guided restoration system 102 inputs a sample digital image into the encoder neural network, which encodes or extracts the latent code or feature representation of the sample digital image. Furthermore, the guided restoration system 102 passes the encoded feature representation to the decoder neural network, which decodes the feature representation to generate or reproduce the sample digital image (or its best approximation).
[0073] Furthermore, the guided restoration system 102 utilizes a discriminator neural network to test the generated or reproduced digital images to predict whether the images are real (e.g., from a digital image database) or artificial (e.g., generated by the generator neural network 408). In some embodiments, the guided restoration system 102 utilizes an adversarial loss function to determine an error or loss metric associated with the generator neural network 408 (e.g., between the generator neural network 408 and the discriminator neural network). In effect, the guided restoration system 102 determines an adversarial loss that, in certain cases, indicates how accurate (or inaccurate) the reconstructed digital images are and / or how effective the generator neural network 408 is at deceiving the discriminator neural network into identifying artificial digital images as real digital images. The guided restoration system 102 improves the generator neural network 408 by modifying parameters to reduce the adversarial loss over multiple iterations of generating reconstructed digital images and determining whether they are real or fake.
[0074] Additionally, in some embodiments, the guided repair system 102 also utilizes a supervised learning method with L1 loss to modify parameters and improve the generator neural network 408. For example, the guided repair system 102 inputs a sample digit image into the generator neural network 408, which then generates a predicted depth visual guide (e.g., depth visual guide 410). Furthermore, the guided repair system 102 uses an L1 loss function to compare the predicted depth visual guide with ground truth depth visual guides (e.g., stored as ground truth depth visual guides corresponding to the sample digit image). Using the L1 loss function, the guided repair system 102 determines a loss metric associated with the generator neural network 408. Furthermore, the guided repair system 102 modifies the parameters of the generator neural network 408 to reduce or minimize the L1 loss. The guided repair system 102 repeats the following process: inputting a sample digit image, generating a predicted depth visual guide, comparing the predicted depth visual guide with the ground truth depth visual guide, and modifying the parameters for multiple iterations or epochs (e.g., until the L1 loss meets a threshold loss).
[0075] As described above, in some of the described embodiments, the guided restoration system 102 utilizes a teacher-student neural network framework to generate deep visual guidance. Specifically, the guided restoration system 102 utilizes a student neural network with parameters learned from the teacher neural network to generate deep visual guidance directly from the input digital image. Figure 5 The diagram illustrates how, according to one or more embodiments, a student neural network 512 with parameters learned from a teacher neural network 510 is used to generate a deep visual guide 508.
[0076] In some embodiments, the student neural network includes a neural network that learns or passes parameters from the teacher neural network. On the other hand, the teacher neural network includes a neural network that learns parameters through multiple learning iterations (e.g., generating predictions based on a sample dataset and modifying parameters to reduce the loss metric associated with the prediction).
[0077] like Figure 5 As shown, the guided restoration system 102 utilizes a student neural network 512 to generate a depth visual guide 508 in the form of image depth guidance. Specifically, the guided restoration system 102 utilizes the student neural network 512 to process an incomplete digital image 506 that includes hole regions 507 of pixels to be replaced, and the student neural network 512 generates the depth visual guide 508 based on parameters learned from the teacher neural network 510.
[0078] In practice, the student neural network 512 includes parameters learned from the teacher neural network 510. As shown, the guided restoration system 102 learns parameters for the student neural network 512 by transmitting or adapting parameters from the teacher neural network 510. For example, the guided restoration system 102 uses the teacher neural network 510, pre-trained on a dataset of sample digit images, to label the depth of the digit images. As shown, the guided restoration system 102 inputs a complete digit image 502 (e.g., a digit image without any pixel regions to be replaced) into the teacher neural network 510. The teacher neural network 510 then determines teacher labels for different depths and uses locked or frozen parameters to generate image depth guidance 504.
[0079] like Figure 5 As further shown, the guided restoration system 102 generates an incomplete digital image 506 from the complete digital image 502. The incomplete digital image 506 includes hole regions 507 of pixels to be replaced. The guided restoration system 102 also initializes the weights or parameters of the student neural network 512 with the parameters of the learned teacher neural network 510.
[0080] Furthermore, the guided restoration system 102 inputs the incomplete digit image 506 into a student neural network 512, which then generates depth visual guidance 508 based on the learned parameters. For example, the guided restoration system 102 utilizes the student neural network 512 to generate predictive labels for different depths of the incomplete digit image 506. The guided restoration system 102 uses parameters learned or transmitted from the teacher neural network 510 to generate depth visual guidance 508 from the incomplete digit image 506 with hole regions 507.
[0081] In one or more embodiments, the guidance restoration system 102 modifies the parameters of the student neural network 512 to improve the accuracy of generating outputs similar to those of the teacher neural network 510. More specifically, the guidance restoration system 102 compares predicted labels from the student neural network 512 with teacher labels from the teacher neural network 510. In some cases, the guidance restoration system 102 utilizes a gradient loss function to determine a loss metric associated with the student neural network 512 (or between the student neural network 512 and the teacher neural network 510). Additionally, the guidance restoration system 102 modifies the parameters of the student neural network 512 to reduce the loss metric. In some embodiments, when generating image depth guidance, the guidance restoration system 102 utilizes a Structured Guidance Ranking Loss (“SRGL”) neural network for the teacher neural network 510 and / or the student neural network 512. By iteratively modifying the parameters of the student neural network 512 based on the teacher labels generated by the teacher neural network 510, the guidance restoration system 102 can train the student neural network 512 to generate accurate depth visual guidance from digital images including holes or other replacement regions.
[0082] although Figure 5 The illustration shows the generation of depth visual guidance 508 in the form of image depth guidance, but additional embodiments are also possible. For example, in some embodiments, the guidance restoration system 102 utilizes a teacher-student neural network framework to generate depth visual guidance in the form of segmented image guidance. As another example, the guidance restoration system 102 utilizes a teacher-student neural network framework to generate depth visual guidance in the form of structured image guidance. To generate segmented image guidance, the guidance restoration system 102 utilizes a student neural network 512 to process the incomplete digital image 506 and generates segmented image guidance based on parameters learned from the teacher neural network 510.
[0083] To learn the parameters of the student neural network 512, the guided restoration system 102 processes the complete digital image 502 using the teacher neural network 510. The teacher neural network then generates teacher labels for the segmented image guidance, which instructs on the semantic segmentation of the complete digital image 502. As described above, the guided restoration system 102 also transmits parameters from the teacher neural network 510 to the student neural network 512. Furthermore, the guided restoration system 102 utilizes the student neural network 512 to generate predicted labels for the semantic segmentation of the incomplete digital image 506.
[0084] As described above, the guided restoration system 102 encourages the student neural network 512 (from the incomplete digit image 506) to generate segmentation image guidance similar to that generated by the teacher neural network 510 (from the complete digit image 502). Specifically, the guided restoration system 102 utilizes a negative log-likelihood loss function to determine the loss metric associated with the student neural network 512. Furthermore, the guided restoration system 102 modifies the parameters of the student neural network 512 to reduce the loss metric, thereby generating an output more similar to the output of the teacher neural network 510. In some embodiments, when generating segmentation image guidance, the guided restoration system 102 utilizes an HRNet neural network pre-trained on the ADE20K dataset for both the teacher neural network 510 and / or the student neural network 512.
[0085] As described above, in some embodiments, the guided repair system 102 combines a patch matching model with one or more depth visual guides to identify and select replacement pixels. Specifically, the guided repair system 102 utilizes a cost function as part of the patch matching model to identify and select replacement pixels to fill regions of the digital image. Figure 6 The illustration shows a modification and utilization of the patch matching cost function to merge one or more depth vision guides, according to one or more embodiments.
[0086] like Figure 6 As shown, the guided restoration system 102 determines the weights of the depth visual guidance 602 (or its corresponding visual guidance algorithm). Specifically, the guided restoration system 102 determines the structural weights for the structural image guidance, the depth weights for the image depth guidance, and the segmentation weights for the segmentation image guidance. Based on these weights, the guided restoration system 102 combines the depth visual guidance (or its corresponding visual guidance algorithm) to identify and select replacement pixels. In some cases, the guided restoration system 102 utilizes default weights, where the structural weights, depth weights, and segmentation weights are the same as or different from each other.
[0087] like Figure 6 As further shown, the guided restoration system 102 utilizes an initial cost function 604. Specifically, the guided restoration system 102 utilizes the initial cost function 604 as part of a patch matching model (e.g., patch matching model 218) to identify and select replacement pixels from the digital image. For example, the guided restoration system 102 utilizes the initial cost function 604 to determine the relationship or distance between pixels in the missing region and pixels in a subset of potential replacement pixels. In some cases, the guided restoration system 102 implements the initial cost function 604 to determine the target portion (e.g., defined by two-dimensional point "dst"). xyThe sum of squared differences (“SSD”) between the target pixel or portion (e.g., a potential replacement pixel or portion currently being analyzed by the patch matching model) in the “specified” and the pixel or patch to be replaced by a two-dimensional point (e.g., “srcProposedPt”). For example, the bootstrap repair system 102 utilizes an initial cost function 604 given by the following formula:
[0088]
[0089] SSD(.) specifies two items. and The sum of the squared differences between them, where These are potential (multiple) replacement pixels. The missing or unwanted (multiple) pixels to be replaced.
[0090] like Figure 6 As shown, the guided inpainting system 102 also generates a modified cost function 606. Specifically, the guided inpainting system 102 modifies the initial loss function 604 using the weights of the depth visual guidance 602. For example, the guided inpainting system 102 generates a weighted combination of the initial loss function and weighted versions of the structure image guidance, image depth guidance, and segmentation image guidance. In fact, the guided inpainting system 102 combines two or more depth visual guidances according to the weights assigned to each depth visual guidance (or the corresponding visual guidance algorithm). For example, the guided inpainting system 102 utilizes the modified cost function given by:
[0091]
[0092] in, This represents the initial cost function 604 defined above (e.g., the color item). This represents the color weights of the initial cost function 604. Indicates depth weights, Represents structural weights, Indicates the splitting weight. The cost function represents the sum of squared differences between the pixels used to determine the pixels to be replaced and the pixels indicated by image depth as potential replacement pixels. This represents the cost function that sums the squared differences between the pixels used to determine the pixels to be replaced and the pixels indicated as potential replacement pixels by the structure image. This represents the cost function that sums the squared differences between the pixels used to determine which ones to replace and the pixels indicated as potential replacement pixels by the segmented image.
[0093] By utilizing a modified cost function 606, the guided inpainting system 102 uses more than one depth visual guide (e.g., combined depth visual guides) to guide the patch matching model to identify replacement pixels. Specifically, the guided inpainting system 102 generates a weighted combination of two or more of structural image guides, image depth guides, and / or segmentation image guides to identify replacement pixels (e.g., together with the patch matching model). In some cases, the guided inpainting system 102 automatically determines the appropriate weights for the depth visual guides (e.g., by weighting them uniformly or in a specific ascending or descending order). In these or other cases, the guided inpainting system 102 uses one or more weighted depth visual guides to modify the cost function of the patch matching model (but not necessarily all three). For example, the guided inpainting system 102 uses a modified cost function such as:
[0094]
[0095] As described above, in some embodiments, the guided repair system 102 improves accuracy compared to conventional digital image editing systems. Specifically, the guided repair system 102 can generate modified digital images that are more semantically consistent and include replacement pixels that more accurately fill in missing or unwanted pixel areas. Figures 7-9 An example comparison is shown between a modified digital image generated by the boot repair system 102 and a modified digital image generated by a conventional patch matching system, according to one or more embodiments.
[0096] For example, Figure 7 A modified digital image 708 generated by an example embodiment of the guided repair system 102 using a structural image guide 704 is shown. In fact, the example embodiment of the guided repair system 102 generates a structural image guide from the input digital image 702 and also generates the modified digital image 708 by filling in missing regions of the input digital image 702 using a structural guide patch matching model. Comparing the modified digital image 708 to a modified digital image 706 generated by a conventional patch matching system, the example embodiment of the guided repair system 102 performs better by using replacement pixels that do not add unwanted or semantically incoherent artifacts (which are added to the modified digital image 706). In fact, the modified digital image 706 includes phantom table legs that are not present in the modified digital image 708.
[0097] Figure 8A comparison is shown between the output of a guided repair system 102 using image depth guidance 804 and the output of a conventional patch matching system, according to one or more embodiments. As shown, an example embodiment of the guided repair system 102 generates image depth guidance 804 from an input digital image 802. Furthermore, the guided repair system 102 utilizes a depth-guided patch matching model to generate a modified digital image 808. The modified digital image 808 is compared to a modified digital image 806 generated by a conventional patch matching system, which includes pixels that are undesirably smaller and appear more realistic compared to the modified digital image 806.
[0098] Figure 9 A comparison is shown between the output of a guided image repair system 102 using segmented image guidance 904 and the output of a conventional patch matching system, according to one or more embodiments. As shown, an example embodiment of the guided image repair system 102 generates segmented image guidance 904 from an input digital image 902. Furthermore, the guided image repair system 102 utilizes a segmentation-guided patch matching model to generate a modified digital image 908. The modified digital image 908 is compared to a modified digital image 906 generated by a conventional patch matching system, which includes pixels that better match the sky of the input digital image 902 and does not include pixels with added artifacts (unlike the modified digital image 906, which includes pixels that appear to be part of buildings floating in the sky).
[0099] Now for reference Figure 10 Additional details regarding the components and capabilities of the boot repair system 102 will be provided. Specifically, Figure 10 An example schematic diagram of a boot repair system 102 on an example computing device 1000 (e.g., one or more of client devices 108 and / or servers 104). Figure 10 As shown, the boot repair system 102 includes a deep repair manager 1002, a deep vision boot manager 1004, a replacement pixel manager 1006, a digital image manager 1008, and a storage manager 1010.
[0100] As described above, the guided repair system 102 includes a deep repair manager 1002. Specifically, the deep repair manager 1002 manages, maintains, stores, accesses, applies, utilizes, implements, or identifies a deep repair neural network. In practice, the deep repair manager 1002 generates or creates a preliminary repaired digital image from the input digital image by identifying an initial set of replacement pixels used to fill one or more pixel regions within the input digital image to be replaced.
[0101] Furthermore, the guided restoration system 102 includes a deep visual guidance manager 1004. Specifically, the deep visual guidance manager 1004 manages, maintains, stores, accesses, determines, generates, or identifies deep visual guidance. For example, the deep visual guidance manager 1004 generates deep visual guidance based on a restored digital image using one or more visual guidance algorithms. In some cases, the deep visual guidance manager 1004 generates deep visual guidance directly from an input digital image (e.g., without using a restored digital image) via a generator neural network and / or a teacher-student neural network framework. In some embodiments, the deep visual guidance manager 1004 generates combined deep visual guidance and / or learns parameters for one or more neural networks used to generate deep visual guidance, as described above.
[0102] As shown in the figure, the guided repair system 102 also includes a replacement pixel manager 1006. Specifically, the replacement pixel manager 1006 manages, maintains, stores, identifies, accesses, selects, or identifies replacement pixels from the input digital image (or different digital images). For example, the replacement pixel manager 1006 identifies replacement pixels using a patch matching model along with depth vision guidance, or by notification from depth vision guidance.
[0103] Furthermore, the guided repair system 102 includes a digital image manager 1008. Specifically, the digital image manager 1008 manages, stores, accesses, generates, creates, modifies, fills, repairs, or identifies digital images. For example, the digital image manager 1008 utilizes a patch matching model and depth vision guidance to generate a modified digital image by filling regions of an input digital image with replacement pixels.
[0104] The guided repair system 102 also includes a storage manager 1010. The storage manager 1010 includes or operates in conjunction with one or more storage devices, such as a database 1012 (e.g., database 112) storing various types of data, such as a repository of digital images and various neural networks. The storage manager 1010 (e.g., via non-transient computer memory / one or more storage devices) stores and maintains data associated with generating repaired digital images, generating deep visual guidance, learning parameters of neural networks, and generating modified digital images (e.g., within database 1012) via a guided patch matching model. For example, the storage manager 1010 stores a repair neural network, including a visual guidance algorithm comprising at least one of a structured image model, an image deep neural network, or a segmentation image neural network, a patch matching model, and a digital image including the pixel regions to be replaced.
[0105] In one or more embodiments, each component of the boot repair system 102 communicates with each other using any suitable communication technology. Additionally, components of the boot repair system 102 communicate with one or more other devices, including the aforementioned client devices. It will be appreciated that, although in Figure 10 The components of the boot repair system 102 are shown to be separate, but any sub-component can be combined into fewer components, such as a single component, or divided into more components to serve a specific implementation. Furthermore, although described in conjunction with the boot repair system 102... Figure 10 The components described herein for performing operations in conjunction with the boot repair system 102 may be implemented on other devices within this environment.
[0106] Components of the boot repair system 102 may include software, hardware, or both. For example, components of the boot repair system 102 may include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices (e.g., computing device 1000). When executed by one or more processors, the computer-executable instructions of the boot repair system 102 may cause the computing device 1000 to perform the methods described herein. Alternatively, components of the boot repair system 102 may include hardware, such as a dedicated processing device that performs a particular function or group of functions. Additionally or alternatively, components of the boot repair system 102 may include a combination of computer-executable instructions and hardware.
[0107] Furthermore, components of the boot repair system 102 that perform the functions described herein may be implemented, for example, as part of a standalone application, as a module of an application, as a plug-in to an application including a content management application, as one or more library functions that can be called by other applications, and / or as a cloud computing model. Therefore, components of the boot repair system 102 may be implemented as part of a standalone application on a personal computing device or mobile device. Alternatively or additionally, components of the boot repair system 102 may be implemented in any application that allows the creation and delivery of marketing content to users, including but not limited to applications in Adobe® Experience Manager and Creative Cloud®, such as Adobe® Stock, Photoshop®, Illustrator®, and Indesign®. "ADOBE", "ADOBE Experience Manager", "CREATIVE Cloud", "ADOBE Stock", "Photoshop", "Illustrator", and "Indesign" are registered trademarks or trademarks of Adobe Systems Incorporated in the U.S. and / or other countries.
[0108] Figures 1-10 The corresponding text and examples provide several different systems, methods, and non-transient computer-readable media for generating modified digital images by replacing pixels using a guided patch matching model identifier. In addition to the above, embodiments may also be described according to flowcharts including actions for achieving specific results. For example, Figure 11 A flowchart illustrating an example sequence or series of actions according to one or more embodiments is shown.
[0109] Although Figure 11 Actions according to one embodiment are shown, but alternative embodiments may omit, add, reorder, and / or modify them. Figure 11 Any action shown. Figure 11 The action can be performed as part of a method. Alternatively, a non-transient computer-readable medium can include instructions that, when executed by one or more processors, cause a computing device to perform... Figure 11 The system can perform the following actions. In other embodiments, the system can execute... Figure 11 The actions described herein can be repeated or performed in parallel with each other, or performed in parallel with different instances of the same or other similar actions.
[0110] Figure 11 An example series of actions 1100 is shown to generate a modified digital image by identifying replacement pixels using a guided patch matching model. Specifically, the series of actions 1100 includes an action 1102 to generate the repaired digital image. For example, action 1102 involves generating a repaired digital image from the digital image using a repair neural network, the repaired digital image including an initial set of replacement pixels for the region.
[0111] As shown in the figure, the action series 1100 also includes an action 1104 for generating depth visual guidance. Specifically, action 1104 involves generating depth visual guidance from the repaired digital image using a visual guidance algorithm. For example, action 1104 involves generating a depth visual guidance image that indicates one or more of structural, depth, or semantic segmentation within a region of the digital image. In some embodiments, action 1104 involves generating depth visual guidance from the repaired digital image using a visual guidance algorithm, the depth visual guidance including at least one of structural image guidance, image depth guidance, or segmentation image guidance. In some cases, identifying replacement pixels involves using a patch matching model to identify pixels within the digital image that correspond to structural, depth, or semantic segmentation within the digital image region indicated by the depth visual guidance.
[0112] In some embodiments, action 1104 involves generating a structural image guide from a repaired digital image using a visual guidance algorithm including a structural image model, for replacing pixels from one or more structural identifiers identified within the repaired digital image. In some cases, action 1104 involves generating a structural image guide using a structural image model as a depth visual guide to determine edges between different structural components of the repaired digital image. In these or other embodiments, action 1104 involves generating an image depth guide from the repaired digital image using a visual guidance algorithm including an image deep neural network, for replacing pixels from depth map identifiers of the repaired digital image. In the same or other embodiments, action 1104 involves generating an image depth guide from the repaired digital image using a visual guidance algorithm including an image deep neural network, for replacing pixels from depth map identifiers of the repaired digital image. In some embodiments, action 1104 involves generating a segmented image guide from the repaired digital image using a visual guidance algorithm including a segmentation image neural network, for replacing pixels from semantic segmentation identifiers of the repaired digital image.
[0113] Furthermore, the action series 1100 includes an action 1106 for identifying replacement pixels. Specifically, action 1106 involves identifying replacement pixels for regions of a digital image based on a patch matching model and depth vision guidance. For example, action 1106 involves using a patch matching model to identify one or more of the following: pixels of a digital image belonging to a structure of a pixel region guided by a structure image; pixels of a digital image having a depth corresponding to the depth of a pixel region guided by image depth; or pixels of a segmented digital image belonging to a pixel region guided by a segmented image.
[0114] like Figure 11 As shown, action series 1100 includes action 1108 of generating a modified digital image. Specifically, action 1108 involves generating a modified digital image by replacing regions of a digital image with replacement pixels. For example, action 1108 involves replacing regions of a digital image with replacement pixels, wherein the modified digital image has a higher resolution than the repaired digital image.
[0115] In some embodiments, action series 1100 includes actions that utilize a convolutional neural network to select a visual guidance algorithm from a set of visual guidance algorithms from a digital image for generating depth visual guidance. In the same or other embodiments, action series 1100 includes actions that generate a combination of depth visual guidance, which includes two or more of structural image guidance, image depth guidance, or segmentation image guidance. In these embodiments, action 1108 involves using a patch matching model and the combined depth visual guidance to identify replacement pixels.
[0116] In one or more embodiments, the series of actions 1100 includes actions that utilize a convolutional neural network to select a visual guidance algorithm for generating the deep visual guidance from a set of deep visual guidances (including structural image guidance, image depth guidance, and segmentation guidance) from a digital image. In some cases, the series of actions 1100 includes actions that utilize a convolutional neural network to determine a correspondence metric between the digital image and each of the structural image model, the image depth neural network, and the segmentation image neural network.
[0117] Additionally, the action series 1100 includes actions that assign weights to each of the structured image model, the image deep neural network, and the segmentation image neural network according to a correspondence metric. Furthermore, the action series 1100 includes actions that generate depth vision guidance by combining two or more of the structured image guidance, image deep guidance, or segmentation image guidance according to weights. In some embodiments, the action series 1100 includes actions that learn parameters of a convolutional neural network from a digital image database classified as corresponding to one or more of the structured image model, image deep neural network, or segmentation image neural network.
[0118] In some embodiments, the action series 1100 includes the action of receiving a digital image including the pixel region to be replaced. In these or other embodiments, the action series 1100 includes the action of generating an intermediate digital image indicating a structure of a portion of the digital image outside the region and the action of generating a binary mask indicating the region of the digital image. Furthermore, the action series 1100 includes the action of using a generator neural network to predict the structure within the region of the digital image based on the digital image, the intermediate digital image, and the binary mask. In some cases, the action series 1100 includes the action of generating a binary mask indicating the region of the digital image, and the action of generating depth vision guidance by using a generator neural network to predict the structure within the region of the digital image based on the binary mask.
[0119] In one or more embodiments, the series of actions 1100 includes actions such as using a student neural network to predict one or more depths or segments within a region of a digital image based on parameters learned from a teacher neural network. Furthermore, the series of actions 1100 includes actions such as using a teacher neural network to determine teacher labels for one or more depths or segments of a complete digital image, generating an incomplete digital image including hole regions from the complete digital image, using a student neural network to generate predicted labels for one or more depths or segments of the incomplete digital image, and modifying the parameters of the student neural network by comparing the predicted labels with the teacher labels.
[0120] In some embodiments, action series 1100 includes actions that guide the identification of replacement pixels from digital images and depth vision using a cost function of a patch matching model. For example, identifying replacement pixels may involve using a cost function to identify replacement pixels based on a weighted combination of structural image guidance, image depth guidance, and segmentation image guidance.
[0121] Action series 1100 may include actions that replace pixels from digital image identifiers according to a cost function of a patch matching model utilizing depth vision guidance. In some cases, action series 1100 includes actions that combine depth vision guidance and additional depth vision guidance by assigning weights to two or more of structural image guidance, image depth guidance, or segmentation guidance using the cost function. Replacing pixels from digital image identifiers may include identifying replacement pixels using structural image guidance, image depth guidance, and segmentation image guidance, according to a modified cost function and weights.
[0122] Embodiments of this disclosure may include or utilize a dedicated or general-purpose computer including computer hardware, such as one or more processors and system memory, as discussed in more detail below. Embodiments within the scope of this disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Specifically, one or more processes described herein may be implemented at least in part as instructions implemented in a non-transitory computer-readable medium and executable by one or more computing devices, such as any media content access device described herein. Typically, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., memory, etc.) and executes those instructions to perform one or more processes, including one or more processes described herein.
[0123] Computer-readable media can be any available medium accessible by a general-purpose or special-purpose computer system. A computer-readable medium storing computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium carrying computer-executable instructions is a transmission medium. Therefore, by way of example and not limitation, embodiments of this disclosure may include at least two distinct types of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0124] Non-transient computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives (“SSDs”) (e.g., RAM-based), flash memory, phase-change memory (“PCM”), other types of memory, other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of computer-executable instructions or data structures and is accessible by a general-purpose or special-purpose computer.
[0125] A “network” is defined as one or more data links that enable the transmission of electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer via a network or other communication connection (hardwired, wireless, or a combination of hardwired and wireless), the computer correctly regards that connection as a transmission medium. A transmission medium may include networks and / or data links that can be used to carry desired program code in the form of computer-executable instructions or data structures and are accessible by general-purpose or special-purpose computers. Combinations of the foregoing should also be included within the scope of computer-readable media.
[0126] Furthermore, upon arrival at various computer system components, program code in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to a non-transitory computer-readable storage medium (device) (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be cached in RAM within a network interface module (e.g., a NIC) and then ultimately transferred to the computer system RAM and / or a less volatile computer storage medium (device) at the computer system. Therefore, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize the transmission medium.
[0127] Computer-executable instructions include, for example, instructions and data that, when executed at a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a particular function or set of functions. In some embodiments, executing the computer-executable instructions on a general-purpose computer transforms the general-purpose computer into a special-purpose computer that implements the elements of this disclosure. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions described above. Rather, the described features and actions are disclosed as exemplary forms for implementing the claims.
[0128] Those skilled in the art will understand that this disclosure can be implemented in network computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframes, mobile phones, PDAs, tablets, pagers, routers, switches, etc. This disclosure can also be implemented in distributed system environments, where both local and remote computer systems perform tasks via network links (or via hardwired data links, wireless data links, or a combination of hardwired and wireless data links). In a distributed system environment, program modules can reside in local and remote memory storage devices.
[0129] Embodiments of this disclosure can also be implemented in a cloud computing environment. In this specification, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be used in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization, released with minimal management effort or service provider interaction, and then scaled accordingly.
[0130] Cloud computing models can consist of various features, such as on-demand self-service, broadband network access, resource pooling, rapid elasticity, and measurement services. Cloud computing models can also expose various service models, such as Software as a Service (“SaaS”), Platform as a Service (“PaaS”), and Infrastructure as a Service (“IaaS”). Cloud computing models can also be deployed using different deployment models, such as private cloud, community cloud, public cloud, and hybrid cloud. In this specification and claims, a “cloud computing environment” refers to an environment employing cloud computing.
[0131] Figure 12 An example computing device 1200 (e.g., computing device 1000, client device 108, and / or (multiple) servers 104) is illustrated in block diagram form and can be configured to perform one or more of the processes described above. It will be understood that the boot repair system 102 may include an implementation of the computing device 1200. As... Figure 12 As shown, the computing device 1200 may include a processor 1202, a memory 1204, a storage device 1206, an I / O interface 1208, and a communication interface 1210. Furthermore, the computing device 1200 may include input devices such as a touchscreen, mouse, and keyboard. In some embodiments, the computing device 1200 may include a... Figure 12 The components shown are fewer or more components. A more detailed description will follow. Figure 12 Components of the computing device 1200 shown.
[0132] In a particular embodiment, processor(s) 1202 includes hardware for executing instructions, such as instructions that constitute a computer program. By way of example and not limitation, in order to execute instructions, processor(s) 1202 may retrieve (or fetch) instructions from internal registers, internal caches, memory 1204, or storage device 1206, and decode and execute them.
[0133] Computing device 1200 includes memory 1204 coupled to processor(s) 1202. Memory 1204 can be used to store data, metadata, and programs executed by processor(s). Memory 1204 may include one or more of volatile and non-volatile memory, such as random access memory (“RAM”), read-only memory (“ROM”), solid-state drive (“SSD”), flash memory, phase-change memory (“PCM”), or other types of data storage. Memory 1204 may be internal or distributed memory.
[0134] Computing device 1200 includes storage device 1206, which includes storage for storing data or instructions. By way of example and not limitation, storage device 1206 may include the aforementioned non-transient storage media. Storage device 1206 may include hard disk drives (HDDs), flash memory, universal serial bus (USB) drives, or combinations of these or other storage devices.
[0135] The computing device 1200 also includes one or more input or output (“I / O” devices / interfaces 1208 provided to allow a user to provide input (such as user strokes) to the computing device 1200, receive output from the computing device 1200, and otherwise transmit data to and from the computing device 1200. These I / O devices / interfaces 1208 may include a mouse, keypad or keyboard, touchscreen, camera, optical scanner, network interface, modem, other known I / O devices, or combinations of these I / O devices / interfaces 1208. The touchscreen can be activated using a writing device or a finger.
[0136] I / O device / interface 1208 may include one or more devices for presenting output to a user, including but not limited to a graphics engine, a display (e.g., a screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In some embodiments, device / interface 1208 is configured to provide graphical data to the display for presentation to the user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may serve a particular implementation.
[0137] The computing device 1200 may also include a communication interface 1210. The communication interface 1210 may include hardware, software, or both. The communication interface 1210 provides one or more interfaces for communication (e.g., packet-based communication) between the computing device and one or more other computing devices 1200 or one or more networks. By way of example and not limitation, the communication interface 1210 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks such as Wi-Fi. The computing device 1200 may also include a bus 1212. The bus 1212 may include hardware, software, or both for coupling components of the computing device 1200 to each other.
[0138] In the foregoing description, the invention has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the invention have been described with reference to the details discussed herein, and various embodiments are illustrated in the accompanying drawings. The foregoing description and drawings illustrate the invention and should not be construed as limiting the invention. Numerous specific details have been described to provide a thorough understanding of various embodiments of the invention.
[0139] This invention may be practiced in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered illustrative in all respects only, and not restrictive. For example, the methods described herein may be performed with fewer or more steps / actions, or the steps / actions may be performed in a different order. Furthermore, the steps / actions described herein may be performed repeatedly or in parallel with each other, or in parallel with different instances of the same or similar steps / actions. Therefore, the scope of the invention is indicated by the appended claims, and not by the foregoing description. All modifications within the meaning and equivalent scope of the claims should be included within their scope.
Claims
1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause a computing device to replace pixels within a region of a digital image to generate a modified digital image by: generating, with a inpainting neural network, an inpainted digital image from the digital image, the inpainted digital image comprising an initial set of replacement pixels used to fill the region of the digital image; generating, with a visual guidance algorithm, a deep visual guidance that informs patch matching inpainting from the inpainted digital image; identifying replacement pixels for inpainting the region of the digital image by identifying, with a patch matching model having a weighted cost function, the replacement pixels from a weighted combination of structure image guidance, image depth guidance, and segmentation image guidance indicated by the deep visual guidance for potential replacement pixels, wherein the weighted cost function is based on a modification to an initial cost function of the patch matching model representing a relationship between the pixels within the region of the digital image and the potential replacement pixels within an additional region of the digital image; and generating the modified digital image with the patch matching model by replacing the region of the digital image with the replacement pixels.
2. The non-transitory computer-readable medium of claim 1, wherein generating the deep visual guidance comprises generating a deep visual guidance image indicating one or more of: a structure within the region of the digital image from the structure image guidance, a depth from the image depth guidance, or a semantic segmentation from the segmentation image guidance.
3. The non-transitory computer-readable medium of claim 2, wherein identifying the replacement pixels comprises identifying, with the patch matching model, the potential replacement pixels within the digital image corresponding to the structure, the depth, or the semantic segmentation within the region of the digital image indicated by the deep visual guidance.
4. The non-transitory computer-readable medium of claim 1, wherein generating the deep visual guidance comprises utilizing the visual guidance algorithm comprising a structure image model to generate the structure image guidance from the inpainted digital image for identifying the replacement pixels from one or more structures identified within the inpainted digital image.
5. The non-transitory computer-readable medium of claim 1, wherein generating the deep visual guidance comprises utilizing the visual guidance algorithm comprising an image depth neural network to generate the image depth guidance from the inpainted digital image for identifying the replacement pixels from a depth map of the inpainted digital image.
6. The non-transitory computer-readable medium of claim 1, wherein generating the deep visual guidance comprises utilizing the visual guidance algorithm comprising a segmentation image neural network to generate the segmentation image guidance from the inpainted digital image for identifying the replacement pixels from a semantic segmentation of the inpainted digital image.
7. The non-transitory computer-readable medium of claim 1, further comprising instructions that, when executed by the at least one processor, cause the computing device to generate the weighted cost function of the patch matching model by modifying an initial cost function of the patch matching model to include one or more of the potential replacement pixels indicated by the depth visual guidance for comparison.
8. The non-transitory computer-readable medium of claim 1, further comprising instructions that, when executed by the at least one processor, cause the computing device to utilize a student neural network to predict one or more of depth or segmentation within the region of the digital image from parameters learned from a teacher neural network for the depth visual guidance.
9. A system comprising: one or more memory devices comprising: a inpainting neural network; a visual guidance algorithm comprising at least one of a structure image model, an image depth neural network, or a segmentation image neural network; a patch matching model; and a digital image comprising a region of pixels to be replaced; and one or more computing devices configured to cause the system to: generate, utilizing the inpainting neural network, a inpainted digital image from the digital image, the inpainted digital image comprising an initial set of replacement pixels used to fill a region of the digital image; generate, utilizing the visual guidance algorithm, depth visual guidance informing patch matching inpainting from the inpainted digital image, the depth visual guidance comprising at least one of structure image guidance, image depth guidance, and segmentation image guidance; and identify replacement pixels for inpainting the region of the digital image by utilizing the patch matching model having a weighted cost function representing a weighted combination of structure image guidance, image depth guidance, and segmentation image guidance indicated by the depth visual guidance for potential replacement pixels based on a modification to an initial cost function of the patch matching model between pixels within the region of the digital image and the potential replacement pixels within an additional region of the digital image.
10. The system of claim 9, wherein the one or more computing devices are further configured to cause the system to identify the replacement pixels from the digital image by utilizing the patch matching model to identify one or more of: pixels of the digital image corresponding to a structure of the region of pixels utilizing the structure image guidance; pixels of the digital image having a depth corresponding to a depth of the region of pixels utilizing the image depth guidance; or pixels of the digital image corresponding to a segmentation of the region of pixels utilizing the segmentation image guidance.
11. The system of claim 9, wherein the one or more computing devices are further configured to cause the system to generate a modified digital image by utilizing the replacement pixels to replace the region of the digital image, wherein the modified digital image has a higher resolution than the inpainted digital image. 12. The system of claim 9, wherein the one or more computing devices are further configured to cause the system to generate the image depth guidance from the inpainted digital image according to the image depth neural network of the visual guidance algorithm.
13. The system of claim 12, wherein the one or more computing devices are further configured to cause the system to generate the segmentation image guidance from the inpainted digital image according to the segmentation image neural network of the visual guidance algorithm.
14. The system of claim 9, wherein the one or more computing devices are further configured to cause the system to generate the structural image guidance from the structural image model to determine edges between different structural components of the inpainted digital image.
15. The system of claim 9, wherein the one or more computing devices are further configured to cause the system to: generate a binary mask indicative of the region of the digital image; and generate the depth visual guidance based on the binary mask to predict structures within the region of the digital image by utilizing a generator neural network.
16. A computer-implemented method comprising: receiving a digital image, the digital image including a region of pixels to be replaced; generating, with an inpainting neural network, an inpainted digital image from the digital image, the inpainted digital image including an initial set of replacement pixels used to fill the region of the digital image; generating, with a visual guidance algorithm, depth visual guidance informing patch matching inpainting from the digital image, the depth visual guidance including structural image guidance, image depth guidance, and segmentation image guidance from the inpainted digital image; identifying replacement pixels for inpainting the region of the digital image by utilizing a patch matching model having a weighted cost function that identifies the replacement pixels according to a weighted combination of structural image guidance, image depth guidance, and segmentation image guidance indicated by the depth visual guidance for potential replacement pixels within the region of the digital image, wherein the weighted cost function is based on a modification to an initial cost function of the patch matching model representing a relationship between pixels within the region of the digital image and the potential replacement pixels within an additional region of the digital image; and generating a modified digital image with the patch matching model by replacing the region of the digital image with the replacement pixels.
17. The computer-implemented method of claim 16, wherein generating the depth visual guidance comprises: generating an intermediate digital image indicative of structures for a portion of the digital image outside of the region; and generating a binary mask indicative of the region of the inpainted digital image.
18. The computer-implemented method of claim 17, wherein generating the depth visual guidance comprises utilizing a generator neural network to predict structures within the region of the digital image from the inpainted digital image, the intermediate digital image, and the binary mask.
19. The computer-implemented method of claim 16, wherein generating the deep visual guidance comprises utilizing a student neural network to predict one or more of depth as an image depth guide or segmentation as a segmentation image guide within the region of the inpainted digital image from parameters learned from a teacher neural network.
20. The computer-implemented method of claim 19, further comprising learning the parameters for the student neural network by: determining, utilizing the teacher neural network, a teacher label for one or more of depth or segmentation for a complete digital image; generating, from the complete digital image, an incomplete digital image comprising a hole region; generating, utilizing the student neural network, a predicted label for one or more of the depth or the segmentation for the incomplete digital image; and modifying parameters of the student neural network by comparing the predicted label to the teacher label.