Method, system, and computer readable medium for digital image editing

The Latent Vector Image Editing System generates image difference metrics through web-based mediation and image modification neural networks, solving the problems of excessive computational resource consumption and complex user interaction in existing technologies. It enables real-time image editing on devices with limited computing power, improving system efficiency and flexibility.

CN114972574BActive Publication Date: 2025-11-04ADOBE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111411518.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-02-23
Filing Date
2021-11-25
Publication Date
2025-11-04
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

Existing image editing systems consume excessive computational resources when using generative adversarial networks (GANs) to modify digital images, making it impossible to edit images in real time on devices with limited computing power. Furthermore, the complex user interaction limits the system's flexibility and efficiency.

Method used

The latent vector image editing system utilizes a web-based intermediary to modify the latent vectors of digital images, combines an image modification neural network to generate an image difference metric, and displays the image modifications in real time on client devices through a distributed architecture, reducing user interaction and computational resource consumption.

Benefits of technology

It enables real-time image editing on devices with limited computing power, improving system efficiency and flexibility, reducing user interaction, and supporting real-time image editing based on neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972574B_ABST
    Figure CN114972574B_ABST
Patent Text Reader

Abstract

Real-time editing of digital images based on the web involving latent vector stream renderer and image modification neural network. The present disclosure describes systems, methods, and non-transitory computer-readable media for detecting user interactions to edit a digital image from a client device and modifying the digital image of the client device by using a web-based intermediary of latent vectors that modify the digital image and an image modification neural network to generate a modified digital image from the modified latent vectors. In response to user interactions to modify the digital image, the disclosed systems modify latent vectors extracted from the digital image to reflect the requested modifications. The disclosed systems also use a latent vector stream renderer (as an intermediary device) to generate an image delta that indicates the difference between the digital image and the modified digital image. The disclosed systems then provide the image delta as part of a digital stream to the client device to quickly render the modified digital image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the field of image processing, and in particular, to web-based real-time editing of digital images utilizing latent vector streamer and image modification neural networks. BACKGROUND

[0002] In recent years, computer engineers have developed software and hardware platforms for modifying digital images using various models, such as neural networks, including generative adversarial networks (“GANs”). Based on such developments, some conventional image editing systems are able to modify digital images by extracting features from the digital images and combining the extracted features with features extracted from other digital images. Other conventional systems are able to modify digital images by performing GAN-based operations to adjust certain features corresponding to specific GAN-based visual attributes, such as age, anger, surprise, or happiness. However, despite these advances, many conventional image editing systems often require an excessive amount of computing resources to modify digital images using GANs. As a result, conventional systems are often unable to modify images using GANs in real-time on certain computing devices, and often limit such image editing or generation to devices with powerful processors.

[0003] As just presented, conventional image editing systems often inefficiently consume computing resources when extracting features from digital images (or modifying digital images) using neural networks. Specifically, when modifying digital images using GANs or other neural networks, conventional systems sometimes waste processing time, processing power, and memory. For example, some conventional systems generate (and transmit for display) entirely new digital images for each editing operation, such that each new digital image results from the most recent editing operation. The computational cost of generating and transmitting a single image for a single edit, much less editing tens or hundreds of images in a single session, can require a significant amount of computer processing.

[0004] Due to the computational inefficiencies for each successive editing operation, some conventional digital image editing systems only modify digital images at a low speed when performing neural networks using local processors. In fact, conventional systems that utilize local GANs to modify digital images on particular computing devices, such as mobile or simple laptop devices, are often too slow for real-time applications. Except for very conventional systems running on computers with powerful graphics processing units (“GPUs”), using GANs to modify digital images takes a significant amount of time, thereby eliminating the possibility of performing such modifications as part of an interactive (on-the-fly image editing).

[0005] Due at least in part to their computational intensity and their speed constraints, many conventional digital image editing systems also stubbornly limit image editing to particular types of image editing or other applications. Specifically, due to the computational requirements of GAN- or other neural network-based image operations, conventional systems often strictly limit applications to particularly powerful computing devices. Thus, not only are conventional systems unable to perform real-time editing on many client devices with neural networks, the computational expense of these systems often prevents their application on less powerful devices, such as mobile devices.

[0006] As yet another example of inefficiency, some conventional image editing systems provide inefficient graphical user interfaces that require excessive user interaction to access desired data and / or functionality. For example, to implement GAN-based digital image modification, some conventional systems require multiple user interactions to manually select and edit particular portions (or attributes) of a digital image. In some cases, processing such a large number of user interactions wastes computational resources, such as processing power and memory that could otherwise be conserved with fewer user interactions. SUMMARY

[0007] The present disclosure describes one or more embodiments of systems, methods, and non-transitory computer-readable media that address one or more of the foregoing or other issues in the art. The disclosed systems are able to detect user interactions to edit a digital image from a client device and modify the digital image for the client device by using a web-based intermediary that modifies a latent vector of the digital image and using an image modification neural network to generate a modified digital image from the modified latent vector. In response to user interactions to modify the digital image, the disclosed systems modify a latent vector extracted from the digital image to reflect the requested modifications, for example. Based on the modified latent vector, the disclosed systems utilize an image modification neural network, such as a generative adversarial network (“GAN”), to generate a modified digital image. The disclosed systems also use a latent vector streamer (as an intermediary device) to generate an image delta or difference metric that indicates a difference between the digital image and the modified digital image. The disclosed systems then provide the image delta as part of a digital stream to the client device to quickly render the modified digital image. In some embodiments, the disclosed systems also generate and provide an efficient user image modification interface that requires relatively less user interaction for performing neural network-based operations, such as GAN-based operations, on the digital image.

[0008] Additional features and advantages of one or more embodiments of the present disclosure are outlined in the description that follows, and will in part be apparent from the description, or can be learned by practice of such example embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0009] The present disclosure describes one or more embodiments of the present application with additional specificity and detail through reference to the drawings. The following paragraphs briefly describe a few illustrative embodiments of the drawings, wherein:

[0010] Figure 1 illustrates an example system environment in which a latent vector image editing system operates in accordance with one or more embodiments;

[0011] Figure 2 illustrates an overview of modifying digital images using a latent vector approach to determine image differential metrics in accordance with one or more embodiments;

[0012] Figures 3A-3C illustrates a wired diagram of various actions performed by a client device, a latent vector stream painter, and an image modification neural network in accordance with one or more embodiments;

[0013] Figure 4 illustrates an example process for extracting latent image vectors in accordance with one or more embodiments;

[0014] Figure 5 illustrates an example process for generating modified digital images from modified latent image vectors in accordance with one or more embodiments;

[0015] Figure 6 illustrates an example process for determining and providing image differential metrics in accordance with one or more embodiments;

[0016] Figure 7 illustrates an example distributed architecture of a latent vector image editing system in accordance with one or more embodiments;

[0017] Figures 8A-8B illustrates an image modification interface including an additional digital image grid in accordance with one or more embodiments;

[0018] Figures 9A-9B illustrates an image modification interface including a slider tool in accordance with one or more embodiments;

[0019] Figures 10A-10B illustrates an image modification interface including a timeline tool in accordance with one or more embodiments;

[0020] Figures 11A-11B illustrates an image modification interface including a collage tool in accordance with one or more embodiments;

[0021] Figures 12A-12B illustrates an image modification interface including a sketch tool in accordance with one or more embodiments;

[0022] Figure 13FIG. illustrates a schematic diagram of a latent vector image editing system, in accordance with one or more embodiments;

[0023] Figure 14 FIG. illustrates a flow diagram of a series of acts for generating and providing an image difference metric by comparing a digital image associated with a latent image vector, in accordance with one or more embodiments; and

[0024] Figure 15 FIG. illustrates a block diagram of an example computing device, in accordance with one or more embodiments. DETAILED DESCRIPTION

[0025] The present disclosure describes one or more embodiments of a latent vector image editing system that detects user interactions to edit a digital image on a client device and modifies the digital image for the client device by using a web-based intermediary to modify a latent vector of the digital image and using an image modification neural network to generate a modified digital image from the modified latent vector. In particular, the latent vector image editing system receives an indication of a user interaction to modify a digital image. Based on the user interaction, in some embodiments, the latent vector image editing system also determines an image difference metric that reflects the modification. For example, the latent vector image editing system utilizes a novel latent vector flow renderer to generate an image difference metric by comparing an initial digital image with a digital image modified via an image modification neural network. In some cases, the latent vector image editing system provides the image difference metric to the client device to enable the client device to render and display the modified digital image in real-time (or near real-time) relative to the user interaction, even in cases where the client device is a mobile device.

[0026] As suggested above, in one or more embodiments, the latent vector image editing system receives an indication of a user interaction to modify a digital image using a neural network-based operation (e.g., a GAN-based operation). To support such user interactions, the latent vector image editing system provides a digital image for display within an image modification interface on a client device. In some cases, the latent vector image editing system provides the digital image as part of a digital stream played on the client device (e.g., as a frame in a digital video feed). In some embodiments, the digital stream appears as a still digital image within the image modification interface. The latent vector image editing system receives an indication of a user interaction to modify the digital image from the image modification interface and displays the digital image as part of the digital stream.

[0027] Based on the indication of the user interaction, in some embodiments, the latent vector image editing system generates a modified latent image vector for the digital image. For example, the latent vector image editing system extracts latent features from the digital image using an image modification neural network, such as a GAN, and further modifies the latent image vector based on the user interaction. In some cases, the latent vector image editing system generates an initial latent image vector for an initial digital image selected or uploaded by the client device (e.g., prior to user interactions for modifying the digital image), and subsequently modifies the latent image vector to reflect modifications made to the digital image selected by the user.

[0028] After modifying the latent image vector, in certain embodiments, the latent vector image editing system also generates a modified digital image reflecting changes requested via the user interaction. For example, the latent vector image editing system utilizes the image modification neural network to generate a modified digital image from the modified latent image vector. In one or more embodiments, the latent vector image editing system also determines an image difference metric between the initial digital image and the modified digital image. For example, the latent vector image editing system compares the modified digital image to the initial digital image to generate an image difference metric reflecting changes to the digital image (or latent image vector) resulting from the user interaction.

[0029] In some cases, after generating the image difference metric, the latent vector image editing system provides the image difference metric to the client device. For example, the latent vector image editing system provides the image difference metric to cause the client device to update the digital image and render the modified digital image for display. In one or more embodiments, the latent vector image editing system provides the image difference metric as part of a digital stream including code or instructions causing the client device to render changes to the digital stream. For example, the latent vector image editing system provides the image difference metric as part of a digital stream to instruct the client device to render the modified digital image in place of the initial digital image (e.g., as a subsequent frame) to visually illustrate the changes. Indeed, in certain embodiments, rather than regenerating and providing a completely new digital image for each modification made within the image modification interface, the latent vector image editing system determines and provides the image difference metric to the client device to indicate relatively small changes or increments resulting from the modifications.

[0030] To illustrate an initial digital image or a modified digital image, in some embodiments, the latent vector image editing system provides an image modification interface for display on a client device. For example, the latent vector image editing system provides an image modification interface that includes a digital image (e.g., as part of a digital video feed) and one or more selectable elements for performing a neural network-based operation (e.g., a GAN-based operation) on the digital image. In some cases, the image modification interface includes a grid of additional digital images that are selectable to modify the initial digital image by blending or mixing features associated with a digital image selected from the grid with features of the initial digital image. In other cases, the image modification interface includes slider elements that are selectable to adjust certain image features associated with the initial digital image. In other cases, the image modification interface includes additional or alternative elements for performing a neural network-based operation. Additional details regarding various embodiments of image modification interfaces are provided below with reference to the accompanying figures.

[0031] As suggested above, the latent vector image editing system provides a number of technical advantages over conventional image editing systems. For example, in some embodiments, the latent vector image editing system improves computational efficiency as compared to conventional systems. To elaborate, as compared to conventional systems, the latent vector image editing system is able to generate and provide neural network-based modifications to digital images for display on a client device using less processing time, processing power, and memory. Whereas many conventional systems generate and provide entirely new digital images to visually represent each new neural network-based operation performed, the latent vector image editing system preserves a significant amount of computational resources on the client device by generating and providing image differential metrics that reflect the neural network-based modifications to the client device, rather than using a neural network (such as a GAN) to locally modify the image on the client device. In effect, the latent vector image editing system determines and provides image differential metrics that indicate that the client device can draw modifications to the digital image based on changes or deltas that result from user interactions using less processing power, rather than expending a considerable amount of local processing power to generate or regenerate entirely new digital images.

[0032] To illustrate such improvements in computational efficiency, in some embodiments, the latent vector image editing system utilizes a distributed architecture that includes a latent vector streamer at one computing device and one or more neural network (e.g., GAN) computing devices at another computing device— both separate from the client device. The latent vector streamer stores latent vectors to pass back and forth from the neural network and to facilitate determining image difference metrics to update the digital stream provided to the client device. Thus, the latent vector image editing system saves computational resources by determining and passing image difference metrics to update the digital image for each new operation. In contrast, many conventional systems run GANs or other neural networks on a local GPU, thus requiring significantly more local computational resources to generate and regenerate modified digital images from scratch for each new editing operation. Thus, in contrast to conventional systems that are too slow for real-time interactive editing operations, the latent vector image editing system is able to perform image editing operations for interactive applications on the fly.

[0033] As yet another example of improved efficiency, in some embodiments, the latent vector image editing system provides an efficient graphical user interface that requires less user interaction to access desired data and / or functionality compared to conventional systems. For example, in contrast to conventional systems that require many user interactions to manually select and edit portions of a digital image, the latent vector image editing system reduces the required user interactions by providing an image modification interface that includes a grid of digital images that are selectable to modify an initial digital image using neural network-based operations by blending features. In some embodiments, the latent vector image editing system provides an image modification interface that includes a set of selectable slider elements for modifying various GAN-based image features of a digital image. The image modification interface provided by the latent vector image editing system improves efficiency by reducing user interactions and simplifying the process of modifying a digital image.

[0034] Due to the improved efficiency of the latent vector image editing system, embodiments of the latent vector image editing system also improve speed compared to conventional digital image editing systems. For example, by generating and proving image difference metrics to update a digital image displayed on a client device, the latent vector image editing system not only reduces the computational requirements to provide a modified digital image, but also further improves the speed at which this is done. In some cases, in contrast to many conventional systems, by providing image difference metrics reflecting neural network-based (e.g., GAN-based) image modifications in real-time from user interactions requesting modifications, the latent vector image editing system is fast enough for interactive on-the-fly digital image editing.

[0035] Not only do its more efficient operations provide increased speed, but the latent vector image editing system also provides increased flexibility as compared to conventional digital image editing systems. For example, in contrast to many conventional systems that are limited to operating on particularly powerful computing devices, the latent vector image editing system is able to support neural network-based digital image editing via less powerful computing devices (such as mobile devices) having basic or slow GPUs. Indeed, the latent vector image editing system generates and provides smaller and more easily processed image difference metrics (as compared to entire images used by conventional systems), thereby enabling mobile devices to render neural network-based digital image modifications in an interactive, real-time manner.

[0036] As suggested by the foregoing discussion, the present disclosure utilizes various terminology to describe features and benefits of the latent vector image editing system. Additional details regarding the meaning of these terms as used in the present disclosure are provided infra. In particular, the term“neural network” refers to a machine learning model that can be trained and / or tuned based on inputs to determine a classification or approximate an unknown function. For example, a neural network includes a model of interconnected artificial neurons (e.g., organized in layers) that communicate and learn to approximate a complex function and generate an output (e.g., a generated digital image) based on a plurality of inputs provided to the neural network. In some cases, a neural network refers to an algorithm (or collection of algorithms) that implements deep learning techniques to model high-level abstractions in data.

[0037] Relatedly, the term“image modification neural network” refers to a neural network that extracts latent image vectors from digital images and / or generates digital images from latent image vectors. In particular, an image modification neural network extracts latent or hidden features from a digital image and encodes these features as a latent feature vector. In some cases, an image modification neural network generates or reconstructs a digital image from a latent image vector. In one or more embodiments, an image modification neural network takes the form of a generative adversarial neural network. For example, in some embodiments, an image modification neural network is a generative adversarial network (GAN) as described by Jun-Yan Zhu, Philipp iGAN described in Generative Visual Manipulation on the Natural Image Manifold by Eli Shechtman and Alexei A. Efros, published in European Conference on Computer Vision, pages 597-613, 2016, which is incorporated by reference in its entirety. In other embodiments, the image modification neural network is a StyleGAN, StyleGAN2, RealnessGAN, ProGAN, or any other suitable generative neural network. In some cases, the image modification neural network is a neural network other than a generative neural network, and for example takes the form of a PixelRNN or PixelCNN.

[0038] As used herein, the term “generative adversarial neural network” (sometimes simply “GAN”) refers to a neural network tuned or trained via an adversarial process to generate an output digital image from an input digital image. In some cases, the generative adversarial neural network includes multiple constituent neural networks, such as an encoder neural network and a generator neural network. For example, the encoder neural network extracts a latent code from a digital image. The generator neural network generates a modified digital image by combining the extracted latent code (e.g., from the encoder neural network). In competition with the generator neural network, a discriminator neural network analyzes the digital image generated from the generator neural network to determine whether the generated digital image is real (e.g., from a stored set of digital images) or fake (e.g., not from the stored set of digital images). The discriminator neural network also causes the latent vector image editing system to modify parameters of the encoder neural network and / or the generator neural network to ultimately generate a digital image that fools the discriminator neural network into indicating that the generated digital image is a real digital image.

[0039] As mentioned, the latent vector image editing system extracts a latent image vector from a digital image. As used herein, the term “latent image vector” refers to a vector representing or indicating hidden or latent features of an image feature and / or unobservable attribute of a digital image. For example, the latent image vector includes a digital encoding or representation of the digital image. In some embodiments, the latent image vector includes one or more vector directions. A “vector direction” refers to a direction of the latent image vector that encodes or indicates a particular image feature. For example, one vector direction corresponds to age, while another vector direction corresponds to happiness of a face depicted within the digital image. Thus, modifying the latent image vector in either direction results in a corresponding modification to a neural network-based image feature (e.g., a GAN-based image feature) depicted by the digital image.

[0040] Relatedly, an "initial latent image vector" refers to a latent image vector that is extracted from or otherwise corresponds to an initial (e.g., unmodified) digital image. In contrast, a "modified latent image vector" refers to a latent image vector that corresponds to a modified digital image. For example, a modified latent image vector includes one or more modified features that are the result of user interaction for performing a GAN-based operation to edit a digital image.

[0041] As mentioned, in some embodiments, the latent vector image editing system utilizes an image modification neural network to modify a digital image or generate a modified version of a digital image. In some such cases, the image modification neural network constitutes a GAN that performs a GAN-based operation. As used herein, the term "GAN-based operation" refers to a digital image editing operation that utilizes one or more GANs to perform a requested modification. Specifically, a GAN-based operation includes an operation that performs a "GAN-based modification" to edit or change one or more "GAN-based image features" of a digital image. Example GAN-based image features include, but are not limited to, a measure of happiness, a measure of surprise, a measure of age, a measure of anger, and a measure of baldness. Indeed, GAN-based image features are typically more complex and computationally intensive than more conventional digital image modifications such as changing color, cropping, and adjusting brightness. Additional GAN-based visual features are described below with respect to the accompanying figures.

[0042] In certain described embodiments, the latent vector image editing system determines an image differential metric for a particular modification. As used herein, the term "image differential metric" refers to a measure or indication of a difference between a prior digital image and a subsequent or modified version of the digital image (e.g., after modification). Indeed, in some cases, the image differential metric includes an indication of a change or delta between a previously rendered digital image and a modified digital image to be rendered. In some embodiments, the image differential metric includes instructions or computer code (interpretable by a browser or another client application) to cause the client device to update the digital video feed by implementing the changes included within the image differential metric to modify the current digital image (or current frame) with the modified version (of the current frame). For example, the image differential metric indicates a change in a latent image vector in one vector direction or another vector direction to adjust a particular GAN-based image feature.

[0043] In some embodiments, the latent vector image editing system provides the image difference metrics to the client device as part of a digital stream. As used herein, the term "digital stream" refers to the continuous or sequential transmission and / or receipt of one or more data objects (e.g., data packets) from one computing device to another computing device. In some cases, the latent vector image editing system provides the digital stream to the client device to keep the digital image displayed on the client device up-to-date in real-time with respect to the user interactions that are requesting modifications. For example, the digital stream can include data for one or more digital images of a digital video feed. The digital stream can also or alternatively include image difference metrics that indicate changes to the digital images or the digital video feed. In some cases, the latent vector image editing system utilizes a latent vector streamer to generate and provide the digital stream to the client device.

[0044] Additional details regarding the latent vector image editing system will now be provided with reference to the accompanying drawings. For example, Figure 1 FIGURE 1 illustrates a schematic diagram of an example system environment for implementing the latent vector image editing system 102, in accordance with one or more embodiments. An overview of the latent vector image editing system 102 is described with respect to Figure 1 Thereafter, more detailed descriptions of the components and processes of the latent vector image editing system 102 are provided with respect to subsequent drawings.

[0045] As shown, the environment includes server(s) 104, client device 116, database 112, and network 114. Each of the components of the environment communicate via network 114, and network 114 is any suitable network through which computing devices communicate. Example networks are discussed in more detail below with respect to Figure 15 FIGURE 2.

[0046] As mentioned, the environment includes client device 116. Client device 116 is one of a variety of computing devices, including a smartphone, a tablet computer, a smart television, a desktop computer, a laptop computer, a virtual reality device, an augmented reality device, or another computing device described with respect to Figure 15 Indeed, unlike many conventional systems, the latent vector image editing system 102 can operate on a mobile device for interactive, real-time, GAN-based digital image editing. While Figure 1A single instance of a client device 116 is illustrated, in some embodiments, the environment includes multiple different client devices, each associated with a different user (e.g., a digital image editor). The client device 116 communicates with the server(s) 104 via the network 114. For example, the client device 116 receives user input from a user interacting with the client device 116 (e.g., via the client application 118) to, for example, edit, modify, or generate digital content, such as digital images. Accordingly, the latent vector image editing system 102 on the server(s) 104 receives information or instructions to generate a modified digital image.

[0047] As Figure 1 illustrated, the client device 116 includes a client application 118. Specifically, the client application 118 is a web application, a native application (e.g., a mobile application, a desktop application, etc.) installed on the client device 116, or a cloud-based application with all or partial functionality executed by the server(s) 104. The client application 118 presents or displays information to a user, including an image modification interface. For example, a user interacts with the client application 118 to provide user input to select and / or modify one or more digital images.

[0048] As Figure 1 illustrated, the environment includes the server(s) 104. The server(s) 104 generate, track, store, process, receive, and transmit electronic data, such as indications of digital image modifications and user interactions. For example, the server(s) 104 receive data from the client device 116 in the form of indications of user interactions for modifying a digital image. Additionally, the server(s) 104 transmit data to the client device 116 to provide an image difference metric to cause the client device 116 to display or present a modified digital image. In effect, the server(s) 104 communicate with the client device 116 to transmit and / or receive data via the network 114. In some embodiments, the server(s) 104 include a distributed server, where the server(s) 104 include multiple server devices distributed across the network 114 and located in different physical locations. The server(s) 104 include a content server, an application server, a communication server, a web hosting server, a multidimensional server, or a machine learning server.

[0049] As Figure 1Further shown, the server(s) 104 also include, as part of the digital content editing system 106, a latent vector image editing system 102. The digital content editing system 106 is in communication with the client device 116 to perform various functions associated with the client application 118, such as storing and managing a repository of digital images, modifying digital images, and providing modified digital images for display. For example, the latent vector image editing system 102 is in communication with the database 112 to access the repository of digital images and / or to access or store latent image vectors within the database 112. In fact, as Figure 1 Further shown, the environment includes the database 112. Specifically, the database 112 stores information such as a repository of digital images and latent image vectors generated from the digital images.

[0050] Additionally, the latent vector image editing system 102 includes a latent vector stream painter 108. Specifically, the latent vector stream painter 108 is in communication with the client device 116 and the image modification neural network 110. For example, the latent vector stream painter 108 receives an indication of a user interaction for modifying a digital image. Based on the indication, the latent vector stream painter 108 also generates a modified latent image vector, and provides the modified latent image vector to the image modification neural network 110. Additionally, the latent vector stream painter 108 receives a modified digital image from the image modification neural network 110, and determines an image difference metric by comparing the initial digital image and the modified digital image. The latent vector stream painter 108 also provides the image difference metric to the client device 116 (e.g., as part of a digital stream) to cause the client device 116 to paint the modified digital image.

[0051] As just mentioned, and as Figure 1 illustrated, the latent vector image editing system 102 also includes the image modification neural network 110. Specifically, the image modification neural network 110 receives and / or provides digital images and / or latent image vectors from the latent vector stream painter 108. For example, the image modification neural network 110 extracts a latent image vector from a digital image received from the latent vector stream painter 108. Additionally, the image modification neural network 110 generates a modified digital image from a modified latent image vector received from the latent vector stream painter 108. In some embodiments, the image modification neural network 110 is a GAN that includes an encoder neural network and a generator neural network.

[0052] Although Figure 1particular arrangement of components, but in some embodiments, the environment has a different arrangement of components and / or can have a wholly different number or set of components. For example, in some embodiments, the latent vector image editing system 102 is implemented by (e.g., located in whole or in part on) the client device 116 and / or a third-party device. In some embodiments, the latent vector flow painter 108 and the image modification neural network 110 are located on the same server(s) 104, while in other embodiments, the latent vector flow painter 108 and the image modification neural network 110 are located on different server devices that are remote from one another (albeit with some geographic distance to maintain communication speed). Additionally, in one or more embodiments, the client device 116 communicates directly with the latent vector image editing system 102, bypassing the network 114. Further, in some embodiments, the database 112 is located external to the server(s) 104 (e.g., in communication via the network 114) or on the server(s) 104 and / or the client device 116.

[0053] As mentioned, in one or more embodiments, the latent vector image editing system 102 generates and provides modified digital images for display with a latent vector approach that supports real-time implementation of neural network-based modifications (e.g., GAN-based modifications). Specifically, the latent vector image editing system 102 determines an image difference metric that indicates and causes a client device to present changes from an initial digital image to a modified digital image. Figure 2 An overview of determining an image difference metric for real-time presentation of digital image modifications (e.g., GAN-based modifications) on a client device is illustrated in accordance with one or more embodiments. Additional details regarding various acts of Figure 2 are provided with respect to subsequent figures.

[0054] As Figure 2 illustrated, the latent vector image editing system 102 performs act 202 to receive a digital image. Specifically, the latent vector image editing system 102 receives a digital image from the client device 116 in the form of a selected digital image, a captured digital image, or an uploaded digital image. In some embodiments, the latent vector image editing system 102 identifies or accesses a digital image from a repository of stored digital images (e.g., within the database 112). For example, the latent vector image editing system 102 receives an indication of a user selection of a particular digital image via the client device 116.

[0055] After receiving the digital image, the latent vector image editing system 102 performs act 204 to extract a latent image vector from the digital image. More specifically, the latent vector image editing system 102 utilizes an image modification neural network (e.g., the image modification neural network 110) to extract a latent image vector that includes encoded features of the digital image. As shown, the latent vector image editing system 102 extracts a latent image vector, denoted as [v].

[0056] As Figure 2 Further illustrated, the latent vector image editing system 102 performs act 206 to receive a user interaction for modifying the digital image. Specifically, the latent vector image editing system 102 receives an indication of a user interaction from the client device 116 for performing a GAN-based operation to edit or modify the digital image. As Figure 2 illustrated, for example, the latent vector image editing system 102 receives an indication of a user interaction for having the initial digital image slide over one or more other digital images within a grid underneath the digital image. In some embodiments, the indication includes information reflecting which digital images are underneath the initial digital image (and in what proportion) for blending features of the initial digital image with features of the underlying digital images.

[0057] As depicted in subsequent figures and further described below, in some embodiments, the latent vector image editing system 102 receives indications of different user interactions. For example, in some cases, the latent vector image editing system 102 provides one or more slider tools within the image modification interface that are selectable to adjust image features with a slider element. As another example, the latent vector image editing system 102 provides a timeline tool that includes a slider bar that is selectable to slide over multiple elements at once to simultaneously modify multiple image features. As yet another example, the latent vector image editing system 102 provides a collage tool for selecting portions of additional digital images to assort or combine with the initial digital image (e.g., by combining features). As yet another example, the latent vector image editing system 102 provides a sketch tool for adding different brushstrokes to the initial digital image via a digital brush or smudge, whereupon the latent vector image editing system 102 modifies the additional digital images according to the added brushstrokes.

[0058] As Figure 2Further shown, latent vector image editing system 102 performs action 208 to modify the latent image vector. More specifically, latent vector image editing system 102 modifies or adjusts the latent image vector extracted from the initial digital image based on the user interaction. For example, latent vector image editing system 102 determines, based on the user interaction, a vector direction of the latent image vector to modify or receives an indication thereof. In some embodiments, the vector direction corresponds to a particular image feature (e.g., GAN-based image feature) indicated by the user interaction. In these or other embodiments, latent vector image editing system 102 modifies the latent image vector by combining (e.g., blending, mixing, concatenating, adding, and / or multiplying) one or more portions of the initial latent image vector with one or more portions of an additional latent image vector (e.g., corresponding to an additional digital image).

[0059] For example, latent vector image editing system 102 determines, or receives an indication of, one or more digital images underlying the initial digital image within the digital image grid of the image modification interface. Additionally, latent vector image editing system 102 modifies the latent image vector by combining features of the initial digital image with features of the underlying digital images in the grid.

[0060] In some cases, latent vector image editing system 102 determines, or receives an indication of, respective amounts of overlap or proportions of the initial digital image on one or more underlying digital images. Based on their respective amounts of overlap, latent vector image editing system 102 weights features of the underlying digital images when combining them with the initial latent image vector to generate the modified latent image vector. As Figure 2 As shown, for example, latent vector image editing system 102 weights features of the digital images overlapped by the initial digital image with weights wl, w2, w3, and w4. Latent vector image editing system 102 thus generates the modified latent image vector [v’].

[0061] In some embodiments, latent vector image editing system 102 determines additional or alternative modifications to the initial latent image vector. For example, latent vector image editing system 102 determines to modify the initial latent image vector in a particular vector direction indicated by the user interaction to modify the digital image. In some cases, latent vector image editing system 102 modifies the initial latent image vector in a vector direction corresponding to a particular image feature (e.g., GAN-based image feature). For example, latent vector image editing system 102 receives a user interaction to increase a measure of happiness, and latent vector image editing system 102 adds or multiplies features of the initial latent image vector in a vector direction corresponding to adjusting a smile (while maintaining other features of the latent image vector the same). Additional details regarding various user interactions and corresponding changes to the latent image vector are provided below with reference to subsequent figures.

[0062] like Figure 2 As further illustrated, the latent vector image editing system 102 performs action 210 to generate a modified digital image. Specifically, the latent vector image editing system 102 utilizes an image modification neural network (e.g., image modification neural network 110) to generate the modified digital image. For example, the latent vector image editing system 102 inputs a modified latent image vector ([v']) into the image modification neural network 110, and the image modification neural network 110 generates the modified digital image. Figure 2 As shown, the latent vector image editing system 102 generates a modified digital image from a modified latent image vector, which is modified to reflect the sliding of the initial digital image across four underlying digital images. In fact, Figure 2 The modified digital image depicted includes features from the initial digital image and a (proportional) combination of features from four digital images superimposed on the initial digital image.

[0063] Also Figure 2 As illustrated, the latent vector image editing system 102 performs action 212 to generate an image difference metric. More specifically, the latent vector image editing system 102 generates the image difference metric by comparing an initial digital image (e.g., received in action 202) with a modified digital image (e.g., generated in action 210). In some embodiments, the latent vector image editing system 102 compares the initial digital image by comparing corresponding latent vectors associated with the image (e.g., by subtracting a vector from another vector).

[0064] By comparing digital images, the latent vector image editing system 102 determines a difference α between an initial digital image and a modified digital image. For example, the latent vector image editing system 102 determines a vector difference or a pixel difference. In any case, the image difference metric includes information indicating the apparent difference between the initial digital image and the modified digital image. In some embodiments, the image difference metric includes only information about the differences between the digital images, such as differences in latent image vectors or corresponding pixels. In these or other embodiments, the image difference metric does not include all information from a completely new digital image or a completely new digital video feed.

[0065] Additionally, the latent vector image editing system 102 performs action 214 to provide an image difference metric to the client device 116. Specifically, the latent vector image editing system 102 provides an image difference metric for drawing a modified digital image for display on the client device 116. Thus, the latent vector image editing system 102 causes the client device 116 to change from displaying an initial digital image to displaying a modified digital image.

[0066] In some cases, the latent vector image editing system 102 causes the client device 116 to update a digital video feed from a preceding frame depicting an initial digital image to a subsequent frame depicting a modified digital image. Indeed, the image difference metric can include causing the client device 116 to render updated information by modifying the appearance of the initial digital image with the image difference metric (e.g., without streaming or rendering all information as a brand new digital image). As Figure 2 shown, the image difference metric is provided to cause the client device 116 to change the appearance of the initial digital image to the appearance of the modified digital image.

[0067] While Figures 3A-3C (and subsequent figures) illustrate generating and modifying digital images in a particular domain (e.g., faces), this is merely exemplary. Indeed, the latent vector image editing system 102 is also applicable to other domains. For example, the latent vector image editing system 102 can modify and provide digital images depicting a variety of subjects, including cars, buildings, people, landscapes, animals, furniture, or food. The principles, methods, and techniques of the latent vector image editing system 102 described herein are applicable to any domain or subject of digital images.

[0068] As suggested above, in one or more embodiments, the latent vector image editing system 102 leverages a multi-faceted architecture of computing devices at different network locations (and / or different physical / geographical locations) to work together to provide real-time digital image modification to a client device. Specifically, in some cases, the latent vector image editing system 102 leverages the latent vector stream renderer 108 and the image modification neural network 110 located at the same server or different servers across a network to perform different actions of the latent vector image editing system 102, respectively. Figure 3A An example diagram illustrating various actions performed by the latent vector stream renderer 108, the image modification neural network 110, and the client device 116 is shown in accordance with one or more embodiments.

[0069] As Figure 3A shown, the client device 116 provides data of the digital image 302 (or an indication thereof) to the latent vector stream renderer 108. Specifically, the client device 116 captures, uploads, or selects the digital image 302. Networked, the client device 116 also provides the digital image 302 or an indication of the digital image 302 to the latent vector stream renderer 108. In turn, the latent vector stream renderer 108 provides the digital image 302 to the image modification neural network 110, which then performs the action 304 to extract a latent image vector from the digital image 302. For example, as described above, the image modification neural network 110 extracts or encodes features corresponding to the visible and unobservable qualities of the digital image 302 into a latent image vector.

[0070] As Figure 3A Further illustrated, the latent vector image editing system 102 passes or sends the latent image vectors extracted from the image modification neural network 110 to the latent vector streamer 108. In one or more embodiments, the latent vector streamer 108 performs act 306 to store the latent image vectors within the database 112. In effect, the latent vector image editing system 102 stores the latent image vectors for later use. In some cases, the latent vector image editing system 102 extracts and stores the latent image vectors for a digital image repository (e.g., a repository associated with the digital content editing system 106).

[0071] In addition to storing the latent image vectors, the latent vector streamer 108 provides a digital stream 308 to the client device 116. More specifically, the latent vector streamer 108 provides the digital stream 308 that includes information for rendering the digital image 302 for display within an image modification interface on the client device 116. In some cases, the digital stream 308 includes a digital video feed of the digital image 302. Accordingly, the client device 116 presents or displays the digital image 302 (e.g., within the digital video feed) with the image modification interface.

[0072] As Figure 3A Further illustrated, the client device 116 also performs act 310 to detect a user interaction. Specifically, the client device 116 detects or receives a user input within the image modification interface to modify or edit the digital image 302. For example, the client device 116 receives a user interaction to perform a GAN-based operation to modify the digital image 302.

[0073] In response to the user interaction, the client device 116 provides a user interaction indication 312 to the latent vector streamer 108. In effect, the client device 116 provides an indication of the user interaction to modify the digital image. In some cases, the user interaction indication 312 includes information that requests or indicates a command for one or more GAN-based image features of the modified digital image 302 (e.g., via a GAN-based operation).

[0074] In response to receiving the user interaction indication 312, the latent vector stream painter 108 performs an action 314 to modify the latent image vector. More specifically, the latent vector stream painter 108 modifies the latent image vector of the digital image 302 that was extracted by the image modification neural network 110 and stored within the database 112. In practice, the latent vector stream painter 108 accesses the latent image vector from the database 112 and modifies the latent image vector using one or more modification operations. Such modification operations include adding or multiplying a portion of the latent image vector that corresponds to a particular vector direction and / or combining features of the latent image vector with features of one or more additional latent image vectors.

[0075] In some embodiments, the latent vector stream painter 108 includes logic to progressively project data of the modified digital image while concurrently manipulating the latent image vector. To elaborate, while in communication with the image modification neural network 110 to extract the latent image vector, the latent vector stream painter 108 concurrently modifies the latent image vector. In some cases, it takes approximately 10 to 12 seconds to extract the latent image vector from the digital image for 100 iterations. However, rather than requiring the encoding process to complete for each digital image modification before applying the corresponding transformation and allowing the user to further manipulate the digital image 302, the latent vector stream painter 108 is able to concurrently modify the latent image vector without waiting for the process of extracting a new latent image vector to complete. Thus, even while the image modification neural network 110 is still extracting the latent image vector, the latent vector stream painter 108 enables user interactions via the client device 116 and concurrently modifies the latent image vector based on the user interactions. Additional details regarding the modification operations used to modify the latent image vector are provided below with reference to subsequent figures.

[0076] As Figure 3B Further illustrated, the latent vector image editing system 102 passes or sends the modified latent image vector 316 from the latent vector stream painter 108 to the image modification neural network 110. In turn, the image modification neural network 110 performs an action 318 to generate a modified digital image. To elaborate, the image modification neural network 110 generates the modified digital image from the modified latent image vector 316. In practice, the image modification neural network 110 processes the modified latent image vector 316 by performing a GAN-based operation to generate the modified digital image, which results in a modification to one or more GAN-based image features of the digital image 302.

[0077] As Figure 3BContinuing, the latent vector image editing system 102 passes or sends the modified digital image 320 from the image modification neural network 110 to the latent vector streamer 108. Based on the modified digital image 320, the latent vector streamer 108 performs an action 322 to determine an image differential metric. More specifically, the latent vector streamer 108 compares the modified digital image 320 to the digital image 302. For example, the latent vector streamer 108 determines the differences between the pixels of the modified digital image 320 and the pixels of the digital image 302. In some embodiments, the latent vector streamer 108 compares the latent image vectors to determine the differences between the modified latent image vector 316 and the latent image vector extracted in action 304. Accordingly, the latent vector streamer 108 generates an image differential metric 324 that includes instructions or information for rendering the changes to the digital image 302 that was initially displayed on the client device 116.

[0078] As Figure 3B Further shown, the latent vector streamer 108 provides the image differential metric 324 to the client device 116. By providing the image differential metric 324, the latent vector streamer 108 causes the client device 116 to display or render the modified digital image 320. Specifically, the image differential metric 324 includes instructions for the client device 116 to modify the presentation of the digital image 302 by changing the pixels to resemble the modified digital image 320. In some cases, the latent vector streamer 108 provides the image differential metric 324 as part of the digital stream. For example, the latent vector streamer 108 maintains the same digital stream 308 provided to the client device 116 (continuously or continuously) and updates the digital stream 308 with new information based on user interactions with the modified digital image. For example, the latent vector streamer 108 provides the image differential metric 324 within the digital stream 308, thereby causing the client device 116 to update the presentation of the digital image 302 to display the modified digital image 320.

[0079] Upon receiving the image differential metric 324, the client device 116 performs an action 326 to render the modified digital image. Specifically, the client device 116 receives the image differential metric 324 and renders the modified digital image 320 in place of the digital image 302. For example, the client device 116 interprets the instructions of the image differential metric 324 that indicate how to modify the presentation of the digital image 302 to transform the presentation of the digital image 302 to the presentation of the modified digital image 320.

[0080] As Figure 3BAs illustrated, the client device 116 also performs act 328 to detect additional user interactions. In particular, the client device 116 detects or receives user interactions to further modify the modified digital image 320. In response to the user interactions, the client device 116 provides an indication 330 of the user interactions to the latent vector stream painter 108. In practice, in some embodiments, the client device 116 provides an indication of the user interactions that request further modification of the modified digital image 320, e.g., by performing GAN-based operations.

[0081] As Figure 3C Further, in turn, the latent vector stream painter system 108 performs act 332 to further modify the latent image vector. In particular, the latent vector stream painter 108 modifies the modified latent image vector 316 based on the user interactions. For example, the latent vector stream painter 108 modifies the vector to combine features with one or more other latent image vectors and / or adjust the vector direction corresponding to a particular image feature, e.g., a GAN-based image feature.

[0082] After further modifying the latent image vector, the latent vector image editing system 102 passes or sends the further modified latent image vector 334 from the latent vector stream painter 108 to the image modification neural network 110. In response, the image modification neural network 110 performs act 336 to generate a further modified digital image. In particular, the image modification neural network 110 generates the further modified digital image from the further modified latent image vector 334.

[0083] As Figure 3C Continuing as in the middle, the latent vector image editing system 102 passes or sends the further modified digital image 338 to the latent vector stream painter 108. Similar to the above discussion, the latent vector stream painter 108 performs act 340 to determine an additional image difference metric. In practice, the latent vector stream painter 108 compares the further modified digital image 338 to the modified digital image 320. In some embodiments, the latent vector stream painter 108 determines a perceptual difference between the further modified digital image 338 and the modified digital image 320 and encodes the difference in the additional image difference metric.

[0084] As Figure 3CAs further shown, in some embodiments, the latent vector stream renderer 108 also provides an additional image difference metric 342 to the client device 116. For example, the latent vector stream renderer 108 provides the additional image difference metric 342 as part of the digital stream 308 (e.g., continuously or persistently provided to the client device 116). In effect, the latent vector stream renderer 108 modifies the digital stream 308 to include the additional image difference metric 342, thereby causing the client device 116 to update its digital video feed from displaying the modified digital image 320 to displaying the additionally modified digital image 338. Figures 3A-3C As shown, client device 116 performs action 344 to draw the additionally modified digital image 338. As mentioned, client device 116 updates the digital video feed by modifying the presentation of the modified digital image 320 to present the additionally modified digital image 338.

[0085] The client device 116, the latent vector flow plotter 108, and the image modification neural network 110 are able to repeat the process in a loop regarding... Figures 3A-3C Any or all of the actions described. For example, based on new user interactions that continuously modify digital images, the latent vector flow plotter 108 receives new instructions and generates new modified latent image vectors. The image modification neural network 110 also generates new modified digital images for display on the client device 116. Therefore, Figures 3A-3C The action can be repeated in a loop until the user interaction stops.

[0086] Although Figures 3A-3C The illustrations show a specific sequence of described actions; additional or alternative sequences are possible. For example, in one or more embodiments, the latent vector stream plotter 108 provides the digital stream 308 to the client device 116 before (or simultaneously with) providing the digital image 302 to the image modification neural network 110 to extract latent image vectors. As another example, in response to multiple user interactions for modifying the digital image, the image modification neural network 110 performs multiple actions to modify the latent image vectors, while concurrently performing actions to extract latent image vectors from the digital image. In some cases, the latent vector image editing system 102 repeats actions in the same or different order. Figure 4 The illustrated action involves continuously editing a digital image through user interaction that requests GAN-based modifications.

[0087] As mentioned above, in some of the described embodiments, the latent vector image editing system 102 generates or extracts latent image vectors from digital images. Specifically, the latent vector image editing system 102 utilizes an encoder as a preceding layer or as part of an image modification neural network (e.g., image modification neural network 110) to extract latent image vectors from digital images. Figure 4An initial digital image 402 is illustrated. The initial digital image 402 is a digital image that is provided to the latent vector image editing system 102 for processing. In some embodiments, the initial digital image 402 is provided by a user of a client device 116. For example, the user of the client device 116 captures the initial digital image 402 using a camera of the client device 116, and uploads the initial digital image 402 to the latent vector image editing system 102. As another example, the user of the client device 116 provides an indication of a user selection of the initial digital image 402, and the latent vector image editing system 102 accesses the initial digital image 402 from a digital image repository (e.g., stored within the database 112).

[0088] As illustrated, the latent vector image editing system 102 identifies the initial digital image 402. Specifically, the latent vector image editing system 102 receives the initial digital image 402 (or an indication of the initial digital image 402) from a client device 116. For example, the client device 116 captures the initial digital image 402, and uploads the initial digital image 402 for access by the latent vector image editing system 102. As another example, the client device 116 provides an indication of a user selection of the initial digital image, and the latent vector image editing system 102 accesses the initial digital image 402 from a digital image repository (e.g., stored within the database 112). Figure 4 As illustrated, the latent vector image editing system 102 identifies the initial digital image 402. Specifically, the latent vector image editing system 102 receives the initial digital image 402 (or an indication of the initial digital image 402) from a client device 116. For example, the client device 116 captures the initial digital image 402, and uploads the initial digital image 402 for access by the latent vector image editing system 102. As another example, the client device 116 provides an indication of a user selection of the initial digital image, and the latent vector image editing system 102 accesses the initial digital image 402 from a digital image repository (e.g., stored within the database 112).

[0089] As illustrated, the latent vector image editing system 102 identifies the initial digital image 402. Specifically, the latent vector image editing system 102 receives the initial digital image 402 (or an indication of the initial digital image 402) from a client device 116. For example, the client device 116 captures the initial digital image 402, and uploads the initial digital image 402 for access by the latent vector image editing system 102. As another example, the client device 116 provides an indication of a user selection of the initial digital image, and the latent vector image editing system 102 accesses the initial digital image 402 from a digital image repository (e.g., stored within the database 112). Figure 5 As illustrated, the latent vector image editing system 102 utilizes the encoder 400 (e.g., as a previous layer or portion thereof of the image modification neural network 110) to analyze the initial digital image 402. More specifically, the latent vector image editing system 102 processes the initial digital image 402 with the encoder 400 to extract features for inclusion within the latent image vector 404. In effect, the latent vector image editing system 102 generates the latent image vector 404 from the initial digital image 402. Thus, the latent image vector 404 includes features that represent visible and / or non-observable hidden features of the digital image 402. In one or more embodiments, the latent vector image editing system 102 also stores the latent image vector 404 in the database 112.

[0090] As mentioned, in some embodiments, the latent vector image editing system 102 generates a modified digital image from the modified latent image vector. Specifically, the latent vector image editing system 102 generates the modified latent image vector based on user interaction with the modified digital image, and the latent vector image editing system 102 also generates the modified digital image from the modified latent image vector. Figure 5 An example process for generating a modified digital image from a modified latent image vector is illustrated.

[0091] As illustrated, the latent vector image editing system 102 receives an indication of user interaction for modifying the initial digital image 504. For example, Figure 5 As illustrated, the latent vector image editing system 102 receives an indication of user interaction for modifying the initial digital image 504. For example, Figure 5The user interaction is illustrated as causing the initial digital image 504 to slide from an initial position on the grid 502 of additional digital images to a new position. In the initial position, the initial digital image 504 includes features of the digital images within the grid 502 that are covered by the initial digital image 504. In response to the user interaction causing the initial digital image 504 to slide to the new position, the latent vector image editing system 102 generates a modified digital image (e.g., a modified version of the initial digital image 504) to include features of the digital images within the grid 502 that are under the digital image at the new position.

[0092] In practice, upon receiving an indication of a user interaction to modify a digital image, the latent vector image editing system 102 performs the action 506 to generate a modified latent image vector. In particular, the latent vector image editing system 102 modifies the latent image vector in accordance with the user interaction. In some cases, the latent vector image editing system 102 modifies the latent image vector by combining features of the latent image vector with features of one or more additional latent image vectors corresponding to additional digital images. For example, the latent vector image editing system 102 combines the latent image vector with one or more additional latent image vectors corresponding to digital images covered by the initial digital image 504 within the grid 502.

[0093] In certain embodiments, the latent vector image editing system 102 proportionally combines the latent image vectors. To elaborate, the latent vector image editing system 102 determines (or receives an indication of) portions, regions, or amounts of the digital images covered by the initial digital image 504 within the grid 502. For example, the latent vector image editing system 102 determines that the initial digital image 504 covers a first portion of a first additional digital image, a second portion of a second additional digital image, a third portion of a third additional digital image, and a fourth portion of a fourth additional digital image (e.g., where the respective portions are the same size or different sizes). Based on the respective covered portions, the latent vector image editing system 102 weights the latent image vectors accordingly.

[0094] For example, in certain implementations, the latent vector image editing system 102 accesses latent image vectors LI, L2, L3, and L4 corresponding to the digital images covered by the initial digital image 504 in the grid 502. In practice, in some cases, the latent vector image editing system 102 accesses a repository of latent image vectors corresponding to digital images within the grid 502. For example, the latent vector image editing system 102 generates and stores the latent image vectors with the image modification neural network 110. In any case, the latent vector image editing system 102 weights the vectors LI, L2, L3, and L4 in accordance with respective covered portions of the vectors (e.g., where image vectors that overlap more with the initial digital image 504 have a higher weight).

[0095] The latent vector image editing system 102 also generates a modified latent image vector by combining the weighted vector of the overlay digital image with the initial latent image vector of the initial digital image 504. For example, in some embodiments, the latent vector image editing system 102 generates a modified latent image vector by combining the latent image vectors according to the following function:

[0096] [v] + w1L1 + w2L2 + w3L3 + w4L4 = [v']

[0097] where [v] represents the initial latent image vector, w i represents the weight corresponding to the additional latent image vector L i of the additional digital image, and [v'] represents the modified latent image vector. In some embodiments, the latent vector image editing system 102 modifies the latent image vector by multiplying the portions of the vector together. For example, instead of adding the latent image vectors together, the latent vector image editing system 102 multiplies the latent image vectors to generate the modified latent image vector.

[0098] While Figure 5 While the particular user interaction of modifying the initial digital image 504 by sliding it to a new location on the grid 502 is illustrated, the latent vector image editing system 102 can implement additional or alternative user interactions. Indeed, as suggested above, in some embodiments, the latent vector image editing system 102 receives an indication of a user interaction for modifying a particular image feature (e.g., a GAN-based image feature) via a slider element or other user interface element. For example, the latent vector image editing system 102 receives an indication of a user interaction via a slider element for increasing or decreasing a measure of happiness of a face within a digital image. In response, the latent vector image editing system 102 determines a vector direction corresponding to the measure of happiness and either multiplies the vector direction or adds or subtracts based on an increase measure requested via the user interaction (e.g., from an initial measure of happiness of 3 to a modified measure of happiness of 10).

[0099] Modifying a measure of happiness is one example of a GAN-based image feature that the latent vector image editing system 102 can modify based on a user interaction with a slider element. Other examples are mentioned above and illustrated in subsequent figures. In any case, the latent vector image editing system 102 determines a vector direction corresponding to the image feature adjusted via the user interaction and modifies the feature of the latent image vector associated with the vector direction in a measure corresponding to the adjustment measure indicated by the user interaction.

[0100] In one or more embodiments, the latent vector image editing system 102 receives an indication of a user interaction with the interactive timeline for modifying multiple neural network-based image features at once. For example, the latent vector image editing system 102 receives an indication that a user adjusted a slidable bar on multiple slider elements at once, where each slider element corresponds to a different GAN-based image feature. Accordingly, the latent vector image editing system 102 generates a modified latent image vector by multiplying or adding or subtracting vector directions of latent image vectors corresponding to adjustments to GAN-based image features requested via the user interaction.

[0101] In certain embodiments, the latent vector image editing system 102 receives an indication of a user interaction with a collage tool for selecting features from one or more additional digital images to combine with the initial digital image 504. For example, the latent vector image editing system 102 receives a selection of a nose from an additional digital image to combine with facial features of the initial digital image 504. In response to the user interaction with the collage tool, the latent vector image editing system 102 generates a modified latent image vector by combining one or more portions of an additional latent image vector from the additional digital image (e.g., portions corresponding to the nose or other selected region) with the latent image vector of the initial digital image.

[0102] In one or more embodiments, the latent vector image editing system 102 receives an indication of a user interaction with a sketch tool of the image modification interface. In particular, the latent vector image editing system 102 receives an indication of one or more strokes made via the sketch tool to add or otherwise alter the initial digital image with a digital paintbrush. For example, the latent vector image editing system 102 determines that the user interaction includes a stroke to add glasses around an eye depicted within the initial digital image 504. In response, the latent vector image editing system 102 generates a modified digital image vector. For example, the latent vector image editing system 102 searches the database 112 to identify digital images depicting glasses (e.g., glasses within a threshold similarity to those added via the stroke of the sketch tool). Additionally, the latent vector image editing system 102 combines features of the additional digital images (e.g., portions of features corresponding to the glasses) with features of the initial latent image vector.

[0103] As Figure 6As further illustrated, in addition to generating the modified latent image vector, the latent vector image editing system 102 generates the modified digital image 508 utilizing the image modification neural network 110. Specifically, the latent vector image editing system 102 generates the modified digital image 508 from the modified latent image vector utilizing the image modification neural network 110. For example, the latent vector image editing system 102 generates the modified digital image 508 that resembles or depicts the modifications requested via the user interaction.

[0104] As mentioned, the latent vector image editing system 102 also causes the client device 116 to render and display the modified digital image 508. Specifically, the latent vector image editing system 102 determines an image differential metric based on the modified digital image 508 and provides the image differential metric to the client device 116 to cause the client device 116 to render the changes to the digital image within the image modification interface. Figure 6 generating an image differential metric and including the image differential metric as part of a digital stream, in accordance with one or more embodiments.

[0105] As Figures 5-6 As illustrated, the latent vector image editing system 102 performs a comparison 606 between the initial digital image 602 and the modified digital image 604. As described above, the latent vector image editing system 102 identifies the initial digital image 602 as the digital image received or indicated by the client device 116. Additionally, as also described above, the latent vector image editing system 102 generates the modified digital image 604 from the initial digital image 602 utilizing the image modification neural network 110.

[0106] Additionally, the latent vector image editing system 102 performs the comparison 606 to determine the differences between the initial digital image 602 and the modified digital image 604. More specifically, the latent vector image editing system 102 determines visible and / or non-observable pixel differences and / or instruction differences between displaying the initial digital image 602 and the modified digital image 604. In some cases, the latent vector image editing system 102 determines the differences by comparing the initial latent image vector to the modified latent image vector (e.g., by subtracting the initial image vector from the modified image vector, or vice versa). As illustrated, the latent vector image editing system 102 performs the comparison 606 by comparing the initial digital image 602 to the modified digital image 604 according to the following function:

[0107] gid_WPS_new = gid_WPS * AWPS + WP (1 - AWPS)

[0108] where gid WPS new represents an array of latent image vectors corresponding to the modified digital image 604, gid WPS represents an array of latent image vectors corresponding to the initial digital image 602, WP represents the initial digital image 602, and AWPS represents the image difference metric 608.

[0109] In fact, as a result of the comparison 606, the latent vector image editing system 102 generates the image difference metric 608. Specifically, the latent vector image editing system 102 generates the image difference metric 608 that indicates or reflects the difference between the initial digital image 602 and the modified digital image 604. As shown, the latent vector image editing system 102 also includes the image difference metric 608 within the digital stream 610.

[0110] More specifically, the latent vector image editing system 102 modifies the digital stream 610 provided to the client device 116 (e.g., as a stream of digital images and / or other data for presenting and editing a digital video feed) by including the image difference metric 608. As illustrated, the latent vector image editing system 102 provides the image difference metric 608 as part of the digital stream 610 to cause the client device 116 to present the transformation of the initial digital image 602 to the modified digital image 604. For example, in some embodiments, the latent vector image editing system 102 causes the client device 116 to render the modified digital image 604 according to the following function:

[0111] Image = Transform(WPS) - Images stream = Transform(WPS + AWPS)

[0112] where Image represents the modified digital image 604, Image stream represents the digital stream 610 provided to the client device 116 to present the digital image (e.g., as part of a digital video feed), and Transform(.) represents a GAN-based operation. In some cases, Transform(.) is performed by an embodiment of the image modification neural network 110, such as by combining features with features of an additional digital image and / or adjusting features to increase / decrease a measure of happiness or other GAN-based image features. The latent vector image editing system 102 thereby transforms the array of latent image vectors WPS modified by the image difference metric AWPS.

[0113] As mentioned, in one or more embodiments, the latent vector image editing system 102 provides the initial digital image 602 and the modified digital image 604 as part of a digital video feed or other digital data stream. For example, the latent vector image editing system 102 provides the digital video feed as part of a digital stream (e.g., digital stream 308 or 610). In certain embodiments, each latent image vector in the latent image vector array corresponds to a single frame of the digital video feed. Thus, the latent vector image editing system 102 provides the initial digital image 602 as a set of frames of the digital video feed, where each frame depicts the same initial digital image 602. To update the digital video feed with the modified digital image 604, as described above, the latent vector image editing system 102 modifies the set of frames of the digital video feed via the image modification neural network 110. Additionally, the latent vector image editing system 102 generates the image difference metric 608 and provides it to the client device 116 as part of the digital stream 610 to cause the client device 116 to update the digital video feed to present the modified digital image 604.

[0114] In some embodiments, the latent vector image editing system 102 performs steps for generating an image difference metric corresponding to a user interaction. Figure 5 Figure 6 The description of the acts 506 and 508 of FIG. 6 and Figure 6 The description of the comparison 606 of FIG. 6) provide various embodiments for performing steps for generating an image difference metric corresponding to a user interaction as well as supporting acts and algorithms.

[0115] For example, in some embodiments, the steps for generating an image difference metric corresponding to a user interaction include generating a modified latent image vector by combining, adding, subtracting, multiplying, or otherwise manipulating the initial latent image vectors described with respect to the acts 506. In some embodiments, the steps for generating an image difference metric corresponding to a user interaction include utilizing the image modification neural network to generate a modified digital image reflecting an image modification to the initial digital image based on changes within the latent image vectors corresponding to the user interaction. In these or other embodiments, the steps for generating an image difference metric corresponding to a user interaction include determining a difference between the initial digital image and the modified digital image as described with respect to the comparison 606 of FIG. 6. Figure 7

[0116] As mentioned above, in some embodiments, the latent vector image editing system 102 utilizes a distributed architecture with multiple devices located at different network sites. Specifically, in certain implementations, the latent vector image editing system 102 includes a latent vector stream plotter and an image modification neural network, where the latent vector stream plotter is also in communication with a client device. Figure 7 ​​FIGURE 1 illustrates an example architecture of a latent vector image editing system 102, according to one or more embodiments.

[0117] As Figure 7 illustrated, the latent vector image editing system 102 includes a latent vector stream painter 108 at a first network location and one or more image modification neural networks 110 (e.g., image modification neural network 110a and image modification neural network 110b) at another network location. In practice, as illustrated, the image modification neural network 110a and the image modification neural network 110b are stored on a server 104a (e.g., as part of the server(s) 104 or separate therefrom), while the latent vector stream painter 108 is stored on a server 104b (e.g., as part of the server(s) 104 or separate therefrom but separate from the server 104a). Additionally, the latent vector stream painter 108 is in communication with a client device 116 located at a third network location.

[0118] As Figure 7 further illustrated, the latent vector stream painter 108 includes logic for state management of a connection with the client device 116. Additionally, the latent vector stream painter 108 includes logic for user management, setting up and managing a digital stream connection (e.g., via Web Real-Time Communication or WebRTC), storing latent image vectors (e.g., within a database 112), receiving input from a user via a channel associated with the digital stream, and encoding frames or digital images as a digital video feed. With respect to the digital video feed, the latent vector stream painter 108 includes logic for establishing a WebRTC connection with the client device 116 and providing the digital stream including the digital video feed to the client device 116 via the connection.

[0119] As Figure 7 depicted, the latent vector stream painter 108 receives data from the client device 116 to control or trigger various video feed events. For example, the latent vector stream painter 108 receives an indication of a user interaction for modifying a digital image modification of a frame of a digital video. In practice, in some embodiments, the client device 116 accesses an application programming interface (“API”) associated with the latent vector image editing system 102 to request or access operations for digital image editing. In response, the latent vector stream painter 108 communicates with the image modification neural network 110b to perform any modifications to the digital image of the video feed corresponding to the user interaction. The latent vector stream painter 108 also provides a video feed event to the client device 116 to update the video feed for rendering and displaying the modified digital image resulting from the user interaction. As mentioned above, the latent vector stream painter 108 modifies the video feed in real-time with the user interaction.

[0120] AsFigures 8A-8B Further to illustrate, the latent vector image editing system 102 includes an image modification neural network 110a for extracting latent image vectors from digital images (sometimes referred to as projecting digital images into latent space). For example, the latent vector streamer 108 receives digital images from a video feed on a client device 116 and provides the digital images to the image modification neural network 110a to extract latent image vectors. As described above, the latent vector streamer 108 also modifies the latent image vectors based on user interactions and generates modified digital images from the modified latent image vectors with the image modification neural network.

[0121] Indeed, in some embodiments, the latent vector streamer 108 utilizes the image modification neural network 110a for projection and also utilizes an image modification neural network 110b to generate modified digital images from the modified latent image vectors (sometimes referred to as transformation). To elaborate, the image modification neural network 110b analyzes the modified latent vectors to transform the latent image vectors back into image space. In some cases, the image modification neural network 110b is the same (or at least the same type) of neural network as the image modification neural network 110a.

[0122] In certain embodiments, the latent vector image editing system 102 utilizes a batch drawing technique. More specifically, the latent vector image editing system 102 processes multiple latent image vectors for a given digital video feed in parallel. For example, the latent vector image editing system 102 utilizes the image modification neural network 110b to generate modified digital images for multiple digital images of a digital video feed (e.g., across multiple threads) simultaneously or concurrently using a GPU. For example, the latent vector image editing system 102 utilizes a GPU to implement the image modification neural network 110b to generate multiple modified digital images in parallel, each image corresponding to a different latent image vector within a vector array of the digital video feed. For example, the latent vector image editing system 102 generates modified digital images using the following function:

[0123] Images = G(WPS)

[0124] where Images represents digital images or frames of a digital video feed, WPS represents an array of latent image vectors of the digital video feed, and G represents a generative model or decoding neural network layer from the image modification neural network 110. Experimenters have demonstrated that the batch technique increases GPU utilization (e.g., a fraction or percentage of time that one or more GPU cores are running in the last second) from 5% to 95% compared to adjusting batch size alone.

[0125] As mentioned, in some embodiments, the latent vector image editing system 102 generates an image modification interface and provides it to the client device 116. Specifically, the latent vector image editing system 102 provides various interactive tools or elements for interacting with and modifying a digital image within the image modification interface. Figure 8A FIGURE 13 illustrates a client device 116 displaying an image modification interface including a grid of additional digital images for combining features with an initial digital image, in accordance with one or more embodiments.

[0126] As Figure 8B As illustrated, the client device 116 displays an image modification interface including a grid of digital images 804 and an initial digital image 802 overlaid on the images within the grid of digital images 804. As shown, the initial digital image 802 includes features from four digital images overlaid by the initial digital image 802 from the grid of digital images 804.

[0127] In some embodiments, the latent vector image editing system 102 arranges the additional images within the grid 804 according to one or more criteria. For example, the latent vector image editing system 102 arranges the additional digital images in the grid 804 according to the gradients of the latent image vectors corresponding to the additional digital images. For example, the latent vector image editing system 102 organizes the grid 804 by grouping images with similar gradients together. In some cases, the latent vector image editing system 102 utilizes a particular method for calculating the gradients, such as Uniform Manifold Approximation (“UMAP”) or t-distributed Stochastic Neighbor Embedding (“TSNE”).

[0128] In contrast, in one or more embodiments, the latent vector image editing system 102 arranges the additional digital images within the grid 804 according to spatial locality. To elaborate, the latent vector image editing system 102 arranges the additional digital images according to the similarity of the depicted content. For example, the latent vector image editing system 102 compares the latent image vectors associated with a plurality of stored digital images and generates the grid 804 by arranging the digital images in a line-by-line manner (e.g., row-by-row or column-by-column) based on the vector similarity. For example, by way of aligned face images, one scan line consisting of many face images is more likely to have similar content.

[0129] As a supplement or instead of using gradient or spatial locality, in some embodiments, the latent vector image editing system 102 arranges additional digital images within the grid 804 according to temporal locality. Since the latent vector space is regularized according to a perceptual path length (“PPL”), small changes in latent vectors cause small changes in pixels over time. In effect, the latent vector image editing system 102 determines PPL as a measure of distance or an amount of perceptual change (e.g., measured as a visual geometry group or “VGG” embedding distance). For example, the latent vector image editing system 102 determines a perceptual change in generated images for corresponding changes in latent image vectors. The latent vector image editing system 102 further arranges the grid 804 according to temporal locality by placing images with smaller perceptual changes closer together. By utilizing spatial locality and / or temporal locality to generate and arrange the grid 804, the latent vector image editing system 102 maintains a high compression ratio of bulk generated digital images for streaming to client devices 116 for video.

[0130] As mentioned, in one or more embodiments, the latent vector image editing system 102 receives an indication of a user interaction that moves the initial digital image 802 within the grid 804. Figures 9A-9B The user interaction is illustrated as sliding the initial digital image 802 to a new location within the grid 804. Based on the user interaction, the latent vector image editing system 102 generates a modified digital image 806 from a modified latent image vector that includes features of the initial digital image 802 as well as features of four digital images that overlap at the new location within the grid 804. For example, the latent vector image editing system 102 generates the modified digital image 806 to depict a combination (e.g., proportionally weighted combination according to respective overlapping areas) of image features from the initial digital image 802 as well as the underlying digital images at the new location. In effect, as a result of combining features of additional digital images, the modified digital image 806 appears different from the initial digital image 802.

[0131] In certain embodiments, the latent vector image editing system 102 updates the grid 804 based on user interactions that slide the digital image to different locations. For example, upon releasing a user input (e.g., releasing a mouse or removing a finger from a touchscreen) at a new location within the grid 804, the latent vector image editing system 102 re-arranges the digital images within the grid 804. In some cases, the latent vector image editing system 102 re-arranges the grid 804 based on changes to the digital images. For example, the latent vector image editing system 102 determines distances of latent image vectors to (vectors of) the initial digital image 802 to generate the initial grid 804. Additionally, the latent vector image editing system 102 updates the distances based on the modified digital image 806 and re-arranges the grid 804 to, for example, depict digital images with features more similar to the modified digital image 806 that are closer to the new location.

[0132] In certain embodiments, the latent vector image editing system 102 provides grid controls within the image modification interface. For example, the latent vector image editing system 102 provides selectable zoom controls to zoom in and out of the grid 804 to display more or fewer digital images. Accordingly, the latent vector image editing system 102 is able to present (or cause the client device 116 to present) different levels of detail to tailor the features of the digital images that are more or less similar to the initial digital image.

[0133] In one or more embodiments, the latent vector image editing system 102 provides an image modification interface that includes selectable slider elements for modifying particular image features (e.g., GAN-based image features) of the digital image. Specifically, the image modification interface includes a slider element for each of a plurality of image features. Figure 9A FIGURE 13 illustrates a client device 116 displaying an image modification interface that includes slider elements, in accordance with one or more embodiments.

[0134] As Figure 9B illustrated, the client device 116 presents an image modification interface that includes an initial digital image 902 and a modification pane 904. Within the modification pane 904, the image modification interface includes separate slider elements for modifying various image features, including an age measure, an anger measure, a surprise measure, a shake measure, a happiness measure, a baldness measure, and a glasses measure depicted in the initial digital image 902. For example, the position of the slider element 906 indicates an age measure corresponding to a -3 position along the slider. In some embodiments, the slider elements provide discrete changes in attributes, while in other embodiments, the changes along the slider are continuous.

[0135] As mentioned, the latent vector image editing system 102 receives an indication of a user interaction with the slider elements for modifying the initial digital image 902. Figures 10A-10BThe user interaction is illustrated as moving the age slider element 906 from a -3 position to a 10 position. In response to the user interaction, the latent vector image editing system 102 determines a vector direction corresponding to the image feature of age, and also determines a change measure corresponding to the vector direction. For example, the latent vector image editing system 102 determines an amount or degree to modify the latent image vector (e.g., the vector of the initial digital image 902) in the direction based on the change from -3 to 10. Accordingly, the latent vector image editing system 102 generates a modified latent image vector, generates a modified digital image 908, and provides the image differential measure to the client device 116 for rendering the modified digital image 908. As illustrated, the modified digital image 908 depicts a face that is older than the face of the initial digital image 902, while preserving other features of the image to maintain other elements of the appearance.

[0136] In a similar manner, for other slider elements, the latent vector image editing system 102 performs the same process. That is, the latent vector image editing system 102 determines a vector direction corresponding to the image feature adjusted by the user interaction, modifies the latent image vector, generates a modified digital image from the modified latent image vector, determines an image differential measure, and provides the image differential measure to the client device 116 for rendering the change.

[0137] In certain described embodiments, the latent vector image editing system 102 provides an image modification interface that includes a timeline element. In particular, the latent vector image editing system 102 provides a time element that is selectable to simultaneously modify multiple image features with a single user interaction. Figure 10A FIGURE 13 illustrates a client device 116 displaying an image modification interface that includes a timeline element, in accordance with one or more embodiments.

[0138] As Figure 10B illustrated, the client device 116 displays an initial digital image 1002 and a timeline element that includes multiple slider elements 1004 and a slidable bar 1006. As illustrated, the slidable bar 1006 within the timeline element indicates values along multiple vector directions for modifying the initial digital image 1002. Accordingly, in response to a user interaction adjusting the size and / or position of any of the slider elements 1004 relative to the slidable bar 1006, the latent vector image editing system 102 modifies the corresponding individual image feature (e.g., by modifying the initial latent image vector in the appropriate vector direction).

[0139] Additionally, in response to the user interaction moving the slidable bar 1006, the latent vector image editing system 102 modifies all of the image features of the plurality of slider elements 1004. In effect, by moving the slidable bar 1006, the value of each of the image features associated with the individual slider elements is adjusted. In response, the latent vector image editing system 102 generates a modified latent image vector and a modified digital image 1008. In effect, the latent vector image editing system 102 generates an image difference metric and provides it to the client device 116 to cause the client device 116 to render the modified digital image 1008. As shown, Figures 11A-11B the modified digital image 1008 includes the eyeglasses as well as some other image feature changes based on the user interaction moving the slidable bar 1006 to the right along the timeline.

[0140] In some embodiments, the latent vector image editing system 102 provides an image modification interface that includes a collage tool. In particular, the latent vector image editing system 102 provides a collage tool for selecting image features from one or more additional digital images to combine with the initial digital image. Figure 11A FIGURE 13 illustrates a client device 116 displaying an image modification interface that includes a collage tool, in accordance with one or more embodiments.

[0141] As Figure 11B illustrated, the client device 116 presents an image modification interface that includes an initial digital image and a collage tool that enables selection of various features of an additional digital image. As shown, the collage tool includes a plurality of selectable features corresponding to different portions of the additional digital image. For example, the collage tool includes a nose element 1104 that can be selected to combine a nose feature of the additional digital image within the initial digital image and a mouth element 1106 that can be selected to combine a mouth feature of the additional digital image with the initial digital image.

[0142] Based on the user interaction selecting the nose element 1104 and the mouth element 1106, the latent vector image editing system 102 generates a modified latent image vector by combining the corresponding latent features of the additional digital image (e.g., those features relative to the selected nose element 1104 and mouth element 1106). Additionally, the latent vector image editing system 102 generates a modified digital image 1108 and causes the client device 116 to render the modified digital image 1108 (e.g., by providing an image difference metric).

[0143] In effect, Figures 12A-12BThe client device 116 is illustrated presenting a modified digital image 1108 in which the nose and mouth are modified by combining the corresponding features with features of the additional digital image. To generate the modified digital image 1108, in some embodiments, the latent vector image editing system 102 modifies the latent image vector to reflect some constraints. For example, the latent vector image editing system 102 includes portions of the initial digital image 1102 and the additional digital image within the modified latent image vector to ensure that the resulting image has facial portions from the initial digital image 1102 and the selected portions of the additional digital image.

[0144] In one or more embodiments, the latent vector image editing system 102 provides an image modification interface that includes a sketch tool. In particular, the latent vector image editing system 102 provides an image modification interface that includes a sketch tool for drawing strokes on the initial digital image to add (or remove) various features. Figure 12A The client device 116 is illustrated displaying an image modification interface that includes a sketch tool in accordance with one or more embodiments.

[0145] As Figure 12B illustrated, the client device 116 displays an image modification interface that includes the initial digital image 1202. Additionally, the image modification interface includes a stroke 1204 (drawn via a digital brush) in the shape of glasses around the eyes of the face in the initial digital image 1202. The latent vector image editing system 102 receives an indication of the stroke 1204 and generates a modified digital image.

[0146] In particular, the latent vector image editing system 102 searches the digital image repository within the database 112 to identify one or more additional digital images that depict a face with glasses similar to the stroke 1204. The latent vector image editing system 102 also extracts or accesses latent image vectors corresponding to the identified digital images from the database 112.

[0147] Additionally, the latent vector image editing system 102 combines features of the identified digital image(s) (e.g., those corresponding to the glasses) with features of the initial digital image 1202. For example, the latent vector image editing system 102 generates a modified latent image vector by combining latent features of the initial latent image vector with latent features of the latent feature vector of the identified image with glasses corresponding to the stroke 1204. Thus, as Figure 13 illustrated, the latent vector image editing system 102 generates a modified digital image 1206 that depicts glasses on the face, otherwise similar to the face in the initial digital image 1202.

[0148] Now seeing Figure 13Additional details regarding the components and capabilities of the latent vector image editing system 102 will be provided. In particular, Figure 13 An example schematic of the latent vector image editing system 102 on an example computing device 1300 (e.g., a client device 116 and / or one or more of the server(s) 104) is illustrated. In some embodiments, the computing device 1300 refers to a distributed computing system in which different managers are located on different devices, as described above. For example, in certain embodiments, the computing system includes multiple computing devices (e.g., servers), one for the latent vector streamer 108, and a separate device for the image modification neural network 110. As Figure 13 illustrated, the latent vector image editing system 102 includes a digital stream manager 1302, a latent image vector manager 1304, a digital image modification manager 1306, an image difference manager 1308, and a storage manager 1310.

[0149] As just mentioned, the latent vector image editing system 102 includes the digital stream manager 1302. In particular, the digital stream manager 1302 manages, maintains, generates, provides, streams, transmits, processes, modifies, or updates digital streams provided to a client device. For example, the digital stream manager 1302 provides a digital stream for rendering a digital image as part of a digital video feed. Additionally, the digital stream manager 1302 updates the digital stream to provide an image difference metric to the client device for rendering a modification to the digital image and / or digital video feed.

[0150] Additionally, the latent vector image editing system 102 includes the latent image vector manager 1304. In particular, the latent image vector manager 1304 manages, maintains, extracts, encodes, determines, generates, modifies, updates, provides, receives, transmits, processes, analyzes, or identifies latent image vectors. For example, the latent image vector manager 1304 extracts a latent image vector from a digital image. Additionally, the latent image vector manager 1304 modifies a latent image vector based on user interaction for editing the digital image (e.g., via GAN-based operations).

[0151] As Figure 13 illustrated, the latent vector image editing system 102 also includes the digital image modification manager 1306. In particular, the digital image modification manager 1306 manages, maintains, generates, modifies, updates, provides, receives, or identifies modified digital images. For example, the digital image modification manager 1306 generates a modified digital image from a modified latent image vector. Additionally, the digital image modification manager 1306 provides the modified digital image to the image difference manager 1308 to generate an image difference metric.

[0152] In practice, as shown, the latent vector image editing system 102 also includes an image difference manager 1308. Specifically, the image difference manager 1308 manages, maintains, determines, generates, updates, provides, transmits, or identifies image difference metrics. For example, the image difference manager 1308 compares the initial digital image to the modified digital image to determine an image difference metric that reflects a difference between the initial digital image and the modified digital image. Additionally, the image difference manager 1308 provides the image difference metric to the digital stream manager 1302, which in turn provides the image difference metric to the client device (e.g., to cause the client device to render the modified digital image).

[0153] The latent vector image editing system 102 also includes a storage manager 1310. The storage manager 1310 operates in conjunction with or includes one or more memory devices, such as a database 1312 (e.g., the database 112) that stores various data, such as a repository of digital images and a repository of latent image vectors. The storage manager 1310 (e.g., via the non-transitory computer memory / memory device(s)) stores and maintains data (e.g., within the database 1312) associated with modifying latent image vectors, generating modified digital images, and determining image difference metrics.

[0154] In one or more embodiments, each of the components of the latent vector image editing system 102 communicate with each other using any suitable communication techniques. Additionally, the components of the latent vector image editing system 102 communicate with one or more other devices, including the one or more client devices described above. It will be recognized that, although the components of the latent vector image editing system 102 are shown in Figure 13 separation, any sub-component can be combined into a fewer component, such as into a single component, or divided into a greater number of components, as can serve a particular implementation. Further, although the components of the latent vector image editing system 102 are described with respect to the latent vector image editing system 102, at least some of the components for performing operations in connection with the latent vector image editing system 102 described herein can be implemented on other devices within the environment. Figures 1-13

[0155] ​The components of the latent vector image editing system 102 can include software, hardware, or both. For example, the components of the latent vector image editing system 102 can include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices, such as the computing device 1300. When executed by the one or more processors, the computer-executable instructions of the latent vector image editing system 102 can cause the computing device 1300 to perform the methods described herein. Alternatively, the components of the latent vector image editing system 102 can include hardware, such as a specialized processing device that performs a particular function or group of functions. Additionally or alternatively, the components of the latent vector image editing system 102 can include a combination of computer-executable instructions and hardware.

[0156] Further, the components of the latent vector image editing system 102 that perform the functions described herein can be, for example, implemented as part of a standalone application, a module of an application, a plug-in of an application including a content management application, a library function or function that can be called by other applications, and / or a cloud computing model. Thus, the components of the latent vector image editing system 102 can be implemented as part of a standalone application on a personal computing device or mobile device. Alternatively or additionally, the components of the latent vector image editing system 102 can be implemented in any application that allows for the creation and delivery of marketing content to users, including but not limited to applications in EXPERIENCE MANAGER and CREATIVE Cloud, such as Stock, and “ADOBE,” “ADOBE EXPERIENCE MANAGER,” “CREATIVE CLOUD,” “ADOBE STOCK,” “PHOTOSHOP,” “LIGHTROOM,” and “INDESIGN” are registered trademarks or trademarks of Adobe Inc. in the United States and / or other countries.

[0157] Figure 14 The systems, methods, and non-transitory computer-readable media described in this document, among other things, provide for generating and providing an image differential measure by comparing digital images associated with latent image vectors. In addition to the foregoing, an embodiment can be described in the form of a flow diagram including acts for accomplishing a particular result. E.g., Figure 14 FIG. illustrates a flow diagram of an example sequence or series of acts in accordance with one or more embodiments.

[0158] Although Figure 14 FIG. illustrates acts in accordance with one embodiment, alternative embodiments can omit, add to, reorder, and / or modify Figure 14any of the actions shown. Figure 14 The actions of can be performed as part of a method. Alternatively, a non-transitory computer-readable medium can include instructions that, when executed by one or more processors, cause a computing device to perform the actions of Figure 14 In other embodiments, a system can perform the actions of Figure 14 Additionally, the actions described herein can be repeated or performed in parallel with each other or with different instances of the same or other similar actions.

[0159] Figure 15 A series of example actions 1400 is illustrated for generating and providing an image differential measure by comparing digital images associated with latent image vectors. Specifically, the series of actions 1400 includes an action 1402 of extracting a latent image vector from an initial digital image. For example, the action 1402 involves extracting a latent image vector from an initial digital image displayed via a client device. In some cases, the action 1402 involves extracting an initial latent image vector from an initial digital image displayed via a mobile device that is a client device.

[0160] Additionally, the series of actions 1400 includes an action 1404 of receiving an indication of a user interaction to modify the initial digital image. Specifically, the action 1404 involves receiving an indication of a user interaction to modify the initial digital image. For example, the action 1404 involves receiving an indication of a user interaction to modify the initial digital image by receiving an indication of a user interaction to select one or more additional digital images to combine with the digital image. Additionally, the series of actions 1400 includes an action of generating a modified latent image vector by combining the latent image vector with one or more additional latent image vectors corresponding to the one or more additional digital images. Further, the series of actions 1400 includes an action of generating, with an image modification neural network, a modified digital image depicting a combination of image features from the initial digital image and the one or more additional digital images based on the modified latent image vector.

[0161] In some embodiments, the action 1404 involves receiving an indication of a user interaction to modify an image feature of the initial digital image. Additionally, the series of actions 1400 includes an action of modifying the latent image vector to represent the modified image feature in the modified latent image vector. Further, the series of actions 1400 includes an action of generating the modified digital image based on the modified latent image vector with a generative adversarial neural network (GAN) as the image modification neural network. In some cases, the series of actions 1400 includes an action of providing the image differential measure to the mobile device in response to the user interaction (in real-time).

[0162] In at least one embodiment, series of actions 1400 includes the following actions: providing, for display on a client device, an image modification interface including a sketch tool for sketching strokes on an initial digital image. Additionally, series of actions 1400 includes the following actions: based on receiving an indication of a stroke sketched on the initial digital image, providing, to the client device, an image differential metric for use in rendering a modified digital image including an overlay of additional image features corresponding to the stroke. In some cases, action 1404 involves: receiving an indication of a movement of the initial digital image to overlay an additional digital image within a digital image grid.

[0163] As illustrated, series of actions 1400 includes action 1406: determining an image differential metric. Specifically, action 1406 involves: in response to an indication of a user interaction, determining, based on a change within a latent image vector corresponding to the user interaction, an image differential metric reflecting an image modification of an initial digital image performed by an image modification neural network. For example, action 1406 involves: in response to an indication of a user interaction to modify a digital image, modifying, with a first computing device, a latent image vector; generating, with an image modification neural network on a second computing device, a modified digital image from the modified latent image vector; and determining, with the first computing device, a difference between the initial digital image and the modified digital image. Modifying the latent image vector can involve: determining a vector direction and a change measure corresponding to the vector direction associated with the user interaction; and modifying the latent image vector in the vector direction according to the change measure.

[0164] In some embodiments, action 1406 involves: generating, with the image modification neural network, a modified digital image reflecting an image modification of the initial digital image based on a change within the latent image vector corresponding to the user interaction; and determining a difference between the initial digital image and the modified digital image. Generating the modified digital image can involve: processing, with the image modification neural network on a server device at a first location, a modified latent image vector corresponding to the change within the latent image vector. In some cases, determining the image differential metric involves: determining, via a computing device at a second location, the difference between the initial digital image and the modified digital image.

[0165] The series of acts 1400 further includes act 1408: providing the image differential metric for rendering the modified digital image. In particular, act 1408 involves providing the image differential metric to the client device for rendering the modified digital image that depicts the image modification. In at least one embodiment, the series of acts 1400 includes acts of modifying the latent image vector in a direction of the vector corresponding to the user interaction. Additionally, act 1408 can involve generating the modified digital image from the modified latent image vector by performing a GAN-based operation with the image modification neural network. In some cases, act 1408 involves providing the image differential metric for rendering the modified digital image within a browser on the client device at the third location. For example, act 1408 involves providing the image differential metric to the client device for rendering the modified digital image that depicts a combination of image features from the initial digital image and the additional digital image within the grid of digital images.

[0166] In certain embodiments, the series of acts 1400 includes acts of receiving an indication of a first quantity of first additional digital images under the digital image within the grid of digital images; receiving an indication of a second quantity of second additional digital images under the digital image within the grid of digital images; and generating the modified latent image vector by combining the initial latent image vector with the first additional initial latent image vector and the second additional latent image vector proportionally according to the first quantity and the second quantity, respectively.

[0167] In some cases, the series of acts 1400 includes acts of providing the image modification interface for display on the client device, the image modification interface including a depiction of the digital image and a slider element that is selectable to adjust an image feature of the initial digital image. Further, the series of acts 1400 includes acts of providing the image differential metric to the client device for rendering the modified digital image that includes the adjusted image feature based on receiving an indication of the slider element that adjusts the image feature of the initial digital image. In some embodiments, the series of acts 1400 includes acts of providing the image modification interface for display on the client device, the image modification interface including a depiction of the initial digital image, a slider element that is selectable to adjust an image feature of the initial digital image, and a slidable bar that is selectable to simultaneously adjust multiple image features of the initial digital image.

[0168] In one or more embodiments, the series of actions 1400 includes the following actions: providing, for display on the client device, an image modification interface including a collage tool for selecting an image feature from an additional digital image to combine with the initial digital image. Further, the series of actions 1400 can include the following actions: based on receiving an indication of the selected image feature from the additional digital image, providing the image differential metric to the client device for rendering a modified digital image depicting the combination of the image feature from the initial digital image and the selected image feature from the additional digital image.

[0169] In some cases, the series of actions 1400 includes the following actions: receiving, from a computing device including an image modification neural network, a latent image vector for a digital image displayed via the client device; and providing the digital image to the computing device to extract the latent image vector from the digital image with a generative adversarial neural network (GAN). In these or other cases, the series of actions 1400 includes the following actions: based on the latent image vector, providing the initial digital stream to the client device to cause the client device to display the initial digital image; and providing the image differential metric as part of the modified digital stream to the client device to cause the client device to display the modified digital image in place of the initial digital image.

[0170] In some embodiments, the series of actions 1400 includes the following actions: based on the latent image vector, providing the initial digital stream to the client device to cause the client device to display the initial digital image. In these or other embodiments, the series of actions 1400 includes the following actions: providing the image differential metric as part of the modified digital stream to the client device to cause the client device to display the modified digital image in place of the initial digital image.

[0171] In one or more embodiments, the series of actions 1400 includes the following actions: receiving an additional indication of an additional user interaction to modify the initial digital image. Additionally, the series of actions 1400 includes the following actions: in response to the additional indication of the additional user interaction, generating an additional image differential metric indicating a difference between the modified digital image and an additionally modified digital image. Further, the series of actions 1400 can include the following actions: providing the additional image differential metric as part of the digital stream to the client device for rendering the additionally modified digital image in place of the modified digital image.

[0172] Embodiments of the present disclosure can include or utilize special-purpose or general-purpose computers, including computer hardware such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer- executable instructions and / or data structures. In particular, one or more of the processes described herein can be implemented at least in part as instructions embodied in a non-transitory computer- readable medium and executable by one or more computing devices, such as any of the media content access devices described herein. Generally, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium, (e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.

[0173] A computer-readable medium can be any available medium or means that can be accessed by a general purpose or special purpose computer. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Accordingly, embodiments of the present disclosure can include at least two distinct kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.

[0174] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid state drives ("SSDs") (e.g., based on RAM), Flash memory, phase- change memory ("PCM"), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.

[0175] A "network" is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.

[0176] Further, upon reaching various computer system components, program code in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received by way of network or data link can be buffered in RAM within a network interface module (e.g., a "NIC"), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.

[0177] For example, computer-executable instructions include instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general purpose computer to transform the general purpose computer into a special purpose computer that implements elements of the present disclosure. For example, computer-executable instructions can be binary, intermediate format instructions such as produced by a compiler, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0178] Those skilled in the art will appreciate that the disclosure can be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure can also be practiced in distributed system environments where local and remote computer systems that are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules can be located in both local and remote memory storage devices.

[0179] Embodiments of the present disclosure can also be practiced in a cloud computing environment. In this description and the following claims, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer on-demand access to shared pools of configurable computing resources. In this context, a "cloud" refers to a set of computing resources available via the Internet, which can be used to perform computing tasks. In some examples, the cloud can comprise a plurality of computing resources, which can be provided by a service provider and accessed via a network, such as the Internet.

[0180] A cloud computing model can consist of various features, such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, etc. A cloud computing model can also expose various service models, such as, for example, Software as a Service (“SaaS”), Platform as a Service (“PaaS”), and Infrastructure as a Service (“IaaS”). A cloud computing model can also be deployed using different deployment models, such as private cloud, community cloud, public cloud, hybrid cloud, etc. In this description and the claims, a “cloud computing environment” is an environment in which cloud computing is employed.

[0181] Figure 15 An example computing device 1500 (e.g., computing device 1300, client device 116, and / or server(s) 104) is illustrated in block diagram form, which can be configured to perform one or more of the processes described above. It is to be appreciated that potential vector image editing system 102 can include an implementation of computing device 1500. As shown by Figure 15 The computing device can include a processor 1502, a memory 1504, a storage device 1506, an I / O interface 1508, and a communication interface 1510, as shown. In addition, the computing device 1500 can include input devices, such as a touchscreen, mouse, keyboard, etc. In certain embodiments, the computing device 1500 can include fewer or more components than those shown. Figure 15 The computing device 1500 can include fewer or more components than those shown. ​ The components of the computing device 1500 shown will now be described in additional detail.

[0182] In particular embodiments, the processor(s) 1502 include hardware, such as a collection of logic gates, to execute instructions. As an example and not by way of limitation, to execute instructions, the processor(s) 1502 can retrieve (or fetch) the instructions from an internal register, an internal cache, memory 1504, or storage 1506 and decode and execute them.

[0183] The computing device 1500 includes memory 1504 coupled to the processor(s) 1504. The memory 1504 can be used for storing data, metadata, and programs for execution by the processor(s). The memory 1504 can include one or more of volatile and non-volatile memories, such as read-only memory (“ROM”), random access memory (“RAM”), solid state drives (“SSDs”), flash memory, phase change memory (“PCM”), or other types of data storage. The memory 1504 can be internal or distributed.

[0184] Computing device 1500 includes a storage device 1506, comprising storage for storing data or instructions. As an example and not by way of limitation, storage device 1506 can comprise a non-transitory storage medium described above. Storage device 1506 can include a hard disk drive (HDD), flash memory, universal serial bus (USB) drive or combination or such storage devices.

[0185] Computing device 1500 also includes one or more input or output ("I / O") devices / interfaces 1508 provided to allow a user to provide input to, receive output from, and otherwise transfer data to and from computing device 1500. These I / O devices / interfaces 1308 can include a mouse, a keypad or keyboard, a touch screen, a camera, an optical scanner, a network interface, a modem, other I / O devices, or a combination of such I / O devices / interfaces 1508. The touch screen can utilize a stylus or a finger for writing.

[0186] I / O devices / interfaces 1508 can include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers and one or more audio drivers. In certain embodiments, device / interface 1508 is configured to provide graphical data to a display for presentation to a user. The graphical data can be representative of one or more graphical user interfaces and / or any other graphical content as can serve a particular implementation.

[0187] Computing device 1500 can also include a communications interface 1510. The communications interface 1510 can include hardware, software, or both. The communications interface 1510 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device and one or more other computing devices 1500 or one or more networks. As an example and not by way of limitation, communications interface 1510 can include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. The computing device 1500 can also include a bus 1512. The bus 1512 can include hardware, software, or both implementing a bus standard including, e.g., PCI, PCI Express (PCIe), HyperTransport, InfiniBand, NuBus, Amazon Web Services (AWS) Direct Connect, a virtual network, a virtual bus, or the like.

[0188] In the foregoing specification, the application has been described with reference to specific examples of embodiments thereof. Various embodiments and aspects of the application are described with reference to details discussed herein, and the accompanying drawings illustrate various embodiments. The above description and drawings illustrate the application and are not to be construed as limiting the application. Numerous specific details are described to provide a thorough understanding of various embodiments of the application.

[0189] The application can be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein can be performed with fewer or additional steps / actions, or the steps / actions can be performed in a different order. Additionally, the steps / actions described herein can be repeated or performed in parallel with each other or with different instances of the same or similar steps / actions. Accordingly, the scope of the application is indicated by the appended claims, rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1. A non-transitory computer-readable medium for digital image editing comprising instructions that, when executed by a processing device, cause the processing device to perform operations comprising: extracting a latent image vector from an initial digital image displayed via a client device; receiving an indication of a user interaction to modify the initial digital image; determining, in response to the indication of the user interaction, an image difference metric based on a change within the latent image vector corresponding to the user interaction, the image difference metric reflecting a difference between the initial digital image and a modified digital image generated by an image modification neural network; and providing the image difference metric to the client device for rendering the modified digital image.

2. The non-transitory computer-readable medium of claim 1, further comprising instructions that, when executed by the processing device, cause the processing device to perform operations comprising: providing, based on the latent image vector, an initial digital stream to the client device to cause the client device to display the initial digital image; and providing the image difference metric to the client device as part of a modified digital stream to cause the client device to display the modified digital image in place of the initial digital image.

3. The non-transitory computer-readable medium of claim 1, wherein determining the image difference metric comprises: modifying, in response to the indication of the user interaction to modify the initial digital image, the latent image vector with a first computing device; generating, with the image modification neural network on a second computing device, the modified digital image from the modified latent image vector; and determining, with the first computing device, the difference between the initial digital image and the modified digital image.

4. The non-transitory computer-readable medium of claim 3, wherein modifying the latent image vector comprises: determining a vector direction and a measure of change corresponding to the vector direction associated with the user interaction; and modifying the latent image vector in the vector direction according to the measure of change.

5. The non-transitory computer-readable medium of claim 1, further comprising instructions that, when executed by the processing device, cause the processing device to perform operations comprising: modifying the latent image vector in a vector direction corresponding to the user interaction; and generating the modified digital image from the modified latent image vector by performing a GAN-based operation with the image modification neural network.

6. The non-transitory computer-readable medium of claim 1, further comprising instructions that, when executed by the processing device, cause the processing device to perform operations comprising: receiving the indication of the user interaction to modify the initial digital image by receiving an indication of a user interaction to select one or more additional digital images to combine with the initial digital image. ​ ​ ​ ​ ​ generating a modified latent image vector by combining the latent image vector with one or more additional latent image vectors corresponding to the one or more additional digital images; and generating, with the image modification neural network, the modified digital image depicting a combination of image features from the initial digital image and the one or more additional digital images based on the modified latent image vector.

7. The non-transitory computer-readable medium of claim 6, further comprising instructions that, when executed by the processing device, cause the processing device to perform operations comprising: receiving an indication of a user interaction for modifying an image feature of the initial digital image by receiving the indication of the user interaction for modifying the image feature of the initial digital image; modifying the latent image vector to represent the modified image feature in a modified latent feature vector; and generating the modified digital image based on the modified latent feature vector with a generative adversarial neural network (GAN) as the image modification neural network.

8. The non-transitory computer-readable medium of claim 1, further comprising instructions that, when executed by the processing device, cause the processing device to perform operations comprising: providing, for display on the client device, an image modification interface including a sketch tool for drawing a stroke on the initial digital image; and based on receiving an indication of a stroke drawn on the initial digital image, providing the image difference metric to the client device for rendering the modified digital image including an overlay of an additional image feature corresponding to the stroke.

9. A system for digital image editing, comprising: one or more memory devices including an image modification neural network; and one or more processors configured to cause the system to: extract an initial latent image vector from an initial digital image displayed via a client device; receive an indication of a user interaction for modifying the initial digital image; in response to the indication of the user interaction, generate an image difference metric reflecting a difference between the initial digital image and a modified digital image by: generating, with the image modification neural network, the modified digital image reflecting an image modification to the initial digital image based on a change within the initial latent image vector corresponding to the user interaction; and determining the difference between the initial digital image and the modified digital image; and provide the image difference metric to the client device for rendering the modified digital image.

10. The system of claim 9, wherein the one or more processors are further configured to cause the system to: receive an additional indication of an additional user interaction for modifying the initial digital image; in response to the additional indication of the additional user interaction, generate an additional image difference metric indicating a difference between the modified digital image and a further modified digital image; and ​ ​ providing the image differential measure to the client device for rendering the further modified digital image in place of the modified digital image.

11. The system of claim 9, wherein the one or more processors are further configured to cause the system to: extract the initial latent image vector from the initial digital image displayed via the mobile device as the client device; and provide the image differential measure to the mobile device in response to the user interaction.

12. The system of claim 9, wherein the one or more processors are further configured to cause the system to: generate the modified digital image by processing a modified latent image vector corresponding to the changes within the initial latent image vector with the image modification neural network on a server device at a first location; generate the image differential measure by determining the difference between the initial digital image and the modified digital image via a computing device at a second location; and provide the image differential measure for rendering the modified digital image within a browser on the client device at a third location.

13. The system of claim 9, wherein the one or more processors are further configured to cause the system to: provide an image modification interface comprising a grid of digital images for display on the client device; receive the indication of the user interaction by receiving an indication of moving the initial digital image to overlay an additional digital image within the grid of digital images; and provide the image differential measure to the client device for rendering the modified digital image depicting a combination of image features from the initial digital image and the additional digital image within the grid of digital images.

14. The system of claim 13, wherein the one or more processors are further configured to cause the system to: receive an indication of a first quantity of a first additional digital image under the initial digital image within the grid of digital images; receive an indication of a second quantity of a second additional digital image under the initial digital image within the grid of digital images; and generate a modified latent image vector by proportionally combining the initial latent image vector with a first additional initial latent image vector and a second additional latent image vector, respectively, in accordance with the first quantity and the second quantity.

15. The system of claim 9, wherein the one or more processors are further configured to cause the system to: provide an image modification interface for display on the client device, the image modification interface comprising a depiction of the digital image and a slider element selectable to adjust an image feature of the initial digital image; and provide the image differential measure to the client device for rendering the modified digital image comprising the adjusted image feature based on receiving an indication of adjusting the slider element of the initial digital image.

16. The system of claim 9, wherein the one or more processors are further configured to cause the system to: provide, for display on the client device, an image modification interface that includes a depiction of the initial digital image, slider elements that are selectable to adjust image features of the initial digital image, and a slidable bar that is selectable to simultaneously adjust multiple image features of the initial digital image.

17. The system of claim 9, wherein the one or more processors are further configured to cause the system to: provide, for display on the client device, an image modification interface that includes a collage tool for selecting image features from additional digital images to combine with the initial digital image; and based on receiving an indication of image features selected from the additional digital images, provide the image difference metric to the client device for rendering the modified digital image that depicts a combination of image features from the initial digital image and the image features selected from the additional digital images.

18. A computer-implemented method for digital image editing that utilizes a latent vector approach for implementing neural network modifications, the computer-implemented method comprising: extracting a latent image vector from an initial digital image that is displayed via a client device; receiving an indication of a user interaction for modifying the initial digital image; in response to the indication of the user interaction, determining, based on a change within the latent image vector corresponding to the user interaction, an image difference metric that reflects a difference between the initial digital image and a modified digital image that is generated by an image modification neural network; and providing the image difference metric to the client device for rendering the modified digital image.

19. The computer-implemented method of claim 18, wherein extracting the latent image vector comprises: providing the initial digital image to a computing device for extracting the latent image vector from the initial digital image utilizing a generative adversarial neural network (GAN).

20. The computer-implemented method of claim 18, further comprising: based on the latent image vector, providing an initial digital stream to the client device to cause the client device to display the initial digital image; and providing the image difference metric as part of a modified digital stream to the client device to cause the client device to display the modified digital image in place of the initial digital image.

Citation Information

Patent Citations

  • Face image editing method and device and storage medium

    CN111260754A

  • Face attribute editing method based on generative adversarial network

    CN112330759A

  • Techniques for Modifying a Query Image

    US20200356591A1