Automated digital tool identification via rasterized images

Through the tool area detection and parameter estimation network of the visual lens system, tools and parameters in raster image data are automatically identified, which solves the problem of lack of explanation of the creative process of digital art collections and improves efficiency and computing resource utilization.

CN114943676BActive Publication Date: 2025-09-23ADOBE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111355933.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-02-08
Filing Date
2021-11-16
Publication Date
2025-09-23
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

Digital art portfolios lack explanations of the creative process, making it difficult for artists to learn how to achieve the same artistic effect using different tools, leading to inefficient consumption of computing resources and difficulty in understanding.

Method used

Through the visual lens system, the tool region detection network and the tool parameter estimation network are used to automatically identify the image region in the raster image data, and the corresponding digital tools and parameter configurations are provided to generate interactive images to indicate the implementation method of the visual appearance.

Benefits of technology

It simplifies the digital art creation process, improves the efficiency of artists in learning and replicating effects, reduces the waste of computing resources, and provides guidance on tool use and parameter configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114943676B_ABST
    Figure CN114943676B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to automated digital tool identification through rasterized images. A visual lens system is described that automatically and without user intervention identifies digital tool parameters to achieve a visual appearance of an image region in raster image data. To do so, the visual lens system processes the raster image data using a tool region detection network trained to output a mask indicating whether a digital tool can be used to achieve the visual appearance for each pixel in the raster image data. The mask is then processed by a tool parameter estimation network trained to generate a probability distribution indicating estimates of discrete parameter configurations that are suitable for the digital tool to achieve the visual appearance. The visual lens system generates an image tool description for the parameter configuration and incorporates the image tool description into an interactive image of the raster image data. The image tool description enables the transfer of the digital tool parameter configuration to different image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of digital images. Background Art

[0002] Digital artists often browse digital artwork portfolios created by other artists to gain inspiration for their own creations and learn new skills. While many sources of digital artwork portfolios are available for browsing, these portfolios only showcase the artwork as a final product and do not include information indicating how the artwork was created or which steps were taken to achieve the final product, thereby hindering artists' ability to learn how to incorporate similar techniques when creating their own artwork. The proliferation of available tools for creating digital artwork also exacerbates the difficulty of understanding how the final artwork was achieved, as different tools can be used to achieve similar effects.

[0003] As a result, the interpretation of how the final artwork was achieved is subject to underlying human biases and is often ambiguous, with different digital artists perceiving different tools used to create the same artwork. Consequently, because conventional systems lack the ability to analyze digital artwork and connect its various components back to the tools used to create it, artists are forced to experiment with different tools in an effort to replicate the desired result, which inefficiently consumes computing resources through the need for repeated iterations. Summary of the Invention

[0004] A visual lens system is described that automatically and without user intervention identifies digital tool parameters that can be used to achieve a visual appearance for an image region in raster image data. To do so, the visual lens system processes the raster image data using a tool region detection network trained to output a mask for each of a plurality of different digital tools, the mask indicating whether the corresponding digital tool can be used to achieve the visual appearance for each pixel in the raster image data. The mask indicating the available digital tools is then processed by a tool parameter estimation network trained to generate a probability distribution indicating estimates of discrete parameter configurations that are suitable for the digital tools to achieve the visual appearance.

[0005] The visual lens system generates an image tool description indicating a digital tool parameter configuration that can be used to achieve the visual appearance of the image area, and incorporates the image tool description into an interactive image of the raster image data. In response to selection of the image area, the interactive image causes the image tool description to be displayed, thereby indicating how a similar visual appearance can be achieved in other image data. The visual lens system is further configured to provide information related to the digital tool and to support application of the digital tool and its parameter configuration identified in the interactive image to the vector image data.

[0006] This Summary introduces a selection of concepts in a simplified form that are further described below in the Detailed Description. Therefore, this Summary is not intended to identify essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0008] The detailed description is described with reference to the accompanying drawings. In some implementations, entities shown in the drawings refer to one or more entities, and thus the discussion interchangeably refers to entities in the singular or plural.

[0009] Figure 1 is an illustration of an environment in an example implementation operable to employ the automated digital tool identification by rasterizing images described herein.

[0010] Figure 2 Depicts a system in an example implementation, showing in more detail Figure 1 Operation of the visual lens system.

[0011] Figure 3 Depicts an example implementation, showing in more detail the Figure 1 The visual lens system realizes the operation of the tool area detection network.

[0012] Figure 4 Describes the process of generating interactive images from rasterized image data by Figure 1 Example visualization of data utilized by the Visual Lens system.

[0013] Figure 5 Depicts an example implementation, showing in more detail the Figure 1 The visual lens system implements the operation of the tool parameter estimation network.

[0014] Figure 6 Depicts an example user interface, including Figure 1 A visual lens system generates interactive images from rasterized image data.

[0015] Figure 7 Depicts an example user interface, including Figure 1 A visual lens system generates interactive images from rasterized image data.

[0016] Figure 8A and Figure 8B Describes the generation of Figure 1A system in which a tool region detection network and a tool parameter estimation network are implemented by a visual lens system.

[0017] Figure 9 It depicts the training by Figure 1 A flowchart of a process in an example implementation of a visual lens system using a tool region detection network to generate interactive images from rasterized image data.

[0018] Figure 10 It depicts the training by Figure 1 A flowchart of a process in an example implementation of a visual lens system using a tool parameter estimation network to generate interactive images from rasterized image data.

[0019] Figure 11 is a flowchart depicting a procedure in an example implementation in which an interactive image is generated from rasterized image data and output to display information describing digital tool parameters that can be used to achieve the visual appearance of the rasterized image data.

[0020] Figure 12 An example system is illustrated, including various components of an example device to implement reference Figures 1 to 11 Describe the technology. DETAILED DESCRIPTION

[0021] Overview

[0022] Computing devices that implement graphics creation and editing software provide an ever-expanding range of digital tools that can be used to define the visual appearance of portions of graphical image data. Digital graphic designers responsible for creating digital image data often browse through vast collections of digital artwork from various sources to gather inspiration for their own projects and to hone their image creation and manipulation skills. However, digital artwork is typically presented in a rasterized form, with only information describing the color of individual pixels and providing no indication of the specific digital tools used to create the digital artwork.

[0023] Given the ever-expanding range of digital tools used to generate this visual data, even the most experienced digital artists often do not know how to achieve the same visual results and incorporate this inspiration into their own image data. Consequently, digital graphic designers and artists are responsible for manually experimenting with various digital tools, which results in inefficient use of computing resources from repetitive processing in an attempt to achieve the same visual result.

[0024] To address these issues, a technique for automating the identification of digital tools by a computing device that rasterizes image data is described. In one example, a visual lens system generates an interactive image from the raster image data. The interactive image identifies discrete image regions within the raster image data and provides an indication of a specific digital tool and a parameter configuration of the specific digital tool that can be used to achieve the visual appearance of the corresponding image region. In an implementation, the visual lens system generates the interactive image to include descriptive information about the digital tool and automatically applies the identified parameter configuration of the digital tool to different image data.

[0025] The visual lens system 104 is configured to automatically and independently of user intervention generate interactive images from raster image data. To do so, the visual lens system implements a tool region detection network that is trained to process the rasterized image data as input and output a binary mask for each of a plurality of different image processing digital tools indicating whether the corresponding digital tool can be used to achieve the visual appearance of a corresponding pixel in the rasterized image data.

[0026] A tool mask, indicating that a digital tool may be used to achieve a visual appearance of an image region in the rasterized image data, is then used to extract a corresponding image region from the rasterized image data, concatenated with the mask, and provided as input to a tool parameter estimation network. The tool parameter estimation network is trained to generate a probability distribution indicating a confidence level for each parameter configuration in a discrete set of possible parameter configurations of controllable parameters of the digital tool for achieving the visual appearance of the image region.

[0027] The visual lens system generates an image tool description based on the output of the tool region detection network and the tool parameter estimation network. The image tool description provides descriptive information about the digital tool and its specific parameter configuration(s), as well as controls that support storing the tool parameters for subsequent use or directly applying the digital tool parameter configuration to different image data. The functionality of the visual lens system is further described in the context of training the tool region detection network and the tool parameter estimation network to enable the automatic generation of interactive images, including image tool descriptions, from raster image data. Further discussion of these and other examples is included in the following sections and is illustrated in the corresponding figures.

[0028] Terminology Examples

[0029] As used herein, the term "image data" refers to digital content configured for display by a computing device. Image data includes raster or rasterized images and vector images.

[0030] As used herein, "raster / rasterized image" refers to an image composed of a finite set of digital values ​​(e.g., pixels), each including a digital representation of a finite number of discrete quantities that describe the visual appearance of the digital values ​​(e.g., the color, intensity, grayscale value, etc. of the pixel). Alternatively, a raster image is referred to as a bitmap image, which includes a dot matrix structure that describes a grid of pixels, with each pixel in the grid being specified by a number of bits, and the size of the rasterized image being defined by the width and height of the pixel grid.

[0031] In contrast to raster images, the term "vector image" as used herein refers to an image composed of paths and curves specified by mathematical formulas (e.g., points connected by polylines, Bezier curves, etc.). Because the visual appearance of a vector image is defined by mathematical formulas, it can be scaled to accommodate different display resolutions and is not limited to a finite set of digital values ​​representing the smallest individual element in a raster image.

[0032] As used herein, the term "digital tool" refers to a computer-implemented mechanism for altering the visual appearance of image data, such as functionality supported by vector graphics editors, raster graphics editors, and other digital graphics editing software. Examples of digital tools include gradients (e.g., linear, radial, freeform, etc.), drop shadows, glows (inner and outer), noise, blur, grain, halftones, brushes, etc. Digital tools are configured for application to discrete portions of image data as well as to image data as a whole.

[0033] As used herein, the term "tool parameters" refers to configurable settings of a digital tool that define the final visual effect of the digital tool on the image data to which the digital tool is applied. Thus, each digital tool is controlled by one or more tool parameters. For example, the example Linear Gradient tool includes three tool parameters: First Color, Second Color, and Direction, which together specify the gradient color gradient and the direction in which the gradient color gradient progresses from the first color to the second color in the area of ​​image data to which the Linear Gradient tool is applied.

[0034] As used herein, the term "parameter configuration" refers to a specific value or set of values ​​within a limited range of values ​​or sets of values ​​for a corresponding tool parameter. For example, the control range for the first color value and the second color value of the example linear gradient tool parameters described above is a discrete integer value from zero to 256, such that the corresponding parameter configuration for the first color and the second color refers to a corresponding integer value within the discrete integer values. Continuing with this example, the control range for the direction tool parameter of the linear gradient tool is 0 to 360 degrees, such that the corresponding parameter configuration for the direction tool refers to a degree value.

[0035] In the following discussion, an example environment configured to employ the techniques described herein is described. Example procedures configured for performance in the example environment as well as other environments are also described. Therefore, execution of the example procedures is not limited to the example environment, and the example environment is not limited to execution of the example procedures.

[0036] Sample Environment

[0037] Figure 1 is an illustration of a digital media environment 100 in an example implementation operable to employ the automated digital tool identification techniques described herein. The illustrated environment 100 includes a computing device 102, which is configurable in a variety of ways.

[0038] For example, computing device 102 can be configured as a desktop computer, a laptop computer, a mobile device (e.g., assuming a handheld configuration, such as the illustrated tablet computer or mobile phone), etc. Thus, computing devices 102 range from full-resource devices with large amounts of memory and processor resources (e.g., personal computers, game consoles) to low-resource devices with limited memory and / or processing resources (e.g., mobile devices). Additionally, although a single computing device 102 is shown, computing device 102 also represents multiple different devices, such as multiple servers used by an enterprise to perform operations "through the cloud."

[0039] Computing device 102 includes a visual lens system 104. Visual lens system 104 is implemented at least in part in hardware of computing device 102 to process image data 106, which is illustrated as being maintained in a storage device 108 of computing device 102, to generate an image tool description 110 that describes at least one digital tool used to generate a visual appearance of image data 106 and at least one parameter 112 for each digital tool detected in the image data, which specifically describes how the digital tool was used to achieve the visual appearance. Although illustrated as being implemented locally at computing device 102, the functionality of visual lens system 104 may be implemented in whole or in part via functionality available via a network 114, such as part of a web service or "in the cloud," as described below with respect to Figure 12 described in further detail.

[0040] Image data 106 represents raster image data, such as a rasterized graphic or bitmap image. In one example, the image data includes a bit structure describing a pixel grid, with each pixel in the grid specified by a number of bits, and the size of the rasterized image defined by the width and height of the pixel grid. In this manner, the number of bits used for each pixel in image data 106 specifies the color depth of the corresponding raster image, wherein the additional bits per pixel enable image data 106 to express a wider range of colors relative to image data including fewer bits per pixel.

[0041] Computing device 102 is configured to render image data 106 in a user interface 116 of a visual lens system 104 at a display device 118 communicatively coupled to computing device 102. By processing image data 106 using the techniques described herein, image data 106 is rendered as an interactive image 120 that includes at least one selectable user interface component 122 that indicates an area of ​​image data 106 where a digital tool is to be used to achieve a particular visual appearance. Selectable user interface component 122 is configured to cause display of an image tool description 110 generated by visual lens system 104 for the area indicated by selectable user interface component 122.

[0042] Image tool description 110 includes at least an identification of a digital tool that can be used to achieve the visual appearance of the corresponding area and at least one parameter 112 for the identified digital tool, thereby enabling a user of visual lens system 104 to recreate the visual appearance in different image data. For example, in digital media environment 100, image tool description 110 corresponding to the area indicated by selectable user interface component 122 in interactive image 120 describes that a "Gradient" digital tool with an orange parameter can be used to achieve the visual appearance of the image area.

[0043] While the parameters 112 are depicted in the illustrated example as a visual representation of a specific color used in conjunction with the "Gradient" digital tool identified by the image tool description 110, the parameters 112 of the image tool description 110 are presented in various ways within the interactive image 120. Examples of presenting the parameters 112 include a textual description (e.g., the name of the depicted color used with the digital tool, such as "Tanger Orange"), a numeric parameter value (e.g., a hexadecimal code, a decimal code, etc., indicating the depicted color used with the digital tool), combinations thereof, etc. In some implementations, the image tool description 110 includes additional information describing the digital tool, such as a uniform resource locator (URL) pointing to relevant information about the digital tool, such as instructional information for using the digital tool.

[0044] Image tool description 110 enables a user to apply parameters 112 of image tool description 110 to one or more different images in user interface 116 of visual lens system 104. For example, image tool description 110 is configured to be selectable via a drag-and-drop operation and transferred to a region of different image data, which causes visual lens system 104 to apply parameters 112 of a digital tool corresponding to image tool description 110 to the region of the different image data. This transfer is fully parametric because the application of parameters 112 is performed by modifying specific parameters of one or more digital tools available to visual lens system 104, thereby enabling a user of the visual lens system to further adjust parameters 112 to achieve a desired resulting appearance for the different image data. Further discussion of this and other examples is included in the following sections and is illustrated using the corresponding figures.

[0045] Generally, the functionality, features, and concepts described with respect to the examples above and below can be employed in the context of the example processes described in that section. Further, the functionality, features, and concepts described with respect to the different figures and examples in this document are interchangeable with each other and are not limited to implementation in the context of a particular figure or process. Moreover, the blocks associated with the different representative processes and corresponding figures herein are configured to be applied together and / or combined in different ways. Therefore, the various functionality, features, and concepts described with respect to the different example environments, devices, components, figures, and processes herein can be used in any suitable combination and are not limited to the combinations represented by the examples listed in this description.

[0046] Automated digital tool identification via rasterized images

[0047] Figure 2 Depicting system 200 in an example implementation, showing in more detail Figure 1 1. To begin in this example, a tool detection module 202 is employed by the visual lens system 104, which implements a tool region detection network 204. The tool region detection network 204 represents the functionality of the visual lens system 104 to identify one or more image regions 206 in image data 106 created using one or more digital tools. The tool detection module 202 is configured to first process the image data 106 to conform to the input format on which the tool region detection network 204 was trained, as described in further detail below.

[0048] In an example implementation, the tool detection module 202 processes the image data 106 into a multi-channel (e.g., 3-channel) image tensor having an arbitrary width W and height H. For each image region 206 in the image data 106 that is generated using one or more digital tools detected by the tool region detection network 204, the tool detection module 202 outputs a plurality of tool masks 208. In the illustrated example, the plurality of tool masks 208 generated for each image region 206 identified in the image data 106 that is generated using one or more digital tools are represented as tool mask 208(1), tool mask 208(2), and tool mask 208(n), where n represents any suitable integer. As described in further detail below with respect to training the tool region detection network 204, n represents an integer number of different digital tools for which the tool region detection network 204 is trained, such that each tool mask 208 output by the tool region detection network 204 corresponds to a different digital tool.

[0049] Tool mask 208 represents a probability mask for image region 206 that indicates whether a corresponding digital tool was used to achieve the visual appearance of image region 206. For example, in a binary probability mask implementation, a pixel of tool mask 208 is assigned a value of 0 to indicate that the corresponding digital tool was not used, and a value of 1 indicates that the corresponding digital tool was used to achieve the visual appearance of the pixel. In an example implementation, tool mask 208(1) represents a probability mask that indicates whether a "brush" tool was used to achieve the visual appearance of image region 206, tool mask 208(2) represents a probability mask that indicates whether a "glow" tool was used to achieve the visual appearance of image region 206, and tool mask 208(n) represents a probability mask that indicates whether a "gradient" tool was used to achieve the visual appearance of image region 206. In some implementations, image region 206 represents the entire image data 106, such that each tool mask 208 includes a probability mask that indicates whether a digital tool was used for each pixel of image data 106. Alternatively, in some implementations, image region 206 represents a subset of image data 106 that includes a width W and a height H that is smaller than the tensor representing image data 106 .

[0050] like Figure 3, image data 302 representing an example instance of the image data 106 is processed by the tool region detection network 204 to generate a plurality of tool masks 208 represented by tool masks 304, 306, and 308. In the example implementation 300, the image region 206 of the image data 302 of each of the tool masks 304, 306, and 308 encompasses the entire image data 302. The tool masks 304 and 308 are illustrated as corresponding to respective digital tools that are not detected as being used in the image data 302 by the tool region detection network 204. The tool mask 306 indicates that the tool region detection network 204 detected the use of a digital tool in the image data 302, the digital tool corresponding to the region of the tool mask 306 represented in white. For example, tool mask 306 corresponds to a "noise" digital tool and represents the appearance that the tool region detection network 204 detected that the noisy digital tool was used to generate a shape emanating from the head of the person depicted in the image data 302 and extending toward the upper right corner of the image data 302. Thus, tool detection module 202 is configured to generate a tool mask 208 for each digital tool that the tool region detection network 204 is trained on, thereby indicating whether a digital tool is detected in the image data 106 for a plurality of different digital tools.

[0051] According to one or more implementations, each tool mask 208 is a reduced resolution representation of the image area 206, such as a W / 8 and H / 8 resolution probability mask indicating digital tool usage in the image area 206. By generating each tool mask 208 at a reduced resolution relative to the image area 206, the tool detection module 202 reduces the amount of computational resources required to generate the tool masks 208 at the same resolution as the image area 206 without compromising the accuracy of the image tool descriptions 110 included in the interactive image 120 of the image data 106. The reduced resolution of each tool mask 208 relative to the image data 106 and the image area 206 does not negatively impact the accuracy of the visual lens system 104 because the overall detection of digital tool usage in the image data 106 is more important than the precise location of the digital tool usage in the image data 106. The location of the digital tool usage is resolved by the visual lens system 104, as described in further detail below.

[0052] The tool detection module 202 passes each image region 206 generated from the image data 106 and a corresponding plurality of tool masks 208 for each image region 206 to a filling module 210. The filling module 210 is configured to generate a padded image region 212 for each image region 206, wherein each padded image region 212 includes a padded tool mask 214. The number of padded tool masks 214 generated by the filling module 210 is represented as an integer n, such that each tool mask 208 in the image region 206 is represented by a corresponding padded tool mask 214. Each padded image region 212 is generated by the filling module 210 to have a size specified by a tool parameter estimation network 218, which is implemented by a parameter estimation module 216 of the visual lens system 104. For example, in some implementations, the padded image region 212 is configured as a 128x128 pixel region.

[0053] Alternatively, the fill module 210 is configured to generate the filled image region 212 in any size suitable for input to the tool parameter estimation network 218. Regardless of the size of the filled image region 212, the fill module 210 generates the filled image region 212 as an image centered around a portion of the image region 206 that is detected as being processed by one of the digital tools corresponding to the corresponding tool mask 208. In scenarios where the image region 206 is larger than the size of the filled image region 212 to be generated for input to the tool parameter estimation network 218, the fill module 210 is configured to generate the filled image region 212 by cropping the image region 206 while maintaining a portion of the tool mask 208 indicating the presence of a digital tool centered around an aperture of the fill tool mask 214.

[0054] Alternatively, in a scenario where the image area 206 is smaller than the size of the filled image area 212 to be generated, the fill module 210 is configured to generate each fill tool mask 214 by centering an area of ​​the corresponding tool mask 208 (indicating the presence of a digital tool centered about the aperture-filled tool mask 214) and adding pixels to the fill tool mask 214 (where the values ​​indicate the absence of a digital tool) until appropriate width and height pixel dimensions are achieved.

[0055] The foregoing combination is achieved by the fill module 210 in certain circumstances, such as scenarios where the area corresponding to the tool mask 208 is centered (indicating the presence of a digital tool in a corner of the image area 206) and is otherwise larger than necessary to fill the image area 212. In this manner, each fill tool mask 214 generated by the fill module 210 achieves a uniform size, wherein the area corresponding to the tool mask 208 indicates the presence of a digital tool in the image data 106 centered about the aperture of the fill tool mask 214.

[0056] like Figure 4 As shown in the example implementation 400 of , image data 402 represents an example instance of the image data 106. Image regions 404, 406, and 408 are example instances of the image region 206 identified by the tool region detection network 204 as having a digital tool applied thereto in the image data 402. As depicted in the example implementation 400, each of the image regions 404, 406, and 408 has a different size and has a corresponding tool mask 410, 412, and 414 indicating that a corresponding tool is detected at the image regions 404, 406, and 408.

[0057] Filled tool masks 416, 418, and 420 are examples of filled tool masks 214 generated by the filling module 210 from tool masks 410, 412, and 414, respectively. As depicted in the example implementation 400, the filled tool masks 416, 418, and 420 are generated by the filling module 210 to have a uniform size that conforms to the data input size of the tool parameter estimation network 218. For example, the filled tool mask 416 represents the tool mask 410 centered about the aperture of the filled tool mask 416, with the filling around the tool mask 410 indicating that no digital tool extensions corresponding to the tool mask 410 were detected to satisfy the data input size for the tool parameter estimation network 218. The filling module 210 is configured to generate a filled image region 212 corresponding to each of the tool masks 410, 412, and 414 by copying a portion of the image data 402 at the data input size of the tool parameter estimation network 218 (centered about the corresponding image region 206). For example, when generating the filled image region 212 for the image region 408 , the fill module 210 copies a portion of the image data 402 centered around the image region 408 in the image data 402 , with a data input size corresponding to the fill tool mask 420 .

[0058] The parameter estimation module 216 generates a concatenation of each filled image region 212 and the corresponding filled tool mask 214 for the image data 106 and feeds the concatenation as input to a tool parameter estimation network 218. The tool parameter estimation network 218 is configured to automatically identify the parameters 112 corresponding to the corresponding digital tool indicated by the filled tool mask 214 for each filled image region 212. Training the tool parameter estimation network 218 is provided in further detail below, with a description of automatically identifying the parameters 112. Through the training process described herein, the tool parameter estimation network 218 is informed of a discrete set of possible parameters 112 for the digital tool corresponding to the filled tool mask 214, and is also informed of a discrete set of possible parameter configurations for each possible parameter in the discrete set of possible parameters 112. In this manner, the tool parameter estimation network 218 generates the parameters 112 for each image region 206 as a classification problem by outputting a probability of the filled tool mask 214 for each possible parameter configuration representing each possible parameter 112 of the corresponding digital tool.

[0059] The tool parameter estimation network 218 is configured to select, for each image region 206 , a top-ranked configuration (ie, a digital tool parameter configuration with the most likely probability) as the parameters 112 of the digital tool corresponding to the image region 206 . Figure 5 As shown in the example implementation 500 of FIGURE 5, a filled image region 502 and a corresponding fill tool mask 504 are concatenated by the parameter estimation module 216 and provided as input to the tool parameter estimation network 218. Having been trained to recognize digital tools including at least the radial gradient tool and the noise tool, the tool parameter estimation network 218 determines from the fill tool mask 504 that the tool region detection network 204 identifies the use of the radial gradient tool at the image region 206 of the image data 106 corresponding to the filled image region 502. Based on this determination, the tool parameter estimation network 218 predicts a possible parameter configuration from a discrete set of possible parameters for the radial gradient tool that is used to achieve the visual appearance of the filled image region 502. The tool parameter estimation network 218 then selects the most likely parameter configuration(s) and outputs the selected parameter configuration(s) as the parameters 112.

[0060] For example, in the example implementation 500, the tool parameter estimation network 218 identifies that the radial gradient tool includes three discrete parameters: a "color 1" parameter, a "color 2" parameter, and a "direction" parameter. The tool parameter estimation network 218 then identifies that the most likely "color 1" value for achieving the visual appearance of the filled image region 502 is "93," the most likely "color 2" value for achieving the visual appearance of the filled image region 502 is "500," and the most likely "direction" value for achieving the visual appearance of the filled image region 502 is "51." Collectively, these most likely parameter configuration values ​​are output as parameters 506 for the radial gradient tool, represented by the fill tool mask 504.

[0061] In a similar manner, the filled image region 508 and the corresponding fill tool mask 510 are concatenated by the parameter estimation module 216 and provided as input to the tool parameter estimation network 218. The tool parameter estimation network 218 determines that the fill tool mask 510 corresponds to a noise tool detected at the image region 206 of the image data 106 corresponding to the filled image region 508. Based on this determination, the tool parameter estimation network 218 predicts possible parameter configurations for achieving the visual appearance of the filled image region 508 from a discrete set of possible parameters for the noise tool. The tool parameter estimation network 218 then selects the most likely parameter configuration(s) and outputs the selected parameter configuration(s) as the parameters 112. For example, in the example implementation 500, the tool parameter estimation network 218 identifies that the noise tool is associated with a single possible parameter specifying the "intensity" of the noise tool. The tool parameter estimation network 218 then identifies that the most likely “Intensity” value for achieving the visual appearance of filling the image region 508 is “12” and outputs this predicted parameter configuration as parameters 512 for the noise tool represented by the fill tool mask 510 .

[0062] The parameter estimation module 216 then passes the determined parameters 112 for each image region 206 to the tool description module 220, which is configured to generate an image tool description 110 for the image region 206. The image tool description 110 provides information identifying the parameters 112 and the corresponding digital tool used to achieve the visual appearance of the image region 206. The information included in the image tool description 110 is presented in a variety of possible configurations, such as a graphical rendering, a textual description, an auditory description, a combination thereof, etc. In some implementations, the image tool description 110 includes additional controls and / or information related to the digital tool used to achieve the visual appearance of the image region 206. The image tool description 110 is then incorporated into the image data 106 and output by the rendering module 222 as the interactive image 120 of the image data 106.

[0063] For example, Figure 6Example implementation 600 illustrates an interactive image 120 output in user interface 116 of visual lens system 104, as generated from image data 402. Example implementation 600 depicts cursor 602, which represents user input to interactive image 120 at a region of interactive image 120 that includes a display of clouds (e.g., a region of interactive image 120 that includes image region 404). In response to receiving input at a portion of interactive image 120 that includes image region 206, which includes a detected digital tool for achieving the visual appearance of the portion, visual lens system 104 is configured to display image tool description 110 for image region 206. In example implementation 600, image tool description 110 indicates that a radial gradient tool with parameters 112 was used to achieve the visual appearance of the depicted clouds. Image tool description 110 also includes controls 604 and 606, respectively, that are selectable via user input to cause an action associated with the digital tool (e.g., radial gradient tool) identified by image tool description 110 to be performed.

[0064] Control 604 includes an option for accessing “digital tool information” corresponding to the digital tool identified by image tool description 110. For example, in response to selection of control 604, visual lens system 104 presents additional information describing the radial gradient tool in user interface 116. In some implementations, such additional information includes a user manual for the radial gradient tool, links to instructional materials for using the radial gradient tool, visual indications of where to access the radial gradient tool in user interface 116, combinations thereof, and the like.

[0065] In this way, the image tool description 110 not only identifies the digital tool used to achieve the visual appearance of the selected image area 206 in the interactive image 120, but also assists users of the visual lens system 104 in becoming familiar with using the digital tool in their own projects. In addition to educating users of the visual lens system 104 about the potential applications of the identified digital tool, the visual lens system 104 enables users to import parameters 112 directly into their own creations.

[0066] For example, control 606 includes an option to "copy parameters" identified by image tool description 110. In response to selection of control 606, visual lens system 104 stores (e.g., in storage device 108) information describing the digital tool identified by image tool description 110 and the specific parameter configuration of the digital tool represented by parameters 112. For example, in example implementation 600, selection of control 606 causes visual lens system 104 to store the radial gradient tool and its associated parameters "color 1:93," "color 2:500," and "direction:51" for subsequent use by visual lens system 104. In this way, users of visual lens system 104 are able to browse different image data 106 for inspiration to generate their own work, and save aspects of image data 106 for their own work for subsequent access.

[0067] Figure 7 An example implementation 700 of interactive image 120 illustrates an interactive image 120 in which parameters 112 of an image tool description 110 are transferred to different image data. For example, example implementation 700 depicts interactive image 120 displayed in user interface 116 of visual lens system 104. Interactive image 120 is depicted as displaying image tool description 110 for image area 206 selected by user input (e.g., via a cursor click). In example implementation 700, image area 206 is visually indicated as being detected by visual lens system 104 as having a visual appearance generated by a digital tool, such that a user viewing interactive image 120 is informed that interacting with (e.g., selecting) image area 206 resulted in display of image tool description 110. In implementation, image area 206 is indicated as having a visual appearance achievable via a digital tool by visually distinguishing image area 206 from the rest of interactive image 120. For example, the example implementation 700 depicts a selectable icon on the image area 206 in the form of a blue circle with a white “+” sign, indicating that the visual appearance of the image area 206 is achievable via the digital tool. In response to selection of the selectable icon, the example implementation 700 displays a visual representation of the parameters 112 of the linear gradient digital tool used to generate the visual appearance of the image area 206 as part of the image tool description 110 displayed in response to selection of the image area 206.

[0068] In the example implementation 700, user input is depicted as progressing along a path 702 from selection of the image area 206 to selection of the parameter 112, and transferring the parameter 112 to other image data 704 displayed in the user interface 116. The selection of the parameter 112 is implemented by the visual lens system 104 in various ways, such as via selection of the control 606 illustrated in the example implementation 600, via a drag-and-drop input operation, etc. As depicted in the example implementation 700, the selection of the parameter 112 and the transfer to the image data 704 result in the parameter 112 for the linear gradient specified in the image tool description 110 being transferred to the image data 704, which modifies the image data 704 to have the visual appearance indicated by the image data 706.

[0069] In this manner, the visual lens system 104 enables a user to transfer parameters 112 of a digital tool directly from a rasterized image to different image data, while preserving the parameters 112 in a manner that enables further customization of the parameters 112 by the visual lens system 104 (e.g., via user input from the user interface 116). In some implementations, the transfer of the digital tool parameters 112 from the image data 106 to the different image data is accompanied by a display of a widget or control interface for the digital tool associated with the parameters 112. This display of the widget or control interface is automatically populated with the parameters 112, and the parameters 112 can be further refined via respective controls associated with one of the discrete parameters 112 of the digital tool.

[0070] Having considered example implementations of the visual lens system 104 , consider now example techniques for generating the tool region detection network 204 and the tool parameter estimation network 218 to achieve the performance of the techniques described herein.

[0071] Figure 8A and Figure 8B An example implementation 800 is illustrated for generating a tool region detection network 204 trained to automatically identify uses of a plurality of different digital tools in rasterized image data 106 and generating a tool parameter estimation network 218 trained to automatically determine parameter configurations for respective ones of the digital tools used to generate the rasterized image data 106. In some implementations, a training system 802 represents functionality of the visual lens system 104 to generate the trained tool region detection network 204 and the trained tool parameter estimation network 218. Alternatively, in some implementations, the training system 802 is implemented remotely from the visual lens system 104, such as at a computing device different from the computing device 102, which transfers the trained tool region detection network 204 and the tool parameter estimation network 218 to the visual lens system 104.

[0072] First, vector image data 804 is obtained by a training system 802. The vector image data 804 represents a plurality of vector images, each having one or more digital tools applied to at least a portion of the vector image. The training system 802 implements a labeling module 806 that is configured to determine the type of digital tool applied to each vector image in the vector image data 804 and generate a labeled vector image 808 for each image in the vector image data 804 that identifies the applied digital tool. In addition to identifying the type of digital tool applied to each vector image, the labeled vector image 808 includes information identifying one or more parameters of the digital tool used to stylize at least a portion of the vector image, as well as a parameter configuration for each of the one or more parameters. In an implementation, the labeling module 806 is configured to generate the vector image 808 automatically (e.g., independently of user input) via explicit user input or a combination thereof.

[0073] For example, the labeling module 806 is configured to automatically identify the type of digital tool applied to the vector image of the vector image data 804 by examining the metadata of the vector image, which metadata describes the area of ​​the vector image to which the digital tool is applied and the specific type of digital tool applied to the area. In an implementation in which the labeling module 806 automatically identifies the type of digital tool, the labeling module 806 is configured to implement an image segmentation network, such as a Mask R-CNN architecture, a SpineNet architecture, etc. In an example implementation in which the Mask R-CNN architecture is implemented by the labeling module 806, example parameters include 300 regions of interest, a batch size of 4, use of a ResNeXt backbone, and a learning rate of 0.001 using an Adam optimizer. These example parameters are merely illustrative of one example implementation, and different parameters may be used for different implementations.

[0074] In implementations where user input is used to identify the type of digital tool applied to the vector image, a user of a computing device implementing the training system assigns labels to individual images of the vector image data 804 by manually specifying (e.g., selecting from a list, textually describing, etc.) the digital tool used to achieve the visual appearance of the image. In implementations where a combination of automatic and user input-based labeling methods are used, user input is used to verify the accuracy of the labeled vector images 808 generated by the labeling module 806. In this manner, each vector image 808 represents an instance of a vector image included in the vector image data 804, having associated information identifying the digital tool used to achieve the visual appearance of each geometric element in the vector image.

[0075] Example digital tools identified by vector image 808 include, but are not limited to, linear gradient, radial gradient, freeform gradient, drop shadow, glow, noise, Gaussian blur, halftone, and a brush tool. In implementations, the specific digital tool identified by vector image 808 is customized for the specific digital graphics editing software to be implemented with visual lens system 104. In this manner, vector image 808 is not limited to identifying the example digital tools described herein and represents identifying any type of digital tool that can be used to adjust, modify, or define the visual appearance of vector graphics data.

[0076] The marking module 806 provides the plurality of vector images 808 to an enhancement module 810, which is configured to generate at least one enhanced version of each vector image 808. The enhanced versions of the vector images 808 are generated by altering geometric elements applied by digital tools in the vector images 808. For example, the enhancement module 810 generates the enhanced versions of the vector images 808 by changing the position of stylized geometric elements (e.g., image areas of the vector images stylized via application of one or more digital tools). Alternatively or additionally, the enhancement module 810 generates the enhanced versions of the vector images 808 by changing the size of the stylized geometric elements in the vector images 808. Alternatively or additionally, the enhancement module 810 generates the enhanced versions of the vector images 808 by altering one or more underlying colors of the stylized geometric elements, thereby creating a different appearance of the same effect applied to the vector images 808 by one or more digital tools.

[0077] Alternatively or additionally, enhancement module 810 generates an enhanced version of vector image 808 by changing at least one parameter of a digital tool applied to the stylized geometric element (e.g., changing the radius of a Gaussian blur digital tool, the intensity value of a noise tool, etc.). Alternatively or additionally, enhancement module 810 generates an enhanced version of vector image 808 by changing the hierarchical arrangement (e.g., z-order placement) of at least one stylized geometric element in the vector image, such that a portion of the at least one stylized geometric element is visually obscured by at least one other geometric element in the vector image. Alternatively or additionally, enhancement module 810 generates an enhanced version of vector image 808 by copying a stylized geometric element from one vector image in vector image data 804 and inserting the stylized geometric element into another vector image in vector image data 804, thereby enabling visualization of the same stylized geometric element in different contexts. Vector image data 812 thus represents the labeled vector images included in vector image data 804, which are diversified to reflect the different ways in which the same digital tool is applied to achieve different visual appearances. For example, when multiple enhanced instances of vector image 808 are included in vector image data 812, a single instance of vector image 808 is enhanced in a number of different ways as indicated above to represent different visual appearances of the image data that are achievable via digital tools using tags corresponding to vector image 808.

[0078] Each enhanced version of the vector image 808 is assigned the same label as the vector image 808, which identifies the digital tool used to generate the stylized geometric elements that were modified as part of the enhancement. The enhanced and labeled versions of the vector image 808 are then output by the enhancement module 810 as enhanced vector image data 812. From the vector image data 812, the training sample module 814 generates a plurality of paired training samples 816. Each of the training samples 816 includes a labeled vector image 818 obtained from the vector image data 812 and a corresponding rasterized image 820 representing a rasterization of the labeled vector image 818.

[0079] In contrast to labeled vector images 818, which include metadata identifying various geometric elements as constrained by points and connected lines (e.g., splines, Bezier curves, etc.) and one or more attributes defining the visual appearance of the geometric elements, rasterized images 820 include only pixel value information, which specifies, in bits, the color depicted by each pixel. Labeled vector images 818 thus serve as ground truth representations of rasterized images 820, indicating whether the digital tool represented by the label associated with labeled vector images 818 was used to achieve the visual appearance of the pixels in rasterized images 820. The rasterized images 820 of the various training samples 816 are rasterized at different resolutions, such as ranging from 320 pixels to 1024 pixels per rasterized image 820. By varying the resolution of the rasterized images 820 included in the training samples 816, the training system 802 is configured to generate training instances for the tool region detection network 204 and the tool parameter estimation network 218 that accurately adapt to a range of different image data 106 resolutions.

[0080] The training samples 816 are provided to a cropping module 822, which is configured to generate at least one cropped training sample 824 for each training sample 816. Each cropped training sample 824 includes a cropped labeled vector image 826 and a cropped rasterized image 828. The cropped labeled vector image 826 represents a cropped instance of the labeled vector image 818, and the cropped rasterized image 828 represents a similar cropped instance of the rasterized image 820 generated for the corresponding training sample 816. Each of the cropped labeled vector image 826 and the cropped rasterized image 828 includes stylized geometric elements of the base vector image 808, which are generated within a cropped aperture to account for the different sizes and different visual saliences of the image regions 206 included in the image data 106 processed by the visual lens system 104. In an implementation, the aperture sizes of the cropped labeled vector image 826 and the cropped rasterized image 828 comprising the cropped training samples 824 are randomly generated by the cropping module 822 such that the sizes of the cropped training samples 824 are different relative to each other.

[0081] continue Figure 8B, the cropping module 822 passes the cropped training sample 824 to the masking module 830. The masking module 830 is configured to generate a training sample mask 832, which serves as the ground truth for the cropped training sample 824. To do so, the masking module 830 creates a mask that encompasses the geometry of the stylized geometric elements represented in the cropped labeled vector image 826, which is then dilated and rasterized against a black background. In this way, the training sample mask 832 represents the ground truth to be output by the trained tool region detection network 204 in response to receiving input of the cropped rasterized image 828 for which the training sample mask 832 was generated.

[0082] The cropped training sample 824 and the training sample mask 832 are also passed to a parameter module 834 for use in creating a parameter set 836 and an associated parameter configuration 838 for each of the plurality of digital tools represented in the cropped training sample 824. To do so, the parameter module 834 identifies possible parameters 836 that can be used to control the corresponding digital tool for generating the stylized geometric elements included in the cropped labeled vector image 826. For each parameter 836, the parameter module 834 identifies a possible parameter configuration 838 for the parameter 836.

[0083] For example, in an example implementation in which the cropped labeled vector image 826 includes stylized geometric elements generated via application of a "Noise" digital tool, the parameter module 834 identifies a single "Intensity" parameter 836 as being usable for controlling the "Noise" tool. Control of the "Intensity" parameter 836 is identified as being scalable in percentage values ​​ranging from 1 to 100 in single percentile increments. Thus, the parameter module 834 identifies that the "Noise" tool used to generate the cropped labeled vector image 826 includes a single "Intensity" parameter 836 having 100 discrete parameter configurations 838 (e.g., 1, 2, ..., 100).

[0084] As another example, in an implementation where the cropped labeled vector image 826 includes stylized geometric elements generated via application of a "Linear Gradient" digital tool, the parameter module 834 identifies three parameters 836 as being usable for controlling the "Linear Gradient" tool: a first color value, a second color value, and a direction value. Controls for the first color value and the second color value are identified as being assignable as integer values ​​ranging from zero to 256 (e.g., colors displayable with an 8-bit pixel depth), and a control for the direction value is identified as being assignable as degrees as integer values ​​ranging from 1 to 360. Thus, the parameter module 834 identifies that the "Linear Gradient" tool includes three parameters 836, and corresponding parameter configurations 838 for the three parameters include 257 discrete options (e.g., for the first color value and the second color value) and 360 discrete options (e.g., for the direction value).

[0085] As yet another example, in an implementation where the cropped marked image 826 includes a stylized geometric element generated via application of a "drop shadow" digital tool, the parameter module 834 identifies six parameters 836 that can be used to control the digital tool: mode, opacity, X and Y offset values, blur, color, and darkness. For example, mode specifies the blending mode of the drop shadow, opacity specifies the percentage opacity of the drop shadow, X and Y offset values ​​specify the distance to offset the drop shadow from the geometric element, blue specifies the distance from the edge of the shadow where blurring will occur, color specifies the color of the shadow, and darkness specifies the percentage of black to be added to the shadow (e.g., 0% to 100% added black).

[0086] In yet another example, in an implementation where the cropped labeled image 826 includes stylized geometric elements generated via application of an "Inner / Outer Glow" digital tool, parameter module 834 identifies three to five parameters 836 that can be used to control the digital tool: mode, opacity, blur, center, and edge. Mode specifies the blending mode for the glow, opacity specifies the percentage opacity of the glow, and blur specifies the distance from the center or edge of the selection within the geometric element at which blurring will occur. The center and edge parameters are example parameters 836 that only function when the inner / outer glow digital tool is applied as an inner glow, where center applies a glow emanating from the center of the geometric element and edge applies a glow emanating from the inner edges of the geometric element, respectively. Thus, parameter module 834 represents the functionality of training system 802 to identify possible ways in which the digital tool used to generate the cropped labeled vector image 826 can be configured based on parameters 836 and their corresponding parameter configurations 838.

[0087] The cropped training sample 824, training sample mask 832, and parameters 836 generated by the training system 802 are then passed to the network generation module 840 for use in training the tool region detection network 204 and the tool parameter estimation network 218. To train the tool region detection network 204, the network generation module 840 implements a segmentation network 842, which represents a semantic segmentation network architecture, such as a ResNet18 backbone. Alternatively, the segmentation network 842 is implemented as a UNet architecture, a dilated ResNet architecture, a Gated-SCNN, or the like. When training the tool region detection network 204, the network generation module 840 provides the cropped rasterized image 828 of the cropped training sample 824 as input to the segmentation network 842, and causes the segmentation network 842 to output a classification for each pixel of the cropped rasterized image 828, indicating whether a digital tool is detected as being applied to the pixel. This classification is performed for each of a plurality of different digital tool types and is not limited to the example digital tool types described herein.

[0088] In an implementation, the cropped rasterized image 828 is fed to the segmentation network 842 as a 3-channel image tensor with arbitrary width W and height H, and the segmentation network 842 is used to output a single channel for each tool in a plurality of tools T. In some implementations, the single channel output by the segmentation network 842 for each T is generated at a reduced resolution (e.g., W / 8 and H / 8). For example, when the segmentation network 842 is implemented as a ResNet18 backbone, the segmentation network 842 downsamples the input cropped rasterized image 828 to W / 16 and H / 16 before applying the convolution block to generate a tensor of size (batch x H / 8 x W / 8 x T). As an example, training of the tool region detection network 204 is performed using a two-class cross entropy loss for each tool T (e.g., a tool present at a pixel vs. a tool not present at a pixel), with a batch size of 32 and a learning rate of 1e -4 .

[0089] During training, the network generation module 840 filters the output generated by the segmentation network 842 to predict weak detections and smooth detection areas, taking into account the cropped rasterized image 828 as input, until the output of the segmentation network 842 matches the training sample mask 832. For example, the network generation module 840 resizes each output generated by the segmentation network 842 to the original resolution of the vector image 808 from which the cropped rasterized image 828 was generated, and assigns a binary tool detection value to each pixel of the resized binary mask based on a specified threshold. For example, if the specified threshold is set to 0.5, the network generation module 840 assigns a value of 1 to all pixels in the mask output by the segmentation network 842, indicating a "tool present" confidence value greater than or equal to 0.5, and assigns a value of 0 to all pixels that do not have a "tool present" confidence value that meets the 0.5 threshold. The resulting binary mask indicating "tool present" or "tool not present" at each pixel is then dilated and eroded (e.g., dilated by five pixels and eroded by five pixels). Then, based on the determination of whether the pixel counts included therein meet a pixel count threshold (e.g., 25 2 The resulting binary mask for a given tool is compared against the corresponding ground truth training sample mask 832 for that tool, and the result of the comparison is used to guide learning of the tool region detection network 204 (e.g., by adjusting one or more internal weights) until training is complete (e.g., until the tool region detection network 204 generates the training sample mask 832 as output for the corresponding cropped rasterized image 828, which was provided as input).

[0090] To train the tool parameter estimation network 218, the network generation module 840 quantizes each parameter 836 of a particular digital tool into a discrete set of bins, where the number of discrete bins for each parameter 836 is defined by the parameter configuration 838 and determined by the parameter module 834. By quantizing the parameters 836 of a particular tool into discrete bins, the training system 802 simplifies the problem to be solved by the tool parameter estimation network 218 into a classification problem, rather than a regression problem of tool parameter targets. Therefore, to generate the tool parameter estimation network 218, the network generation module 840 implements a classification network 844.

[0091] The classification network 844 represents a family of different image classification networks, such as a ResNet18 backbone pre-trained with the ImageNet classifier used as the segmentation network 842 according to one or more implementations. Alternatively, according to one or more implementations, the network generation module 840 is configured to perform pre-training of a ResNet18 backbone with the ImageNet classifier to generate the classification network 844.

[0092] For each cropped training sample 824, an appropriate digital tool for generating a cropped rasterized image 828 is identified from the cropped labeled vector image 826. The network generation module 840 then identifies associated one or more parameters 836 for the identified digital tool and possible parameter configurations 838 for each of the one or more parameters 836, and passes the identified parameters 836 and their associated parameter configurations 838 along with the training sample mask 832 corresponding to the cropped training sample 824 as input to the classification network 844. Specifically, the network generation module 840 concatenates the cropped rasterized image 828 and the corresponding training sample mask 832, and provides the concatenation as input to the classification network 844.

[0093] The objective function (which the classification network 844 is trained to generate the tool parameter estimation network 218) thus causes the classification network 844 to generate a confidence value (regarding whether a particular parameter configuration 838 of the digital tool parameters 836 was used to implement the training sample mask 832) for each discrete bin represented by the respective parameter configuration 838 of each digital tool parameter 836. In this way, the output of the classification network 844 for each cropped training sample 824 represents a probability distribution over the different possible parameter configurations 838 for the parameters 836 of the digital tool as identified by the cropped labeled vector image 828 used in the cropped rasterized image 828.

[0094] For example, consider an example implementation in which the cropped labeled vector image 826 indicates that a “glow” digital tool was used to achieve the visual appearance of the cropped rasterized image 828, and corresponding parameters 836 for the “glow” digital tool include a radius parameter having 500 different possible parameter configurations 838. In this example implementation, the network generation module 840 causes the classification network 844 to assign confidence values ​​indicating whether the parameter configuration 838 was used to achieve the visual appearance of the cropped rasterized image 828 for each of the 500 different parameter configurations 838.

[0095] In an example scenario where the tool parameter estimation network 218 is generated from a ResNet18 classifier pre-trained using an ImageNet classifier implementation of the classification network 844, the network generation module 840 defines the input data size for the tool parameter estimation network 218 and formats the training data to the appropriate input data size. For example, the network generation module 840 is configured to specify an input data size of 128x128 and processes each cropped rasterized image 828 used as training data into a 128x128 image centered around the stylized geometric element identified by the label in the corresponding cropped labeled vector image 826. The ResNet18 backbone then transforms the input 128x128x4 image into a 2048 feature vector and uses a different fully connected linear layer to match each parameter 836 set and associated parameter configuration 838 reservoir for each digital tool identified by the cropped labeled vector image 826 in the cropped training sample 824.

[0096] Using multiple (e.g., 8) different cropped training samples 824 corresponding to a specific digital tool, multiple cropped rasterized images 828 are provided as input to the classification network 844, and a standard cross entropy loss is used for each tool parameter 836 with a learning rate of 1e -3, the network generation module 840 continues to train the tool parameter estimation network. The confidence values ​​for each bin represented by the different parameter configurations 838 of the corresponding parameters 836 of the digital tool are used by the network generation module 840 to select the top-ranked parameter configurations 838 for each parameter 836 as the predicted digital tool parameters for generating the cropped rasterized image 828, which is used to generate the cropped training sample 824 processed by the classification network 844. The predicted digital tool parameters for the cropped rasterized image 828 are then compared with the metadata included in the cropped labeled vector image 826 corresponding to the cropped rasterized image 828 to verify the prediction of the ground truth digital tool parameters. The results of the ground truth comparison are used to guide learning of the tool parameter estimation network 218 (e.g., by adjusting one or more internal weights) until training is complete (e.g., when a training sample mask 832 is provided as input, until the tool parameter estimation network 218 predicts (multiple) parameters 836 and (multiple) corresponding parameter configurations 838 identified in the cropped labeled vector image 826 corresponding to the training sample mask 832).

[0097] The trained tool region detection network 204 and the trained tool parameter estimation network 218 may then be implemented by the visual lens system 104 to automatically generate the interactive image 120 from the rasterized image data 106 according to the techniques described herein.

[0098] Having considered example systems and techniques for generating interactive images from raster image data and training a tool region detection network and a tool parameter estimation network used therein, consider now an example procedure to illustrate various aspects of the techniques described herein.

[0099] Example Process

[0100] The following discussion describes techniques that are configured to be implemented using the previously described systems and devices. Various aspects of each of the processes are configured to be implemented in hardware, firmware, software, or a combination thereof. The processes are shown as a collection of blocks that specify operations performed by one or more devices and are not necessarily limited to the order shown for performing the operations by the respective blocks. In portions of the following discussion, reference is made to Figures 1 to 8B .

[0101] Figure 9A process 900 is depicted in an example implementation of generating a trained tool region detection network (block 902) according to the techniques described herein that can be used by the visual lens system 104 to generate the interactive image 120 from the image data 106. According to one or more implementations, the process 900 is performed by the training system 802 to generate the tool region detection network 204. To generate the trained tool region detection network, training data is generated (block 904) by first obtaining vector image data comprising a plurality of vector images (block 906). For example, the training system 802 obtains the vector image data 804.

[0102] Each vector image in the plurality of vector images is then labeled with information identifying a digital tool and one or more parameters for stylizing at least one geometric element in the vector image (block 908). For example, the labeling module 806 generates labeled vector images 808. Optionally, as part of generating training data (block 904), the enhancement module 810 generates enhanced vector image data 812 from the labeled vector images 808. According to some implementations, the training system 802 generates training data without generating enhanced vector image data 812.

[0103] A raster image is then generated for each vector image (block 910). For example, the training sample module 814 generates a rasterized image for each labeled vector image 818, where the labeled vector image 818 represents the labeled vector image 808 or the vector image of the enhanced vector image data 812. The rasterized image 820 and the labeled vector image 818 are associated as training samples 816. In some implementations, as part of generating training data (block 904), the cropping module 822 generates cropped training samples 824 for one or more training samples 816 by generating a cropped labeled vector image 826 from the labeled vector image 818 and a cropped rasterized image 828 from the rasterized image 820. The generation of the cropped training samples 824 is optionally performed and omitted in some implementations (e.g., in implementations where the stylized geometric elements of the vector image data 804 from which the training samples 816 are generated are centered around the apertures of the labeled vector image 818 and the rasterized image 820).

[0104] A mask is generated from each raster image, where the mask represents the area where the digital tool of the corresponding vector image is applied (block 912). For example, the mask module 830 generates a training sample mask 832 from each of the rasterized image 820 and the cropped rasterized image 828 generated by the training system 820. The segmentation network is then trained to predict whether a digital tool is applied to each pixel of the mask using the corresponding labeled vector image as ground truth data (block 914). For example, the network generation module 840 trains the segmentation network 842 using the training sample mask 832 and the training sample 816 or the cropped training sample 824 from which each training sample mask 832 was generated. Training continues until the segmentation network 842 is configured to output a mask for each of the multiple different digital tools for which the segmentation network 842 was trained, which accurately indicates whether the corresponding digital tool is applied to each pixel of the input raster image. The trained segmentation network is then output as a trained tool region detection network (block 916). The network generation module 840 , for example, outputs the tool region detection network 204 for use by the visual lens system 104 .

[0105] Figure 10 A process 1000 is depicted in an example implementation of generating a trained tool parameter estimation network (block 1002) according to the techniques described herein that can be used by the visual lens system 104 to generate the interactive image 120 from the image data 106. As part of training the tool parameter estimation network, training data is generated (block 904), such as with respect to Figure 9 Descriptive.

[0106] For each digital tool identified in the plurality of vector images labeled in the training data, one or more parameters configurable to control the digital tool are determined (block 1004). For example, the parameter module 834 determines, for each labeled vector image 808 included in the training data, at least one parameter 836 associated with the digital tool identified by the labeled vector image 808. For each of the one or more parameters, a set of discrete parameter configurations in which the parameter is controlled is determined (block 1006). For example, the parameter module 834 determines a plurality of possible parameter configurations 838 for each parameter 836 associated with the digital tool identified by the labeling module 806.

[0107] From the training data, training input pairs are generated, each of which includes a rasterized image and a mask identifying at least one region of the rasterized image where the digital tool is used (block 1008). For example, the network generation module 840 generates a training pair including a training sample mask 832 and a corresponding rasterized image 820 of the training sample 816 from which the training sample mask 832 was generated. The classification network is then trained to predict, for each pixel in the rasterized image, a parameter configuration for each of one or more parameters of the digital tool used to generate the rasterized image (block 1010).

[0108] For example, the network generation module 840 concatenates the training pairs as input and causes the classification network 844 to generate a probability distribution of possible parameter configurations 838 for each digital tool parameter 836 to be used to achieve the visual appearance of each pixel of the rasterized image 820 in the training pairs. Training continues until the classification network 844 is configured to accurately predict the correct parameter configuration 838 for each parameter 836 of the digital tool used to generate the vector image data 804 from which the training samples 816 were generated as the most likely confidence value. The trained classification network is then output as a trained tool parameter estimation network (block 1012). For example, the network generation module 840 outputs the tool parameter estimation network 218 for use by the visual lens system 104 in generating the interactive image 120 given the raster image data 106 as input.

[0109] Figure 11 A process 1100 is depicted in an example implementation for generating an interactive image from raster image data according to the techniques described herein. Raster image data is received (block 1102). For example, visual lens system 104 receives image data 106 representing a raster image. An interactive image is then generated from the raster image data, the raster image data identifying at least one digital tool and one or more parameters of the digital tool that can be used to achieve the visual appearance of the raster image data (block 1104). As part of generating the interactive image, a plurality of tool masks are generated, wherein each tool mask indicates whether a corresponding digital tool can be used to achieve the visual appearance of the raster image data (block 1106). Tool detection module 202 processes image data 106, for example using tool region detection network 204, to generate a plurality of tool masks 208 for at least one image region 206 of image data 106.

[0110] Digital tool parameters that can be used to achieve the visual appearance of the raster image data are then estimated based on the plurality of tool masks (block 1108). The fill module 210, for example, generates a filled tool mask 214 from each tool mask in the tool masks 208, indicating that the corresponding digital tool can be used to achieve the visual appearance of the image region 206 corresponding to the tool mask 208. Additionally, the fill module 210 generates a filled image region 212 for each filled tool mask 214 from the corresponding image region 206 of the image data 106. The filled image region 212 and the corresponding filled tool mask 214 are concatenated by the parameter estimation module 216 and provided as input to the tool parameter estimation network 218, which causes the tool parameter estimation network 218 to generate parameters 112. As described herein, by training a discrete set of one or more parameters 836 and parameter configurations 838 for each digital tool in the plurality of digital tools, the parameters 112 represent the most likely parameter configuration 838 that can be used by each digital tool parameter 836 to achieve the visual appearance of each image region 206 pixel in the image data 106.

[0111] An image tool description for the raster image data is then generated based on the estimated digital tool parameters (block 1110). For example, the tool description module 220 generates an image tool description 110 that specifies a digital tool and its associated parameters 112 that can be used to achieve the visual appearance of one or more image regions 206 of the image data 106. An interactive image is output (block 1112). For example, the drawing module 222 outputs the image data 106 with the incorporated image tool description 110 as an interactive image 120.

[0112] Input is received at the interactive image (block 1114). For example, input selecting selectable user interface component 122 of interactive image 120 is received by computing device 102, as displayed in user interface 116 of visual lens system 104. In response to receiving the input at the interactive image, an image tool description is displayed as part of the interactive image (block 1116). For example, in response to receiving the input selecting selectable user interface component 122, visual lens system 104 causes display of Figure 1 The illustrated image tool depicts 110. As another example, in response to receiving input at the image area 206 of the interactive image 120 indicated by the cursor 602, the visual lens system 104 causes the display of Figure 6 The illustrated image tool depicts 110 .

[0113] According to one or more implementations, an action associated with the image tool description is performed (block 1118). For example, the visual lens system 104 causes the display of information about the digital tool associated with the image tool description 110 (e.g., in response to the Figure 6Alternatively or additionally, the visual lens system 104 copies the parameters 112 of the image tool description 110 (e.g., in response to receiving input at the control 604 of the illustrated image tool description 110). Figure 6 Alternatively or additionally, the visual lens system 104 applies the parameters 112 of the image tool description 110 to different image data (e.g., in response to receiving input at controls 606 of the illustrated image tool description 110). Figure 7 The image tool description 110 receives input for applying associated parameters 112 to image data 704, as indicated by path 702, resulting in the visual lens system 104 generating image data 706).

[0114] Having described example processes in terms of one or more implementations, consider now example systems and devices that implement the various techniques described herein.

[0115] Example systems and devices

[0116] Figure 12 An example system is illustrated generally at 1200, including an example computing device 1202 that represents one or more computing systems and / or devices that implement the various techniques described herein. This is illustrated by including visual lens system 104. For example, computing device 1202 is configured as a server of a service provider, a device associated with a client (e.g., a client device), a system on a chip, and / or any other suitable computing device or computing system.

[0117] The illustrated example computing device 1202 includes a processing system 1204, one or more computer-readable media 1206, and one or more I / O interfaces 1208 communicatively coupled to each other. Although not shown, the computing device 1202 is also configured to include a system bus or other data and command transmission system that couples the various components to each other. The system bus includes any one or combination of different bus structures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus utilizing any one of a variety of bus architectures. Various other examples are also contemplated, such as control and data lines.

[0118] Processing system 1204 represents functionality that performs one or more operations using hardware. Thus, processing system 1204 is illustrated as including hardware elements 1210 that can be configured as processors, functional blocks, and the like. For example, hardware elements 1210 are implemented in hardware as application-specific integrated circuits or other logic devices formed using one or more semiconductors. Hardware elements 1210 are not limited by the materials from which they are formed or the processing mechanisms employed therein. For example, a processor may alternatively or additionally include (a plurality of) semiconductors and / or transistors (e.g., electronic integrated circuits (ICs)). In this context, processor-executable instructions are electronically executable instructions.

[0119] Computer-readable storage media 1206 is illustrated as including memory / storage 1212. Memory / storage 1212 represents memory / storage capacity associated with one or more computer-readable media. Memory / storage 1212 represents volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical disks, magnetic disks, etc.). Memory / storage 1212 is configured to include fixed media (e.g., RAM, ROM, fixed hard drives, etc.) and removable media (e.g., flash memory, removable hard drives, optical disks, etc.). In some implementations, computer-readable media 1206 is configured in various other ways, as further described below.

[0120] Input / output interface(s) 1208 represent functionality that allows a user to enter commands and information into computing device 1202 and also allows information to be presented to the user and / or other components or devices using various input / output devices. Examples of input devices include a keyboard, a cursor control device (e.g., a mouse), a microphone, a scanner, touch functionality (e.g., a capacitive sensor or other sensor configured to detect physical touch), a camera (e.g., a device configured to use visible or invisible wavelengths such as infrared frequencies to recognize movement as gestures that do not involve touch), and the like. Examples of output devices include a display device (e.g., a monitor or projector), speakers, a printer, a network card, a tactile response device, and the like. Thus, computing device 1202 represents various hardware configurations, described further below, to support user interaction.

[0121] Various techniques are described herein in the general context of software, hardware elements, or program modules. Typically, such modules include routines, programs, objects, elements, components, data structures, etc. that perform specific tasks or implement specific abstract data types. As used herein, the terms "module," "functionality," and "component" generally refer to software, firmware, hardware, or a combination thereof. A feature of the techniques described herein is that they are platform-independent, meaning that these techniques are configured for implementation on a variety of commercial computing platforms having a variety of processors.

[0122] The implementation of the described modules and techniques is stored on or transmitted on some form of computer-readable media. Computer-readable media includes various media accessible by computing device 1202. By way of example and not limitation, computer-readable media includes "computer-readable storage media" and "computer-readable signal media."

[0123] "Computer-readable storage media" refers to media and / or devices that achieve persistent and / or non-transient storage of information, as contrasted only to signal transmissions, carrier waves, or signals themselves. Thus, computer-readable storage media refers to non-signal-bearing media. Computer-readable storage media include hardware such as volatile and non-volatile, removable or non-removable media and / or storage devices, which are implemented in methods or techniques suitable for storing information such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data. Examples of computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage devices, hard disks, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or other storage devices, tangible media, or articles of manufacture suitable for storing desired information for access by a computer.

[0124] "Computer-readable signal media" refers to signal-bearing media that is configured to send instructions to the hardware of the computing device 1202, such as via a network. Signal media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, data signal, or other transport mechanism. Signal media also includes any information delivery media. The term "modulated data signal" means a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media includes wired media, such as a wired network or direct-wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media.

[0125] As previously described, hardware elements 1210 and computer-readable media 1206 represent modules, programmable device logic, and / or fixed device logic implemented in the form of hardware that is employed in some embodiments to implement at least some aspects of the technology described herein (such as executing one or more instructions). In some implementations, the hardware includes integrated circuits or systems on a chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and other implemented components in silicon or other hardware. In this context, the hardware operates as a processing device that performs program tasks defined by hardware-implemented instructions and / or logic, as well as hardware for storing instructions for execution (e.g., the computer-readable storage media described previously).

[0126] The foregoing combinations are used to implement the various technologies described herein. Thus, software, hardware, or executable modules are implemented as one or more instructions and / or logic on some form of computer-readable storage medium and / or implemented by one or more hardware elements 1210. Computing device 1202 is configured to implement specific instructions and / or functions corresponding to software and / or hardware modules. Thus, the implementation of a module as software executable by computing device 1202 is at least partially implemented in hardware, such as by using computer-readable storage medium and / or hardware elements 1210 of processing system 1204. Instructions and / or functions are executable / operable by one or more articles of manufacture (e.g., one or more computing devices 1202 and / or processing system 1204) to implement the technologies, modules, and examples described herein.

[0127] The technology described herein is supported by various configurations of computing device 1202 and is not limited to the specific examples of the technology described herein. This functionality is also configured to be implemented in whole or in part using a distributed system, such as on the "cloud" 1214 via the platform 1216 described below.

[0128] Cloud 1214 includes and / or represents a platform 1216 for resources 1218. Platform 1216 abstracts the underlying functionality of the hardware (e.g., servers) and software resources of cloud 1214. Resources 1218 include applications and / or data utilized when computer processing is executed on servers remote from computing device 1202. Resources 1218 also include services provided over the Internet and / or over a subscriber network such as a cellular or Wi-Fi network.

[0129] The platform 1216 is configured to abstract resources and functionality to connect the computing device 1202 with other computing devices. The platform 1216 is also configured to abstract the scaling of resources to provide a corresponding level of scaling for the demand experienced by the resources 1218 implemented via the platform 1216. Thus, in interconnected device embodiments, the implementation of the functionality described herein is configured to be distributed throughout the system 1200. For example, in some configurations, functionality is implemented partially on the computing device 1202 and via the platform 1216 that abstracts the functionality of the cloud 1214.

[0130] in conclusion

[0131] Although the invention has been described in language specific to structural features and / or methodological acts, the invention defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claimed invention.

Claims

1. A method implemented by a computing device in a digital medium automated digital tool identification environment, the method comprising: receiving, by the visual lens system, image data for the rasterized image; generating, automatically and without user intervention, an interactive image by the visual lens system, the interactive image indicating an area of ​​the rasterized image having a visual appearance achievable using a digital tool, the generating comprising: identifying, by a tool detection module, the digital tool by processing the rasterized image using a segmentation network trained to output a binary mask indicating a probability that each pixel of the rasterized image was generated using the digital tool; determining, by a parameter estimation module, a parameter configuration for controlling the digital tool to achieve the visual appearance; creating, by a tool description module, an image tool description based on the digital tool and the parameter configuration; and incorporating the image tool description into the image data; and The interactive image is output by the visual lens system at a display of a computing device, and a display of the image tool description is output in response to detecting an input at the area.

2. The method of claim 1 , wherein the segmentation network is further trained to output a binary mask for each of a plurality of different digital tools, the binary mask indicating a probability that each pixel of the rasterized image was generated using a corresponding digital tool of the plurality of different digital tools.

3. The method of claim 2, wherein identifying the digital tool comprises: The binary mask corresponding to the digital tool and the region of the rasterized image are concatenated, and a probability distribution of possible parameter configurations for the digital tool is generated by inputting the result of the concatenation into a classification network, the probability distribution indicating whether each of the possible parameter configurations can be used to achieve the visual appearance.

4. The method of claim 3 , wherein concatenating the binary mask corresponding to the digital tool and the region of the rasterized image comprises: A padded instance of the binary mask is generated at an input data size for the classification network, wherein a portion of the binary mask corresponding to the region of the rasterized image is centered at an aperture of the padded instance of the binary mask.

5. The method of claim 1 , wherein outputting the interactive image comprises: An indication is displayed that visually distinguishes the region of the rasterized image from at least one other region of the rasterized image.

6. The method of claim 1 , wherein the digital tool comprises a plurality of different parameters, each of the plurality of different parameters being configurable to control the digital tool, and determining the parameter configuration for controlling the digital tool comprises: A parameter configuration is determined for each of the plurality of different parameters.

7. The method of claim 1 , wherein the image tool description includes a control selectable to display information describing the digital tool, the method further comprising: Display of the information describing the digital tool is caused in response to detecting an input at the control.

8. The method of claim 1 , wherein the image tool description includes controls selectable to copy the parameter configuration for controlling the digital tool, the method further comprising: The parameter configuration is stored in a memory of the computing device in response to detecting an input at the control.

9. The method of claim 1 , wherein the image tool description comprises at least one of a textual description identifying a name of the digital tool and describing the parameter configuration or a visual depiction of the visual appearance achievable using the parameter configuration for the digital tool.

10. The method of claim 1 , wherein the interactive image comprises a plurality of different image regions having visual appearances, each of the visual appearances being achievable using a digital tool, wherein generating the interactive image comprises: An image tool description is generated for each image region of the plurality of different image regions.

11. The method according to claim 1 , further comprising: Vector image data is output at the display of the computing device and the vector image data is modified using the digital tool and the parameter configuration of the digital tool.

12. A system for automated digital tool identification in a digital medium environment, comprising: a labeling module implemented at least in part in hardware of the computing device to receive a plurality of vector images and generate, for each vector image, a labeled vector image specifying a digital tool used in the vector image; a training sample module implemented at least in part in hardware of the computing device to generate a rasterized image from each of the vector images; a mask module implemented at least in part in hardware of the computing device to generate a training sample mask for each rasterized image; a parameter module implemented at least in part in hardware of the computing device to identify a plurality of possible parameter configurations for each digital tool identified in the labeled vector image; a network generation module implemented at least in part in hardware of the computing device to generate a trained segmentation network using the labeled vector image, the rasterized image, and the training sample mask, and to generate a trained classification network using the training sample mask and the plurality of possible parameter configurations; as well as A rendering module is implemented at least in part in hardware of the computing device to generate an interactive image that automatically identifies image tools that can be used to achieve a visual appearance of at least one region in raster image data using the trained segmentation network and the trained classification network. 13 . The system of claim 12 , wherein the labeled vector image further specifies a parameter configuration for the digital tool, the parameter configuration of the digital tool being used to stylize geometric elements in the vector image. 14 . The system of claim 13 , wherein the training sample mask for each rasterized image indicates a region in the rasterized image that depicts the stylized geometric element in a corresponding one of the vector images.

15. The system of claim 12, further comprising: An enhancement module is implemented at least in part in hardware of the computing device to generate at least one enhanced vector image from the labeled vector image, wherein the training sample module is configured to generate a rasterized image from the at least one enhanced vector image.

16. The system of claim 15, wherein the enhancement module is configured to generate the at least one enhanced vector image by at least one of: changing the position of the stylized geometric element in the labeled vector image; changing the size of the stylized geometric element in the labeled vector image; changing the color of the stylized geometric element in the labeled vector image; modifying a parameter configuration of a digital tool applied to the stylized geometric element in the labeled vector image; adjusting a hierarchical placement of the stylized geometric elements in the labeled vector image; or Stylized geometric elements from a different labeled vector image are inserted into the labeled vector image.

17. The system of claim 12, wherein the interactive image visually distinguishes the at least one region of the raster image data from a remainder of the raster image data.

18. A system according to claim 12, wherein the interactive image is configured to display an image tool description in response to input at at least one area of ​​the raster image data, the image tool description providing an indication of the image tool and a parameter configuration for the image tool that can be used to achieve the visual appearance.

19. The system of claim 18, wherein the drawing module is configured to apply the parameter configuration for the image tool to vector image data displayed by the computing device in response to receiving input at the image tool description.

20. A system for automated digital tool identification in a digital medium environment, comprising: at least one processor; as well as A computer-readable storage medium storing instructions executable by the at least one processor, the instructions for performing operations comprising: receiving image data for a rasterized image; Automatically and without user intervention, generating an interactive image indicating an area of ​​the rasterized image having a visual appearance achievable using a digital tool, the generating comprising: identifying the digital tool by processing the rasterized image using a segmentation network, wherein the segmentation network is trained to output a binary mask indicating a probability that each pixel of the rasterized image was generated using the digital tool; determining a parameter configuration for controlling the digital tool to achieve the visual appearance; creating an image tool description based on the digital tool and the parameter configuration; and incorporating the image tool description into the image data; and The interactive image is output at a display device, and a display of the image tool description is output in response to detecting an input at the area.

Citation Information

Patent Citations

  • Selective editing of images using editing tools with persistent tool settings

    US20170132768A1