Visual content modification based on generative models in a computing system
By receiving design goals and natural language prompts from visual assets, a candidate set of visual assets is generated, allowing users to perform forward and backward interactions. This solves the linear interaction limitation of generative AI models in visual content generation, achieving more flexible and accurate generation results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FEGMA CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-07-31
AI Technical Summary
Existing generative AI models cannot effectively support forward and backward iterative interactions when generating visual content, making it difficult for users to flexibly choose the context of the generated results and failing to meet the needs of applications such as graphic design.
By receiving design goals and natural language prompts for visual assets, a generative artificial intelligence model is used to generate a set of candidate visual assets. Users can interact forward and backward, select and reselect the context of the generated results, and modify the visual assets using a control panel.
It enables non-linear interaction of generative artificial intelligence models, improves the flexibility and accuracy of generating visual content, and meets the complex interaction needs of applications such as graphic design.
Smart Images

Figure CN122492844A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure generally relate to graphic design tools, and more specifically, to techniques for generating visual content using generative artificial intelligence models. Background Technology
[0002] Generative AI models can be used to generate various types of content, such as text and images. To generate content, generative AI models typically receive input prompts specifying parameters for the content to be generated, and then output content based on these parameters. These generative AI models can take text prompts and output responses in one or more data modalities, depending on the application for which they were trained. Language models can generate text responses to input prompts, while diffusion models can generate image outputs based on input prompts.
[0003] Typically, interaction with a generative AI model is a linear process, where the model responds to a first textual input cue to generate an output and receives a second input cue related to the generated output. Within this linear process, the output of the generative AI model can be iteratively refined until the desired result is achieved. That is, the content generated in the first round of reasoning is often used as context for the second round of reasoning, which iterates based on the content generated in the first round (e.g., adding more detail to the content generated in the first round, generating content describing the concepts contained within the content generated in the first round of reasoning, etc.).
[0004] For some applications, successive text prompts (referencing content generated by the generative AI model in previous inference rounds) may be sufficient to prompt the model to produce the desired output. However, for other applications, text prompts inputted linearly to the generative AI model may not correspond to how users perform various tasks. For example, the linear nature of generative AI model operation may not provide the appropriate infrastructure to allow content generation using forward iteration and backtracking to previously generated content. In the context of graphic design, where the user-executed design process involves creating visual content (e.g., user interface design) using forward iteration and backtracking to previously generated designs, a generative AI model operating under linear prompt patterns may not be able to generate visual content in a natural and intuitive way.
[0005] As explained above, more effective techniques are needed to generate visual content using generative user interface models. Summary of the Invention
[0006] One embodiment of this disclosure provides a technique for generating candidate visual assets using a generative artificial intelligence model and selecting the visual assets on which the generated candidate visual assets are based. An example method typically includes: receiving an instruction for a visual asset to be modified and an input natural language prompt specifying a design goal to be applied to the indicated visual asset. The design goal is translated into multiple transformations to be applied to the indicated visual asset. Based on applying the multiple transformations to the indicated visual asset, a set of candidate visual assets is generated, and the set of candidate visual assets is output for selection in a user interface.
[0007] One embodiment of this disclosure provides a technique for refining visual assets using a generative artificial intelligence model. An example method typically includes: receiving an instruction to modify a visual asset and an input natural language prompt specifying a design goal to be applied to the indicated visual asset. Based on the specified design goal, visual components in the indicated visual asset and attributes of the identified visual components to be modified are identified. One or more control panels are populated with one or more controls for modifying the identified visual components to achieve the specified design goal. Input is received in at least one of the one or more control panels, and the indicated visual asset is modified based on the received input.
[0008] Compared to existing technologies, a key advantage of the technology disclosed in this paper is that it allows users of generative AI models to interact with these models in a nondeterministic and non-linear manner. Generative AI models can allow forward and backward interaction with the results generated by these models. The ability to generate content through forward and backward interaction with the results generated by generative AI models enhances their content generation capabilities because it allows for flexible selection of the context upon which the results of the generative AI model's inference rounds are based. For example, a generative AI model can generate content based on the context of the results of any previous inference round performed by the model, and is not limited to the context of the results of immediately preceding inference rounds. In the context of graphic design applications, the proposed technology allows users to interact with generative AI models using an interaction paradigm that replicates the forward and backward progression of generating visual content. Furthermore, these generative AI models can generate multiple candidate visual assets in each inference round, allowing for the selection (and reselection) of different visual assets as a starting point for generating subsequent sets of candidate visual assets. Attached Figure Description
[0009] To gain a more detailed understanding of the features of the various embodiments described above, reference can be made to the accompanying drawings for a more specific description of the inventive concept briefly outlined above. However, it should be noted that the drawings illustrate only typical embodiments of the inventive concept and should not be construed as limiting its scope in any way, and that other equally effective embodiments exist.
[0010] Figure 1 The illustration depicts a computer system configured to implement one or more aspects of various embodiments of the present disclosure.
[0011] Figure 2 The illustration shows an initial visual asset and input prompts for specifying the design goals of the visual asset, according to some embodiments.
[0012] Figure 3 The illustration shows a set of candidate visual assets generated by a generative artificial intelligence model based on input prompts that indicate the design goals of a specified visual asset, according to some embodiments.
[0013] Figure 4 The illustration shows a generative artificial intelligence model according to some embodiments, which generates different candidate visual asset sets based on the selection of visual assets from a previous candidate visual asset set.
[0014] Figure 5 This is a flowchart of the method steps for generating visual assets based on input prompts for design goals of specified visual assets and the selection of visual assets on which the generative artificial intelligence model generates visual assets, according to some embodiments.
[0015] Figure 6 The illustration shows a control panel, according to some embodiments, for adjusting the appearance of a visual asset based on input prompts that specify the design goals of the visual asset.
[0016] Figure 7 The illustration shows changes to a visual asset generated based on adjustments to one or more parameters in a control panel used to adjust the appearance of the visual asset, according to some embodiments.
[0017] Figure 8 The illustration shows variations in visual assets generated based on adjustments to parameters for recognizing input prompts according to design goals of a specified visual asset, according to some embodiments.
[0018] Figure 9 This is a flowchart of the steps of a method for modifying visual assets based on changes in parameters identified from input prompts for design goals of specified visual assets, according to some embodiments.
[0019] Figure 10The illustration depicts a network computing system for implementing an interactive graphical application platform, according to some embodiments. Detailed Implementation
[0020] In the following description, numerous specific details are set forth in order to provide a more comprehensive understanding of the various embodiments. However, those skilled in the art will understand that the inventive concepts can be practiced even without one or more of these specific details.
[0021] Figure 1 A computing device 100 configured to implement one or more aspects of various embodiments of the present invention is illustrated. In one embodiment, the computing device 100 includes a desktop computer, laptop computer, smartphone, personal digital assistant (PDA), tablet computer, or any other type of computing device configured to receive input, process data, and optionally display images, and is suitable for practicing one or more embodiments. The computing device 100 is configured to run a generative model 122 and a graphics design engine 124 stored in memory 116.
[0022] It should be noted that the computing device described herein is merely an exemplary description, and any other technically feasible configuration falls within the scope of this disclosure. For example, multiple instances of the generative model 122 or the graphics design engine 124 may execute on a set of nodes in a distributed and / or cloud computing system to implement the functionality of the computing device 100. In another example, the generative model 122 or the graphics design engine 124 may execute in various sets of hardware, device types, or environments to adapt the generative model 122 or the graphics design engine 124 to different use cases or applications. In a third example, the generative model 122 or the graphics design engine 124 may execute on different computing devices and / or different sets of computing devices.
[0023] In one embodiment, computing device 100 includes, but is not limited to, an interconnect bus 112 connecting one or more processors 102, an input / output (I / O) device interface 104 coupled to one or more input / output (I / O) devices 108, a memory 116, a storage device 114, and a network interface 106. Processor 102 can be any suitable processor implemented as a central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), artificial intelligence (AI) accelerator, any other type of processing unit, or a combination of different processing units, such as a CPU configured to work in conjunction with a GPU. Typically, processor 102 can be any technically feasible hardware unit capable of processing data and / or executing software applications. Furthermore, in the context of this disclosure, the computing elements shown in computing device 100 can correspond to a physical computing system (e.g., a system in a data center) or can be a virtual computing instance running in a computing cloud.
[0024] I / O device 108 includes devices capable of providing input, such as a keyboard, mouse, touchscreen, microphone, etc., and devices capable of providing output, such as a display device or speaker. Furthermore, I / O device 108 may include devices capable of receiving input and providing output, such as a touchscreen, Universal Serial Bus (USB) port, etc. I / O device 108 can be configured to receive various types of input from end users of computing device 100 (e.g., designers) and provide various types of output to end users of computing device 100, such as displayed digital images, digital video, or text. In some embodiments, one or more I / O devices 108 are configured to couple computing device 100 to network 110.
[0025] Network 110 is any technically feasible type of communication network that allows computing device 100 to exchange data with external entities or devices, such as web servers or other networked computing devices. For example, network 110 may include wide area networks (WANs), local area networks (LANs), wireless (Wi-Fi) networks, and / or the Internet.
[0026] Storage device 114 includes non-volatile memory for storing applications and data, and may include fixed or removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray disc, HD-DVD, or other magnetic, optical, or solid-state storage devices. Generative model 122 and graphics design engine 124 may be stored in storage device 114 and loaded into memory 116 when executed.
[0027] Memory 116 includes random access memory (RAM) modules, flash memory cells, or any other type of memory cell or a combination thereof. Processor 102, I / O device interface 104, and network interface 106 are configured to read data from memory 116 and write data to memory 116. Memory 116 includes various software programs executable by processor 102 and application data associated with said software programs, including generative models 122 or graphical design engines 124.
[0028] An example of using a generative artificial intelligence model to generate candidate visual assets In graphic design software, for example Figure 1 The software implemented by the illustrated graphics design engine 124 can render and design visual assets, such as user interface components in a user interface being designed within the graphics design engine 124. Visual assets, as used herein, can be a set of visual components rendered in the user interface. Visual assets can be vector assets defined based on mathematical relationships between different components, or they can be raster assets defined based on absolute pixel positions. In some cases, visual assets can include one or more containers defined in vector space, into which raster assets (e.g., images) can be inserted. Visual assets can contain any number of visual components, and visual assets can be defined based on the relative positioning or other spatial relationships between the visual components within the visual asset. Visual assets can be predefined (e.g., as templates or pre-designed components in the graphics design engine 124), or they can be manually designed by the user of the graphics design engine 124. These spatial relationships between visual components in a visual asset typically refer to the spacing and / or orientation of one visual component in the visual asset relative to another visual component in the visual asset.
[0029] To accelerate the process of creating visual assets, the various embodiments described herein use generative artificial intelligence models (e.g., generative artificial intelligence model 122) to translate design goals from input natural language prompts into modifications applied to the identified visual assets. A generative artificial intelligence model generally refers to an artificial intelligence model that generates content in response to input prompts specifying the parameters on which the generative artificial intelligence model bases its content generation. Generative artificial intelligence models can generate content in any of various forms, such as text content, image content, source code that generates a user interface at runtime, or other types of content. Typically, design goals can refer to the characteristics of the visual assets rendered in the user interface (e.g., the desired visual appearance or functionality). In various examples, design goals can typically specify the desired visual appearance of the visual asset in the language that designers use to describe the appearance of the visual asset and the relationships between components within the visual asset, and / or specify the required functionality of the visual asset in the form of natural language input specifying how the visual asset should respond to various inputs or events. For example, natural language prompts might include design goals that indicate the identified visual asset should be relatively taller or shorter than its current design, have a heavier or lighter weight than its current design, and have larger or smaller spacing (or white space) between visual components than its current design, etc. In another example, natural language prompts might include design goals that indicate transformations applied to the visual asset in the event of various specified events, such as a change in the appearance of a button or other clickable interface element when clicked, or movement of the visual asset in the event of a specific event, etc.
[0030] Furthermore, as discussed in detail herein, embodiments of this disclosure can allow the generation of visual assets using forward and backward iterations of the design as an identified baseline design for subsequent iterations in generating candidate visual asset sets. By allowing forward and backward selection of visual assets as a baseline for subsequent iterations in generating candidate visual asset sets, the embodiments proposed herein can enable generative AI models to operate in a non-linear manner, similar to a user designing visual assets or generating other content. That is, unlike generative AI models that iterate solely based on the output of the immediate preceding inference round performed by the generative AI model, the embodiments proposed herein allow the use of the results of any inference round performed by the generative AI model as the basis for generating subsequent outputs of the generative AI model.
[0031] Figure 2 Example 200 illustrates an input prompt for an initial visual asset and a design goal for a specified visual asset, according to some embodiments.
[0032] As shown in the figure, to initialize the generation of candidate visual assets using a generative artificial intelligence model and input natural language prompts, visual assets 212 can be added to the user interface 210. Visual assets 212 can be selected from templates or a predefined set of visual assets, or they can be initial designs created by the user of graphic design software. Figure 2 In Example 200 shown, visual asset 212 is illustrated as a horizontally scrolling card, allowing users of the user interface to view different items presented in the container by horizontal scrolling.
[0033] Natural language prompts for modifying visual asset 212 can be input into control panel 220. As shown, control panel 220 can allow input of natural language prompts specifying the design goals of the selected visual asset 212; as shown, input prompts into control panel 220 specify that the generative AI model should generate candidate visual assets that are smaller in height than visual asset 212. In some embodiments, control panel 220 may also include additional options that the generative AI model can use to adjust how the candidate visual assets to be generated by the generative AI model are to be adjusted. For example, as shown, control panel 220 may include a slider that can be used to define the range of visual assets to be generated based on the generative AI model and the input natural language prompts. For example, the slider may indicate that the set of candidate visual assets generated based on the generative AI model and the input natural language prompts includes a balanced range of candidate visual assets, or a range of candidate visual assets that is biased towards refining the selected visual asset or completely redesigning the selected visual asset.
[0034] Figure 3 The illustration shows a set of candidate visual assets 300 generated by a generative artificial intelligence model based on input prompts for a specified design goal of a visual asset, according to some embodiments.
[0035] It can be based on, for example Figure 2 The control panel 220 shown captures input prompts and selected visual assets 212 to generate a candidate visual asset set 300. Input prompts can be natural language text prompts entered into text input fields in the control panel 220, or they can be interactive user interface elements in the control panel 220 (e.g., ...). Figure 2The input can be obtained through interaction with the slider shown, or presented in any of several modalities (voice, image data, eye tracking, motion tracking, etc.). For non-text input, input cues can be derived from user interactions with interactive user interface elements, or by extracting semantic data or intent from non-textual input, and so on. For example, to generate input cues that a generative AI model can use to generate a candidate visual asset set 300, the same or different generative AI models can transform the intent identified from user interactions with interactive user interface elements or non-textual input modal input into text cues. The candidate visual asset set can include any number of candidate visual assets 314, 324, 334 (and others not listed in the original text). Figure 3 The visual assets displayed in the image can be selected by the user as a basis for further iterations or inserted into the user interface (e.g., ...). Figure 2 (See user interface 210 shown). In some embodiments, the candidate visual asset set 300 may correspond to candidate visual assets generated in the initial inference round based on the initial selection of base visual assets and natural language prompts processed by a generative artificial intelligence model.
[0036] Typically, in order to generate candidate visual assets 314, 324, 334 (and other visual assets), generative artificial intelligence models (e.g. Figure 1The generative AI model 122 shown extracts design goals from input natural language prompts and uses the extracted design goals to generate transformations to be applied to the selected visual asset 212. Typically, the generative AI model can be a design language model trained on a design-related input corpus to translate the design goals identified in the generative AI model into a broad set of transformations to be applied to the visual asset. For example, the generative AI model can be trained to output instructions or other outputs that provide a general direction for the visual asset design. This general direction may include, for example, outputs specifying the relative amount of change in the selected visual asset 212 upon which one or more transformations may be based; parameter values specifying the difference (delta) between the values of parameters in the selected visual asset 212 and the values of corresponding parameters in candidate visual assets 314, 324, 334; effective adjustment ranges applicable to define the parameters of the selected visual asset 212 to achieve the specified design goals; and so on. The design language model can identify parameters to be modified based on the various concepts expressed by words or sets of words related to the design goal or direction in the input prompts. These concepts include size-related concepts (smaller, larger, etc.), spacing-related concepts (more or less white space, higher or lower visual density, etc.), appearance-related concepts (thicker, lighter, etc.), color-related concepts (hue, brightness, etc.), and so on.
[0037] In order to visualize candidate visual assets 314, 324, 334 (and Figure 3 (Other visual assets not shown in the diagram) The graphics design engine 124 can generate a control panel 310 containing visualization options associated with each candidate visual asset generated by a generative artificial intelligence model based on design goals extracted from natural language prompts. The control panel may include buttons or other optional user interface elements that allow the user to visualize a candidate visual asset by hovering over the button associated with it (e.g., visualizing candidate visual asset 314 by hovering over button 312, visualizing candidate visual asset 324 by hovering over button 322, visualizing candidate visual asset 334 by hovering over button 332, etc.).
[0038] Although Figure 3Candidate visual asset 314 is illustrated as the same asset as visual asset 212 on which the candidate visual asset is based; however, it should be understood that candidate visual asset 314 may be a different visual asset generated by modifying visual asset 212. Candidate visual assets 324 and 334 illustrate different visual assets generated based on different transformations derived from design goals extracted from input natural language prompts. Typically, candidate visual assets 324 and 334 may contain the same visual components as visual asset 212, but these visual components may be displayed in different arrangements according to design goals extracted from input natural language prompts. In order to... Figure 2 The visual asset 212 shown and the input natural language prompts generate candidate visual assets 324. One or more transformations generated from the design prompts can specify that the container in which the visual components in which the visual asset 212 is arranged is shorter, while keeping the width of the visual asset 212 unchanged. Meanwhile, in order to... Figure 2 The visual asset 212 shown and the input natural language prompts generate candidate visual assets 334. One or more transformations generated from the design prompts can change the width and height of the container in which the visual components of visual asset 212 are arranged, and can also change the order in which the visual components are arranged. For example, by changing the width of the container in which the visual components are arranged, candidate visual asset 334 can be generated, in which the selectable button is located below the text label, which may differ from the side-by-side arrangement of the text label and selectable button shown in candidate visual asset 324.
[0039] For visual assets defined as vector graphic components, transformations identified by generative AI models may include various changes to the visual properties of these components. These visual properties may include, for example, changes in the size and positioning of one or more visual components associated with the visual asset relative to the size and positioning of those visual components in the selected visual asset 212 upon which the candidate visual asset is based, changes in text content, changes in the color palette used to display the various components, and so on. For example, to reduce the vertical height of the candidate visual asset relative to the selected visual asset 212, the size of the container and certain components within the container may be reduced by a defined scale (e.g., reduced by 50% as shown in candidate visual asset 324; reduced by 25% as shown in candidate visual asset 334, etc.). To adjust the positioning of visual components in the visual asset, the relationship between reference points of different visual components can be adjusted. As shown, the relationship between image components and other components in candidate visual asset 324 may remain unchanged, while the relationship between components in candidate visual asset 334 may change to reflect the vertical alignment of images, text labels, and selectable buttons. It should be understood that the above is merely an example of transformations that can be identified and applied to the selected visual asset 212 to generate candidate visual assets 314, 324, 334 (and other candidate visual assets), and other (spatial or non-spatial) transformations may also be considered.
[0040] Figure 4 The illustration shows an example 400 of a generative artificial intelligence model according to certain embodiments that generates different candidate visual asset sets based on the selection of visual assets from a previous candidate visual asset set.
[0041] Interface 410 displays a control panel where a set of candidate visual assets generated in the first round of inference can be previewed and selected. In interface 410, the candidate visual assets can be presented in a selection panel 412. In some embodiments, selection panel 412 may contain multiple buttons, each corresponding to one of the candidate visual assets generated in the first round of inference. As shown, a user can select one of the candidate visual assets generated in the first round of inference (e.g., the candidate visual asset corresponding to button 414) to use as a basis for further refinement that can be performed by a generative AI model. In some embodiments, selecting (e.g., clicking, hovering, etc.) the button in selection panel 412 corresponding to a candidate visual asset allows the user to preview the appearance of the candidate visual asset in the user interface. The generation of subsequent candidate visual asset sets can be performed based on selecting the button corresponding to a candidate visual asset in selection panel 412 and executing a command to generate more variations of the selected candidate visual asset (e.g., clicking a button to request the generative AI model to generate more examples similar to the selected candidate visual asset corresponding to button 414).
[0042] Subsequently, a second round of inference is performed to generate a set of candidate visual assets presented in selection panel 422. In interface 420, the first set of candidate visual assets presented in selection panel 412 can be retained so that the user can change the selection of visual assets included in the user interface to a different visual asset than the one corresponding to button 414. As shown, selection panel 422, like selection panel 412 discussed above, can contain multiple buttons associated with the set of candidate visual assets generated during the second round of inference. Selecting a button (e.g., button 424) can be used to preview the appearance of the candidate visual assets in the user interface. As described above regarding interface 410, subsequent sets of candidate visual assets can be generated based on selecting a button in selection panel 422 and executing a command to generate more variations of the selected candidate visual asset.
[0043] Interface 430 displays a control panel where candidate visual assets generated in the third round of inference can be previewed and selected. The candidate visual assets generated in the third round of inference (displayed in selection panel 432) can be generated based on the visual asset selected in the second round of inference corresponding to button 424 (and indirectly based on the visual asset selected in the first round of inference corresponding to button 414). As shown, interface 430 may include selection panels 412 and 422 to allow the user to select visual assets from previous inference rounds as the basis for generating subsequent variants in subsequent inference rounds using a generative artificial intelligence model.
[0044] In some embodiments, generative models (e.g., Figure 1 The degree of variation in the candidate visual assets generated in the generative model 122 shown in the inference rounds can vary based on the number of previous inference rounds performed to design a particular visual asset. For example, in the first round of inference, a wide variety of candidate visual assets can be generated, each with a significantly different appearance (e.g., such as...). Figure 3 As shown, candidate visual assets 314, 324, and 334 have significantly different height and / or width parameters and different visual component arrangements. Subsequent inference rounds can be performed to generate refinements of the selected visual assets generated in previous inference rounds. For example, the candidate visual assets included in selection panel 422 may include variants generated from the visual asset corresponding to button 414 using smaller transformations than those used when generating the candidate visual assets included in selection panel 412. Similarly, the candidate visual assets included in selection panel 432 may include variants generated from the visual asset corresponding to button 424 using smaller transformations than those used when generating the candidate visual assets included in selection panel 422.
[0045] In some embodiments, generative models (e.g., Figure 1The generative model 122 shown can further optimize selected candidate visual assets based on another natural language prompt input to the generative AI model. The natural language prompt used to refine the selected candidate visual assets can include further design goal information, based on which a refinement result can be generated. In some embodiments, the design goal information specified in the natural language prompt can reference another visual asset in the user interface. When the natural language prompt references another visual asset in the user interface, the generative AI model can acquire parameters associated with the referenced visual asset and determine transformations for generating candidate visual assets based on the parameters associated with the referenced visual asset. For example, the natural language prompt can specify that a certain visual asset should be designed to be visually less prominent than the referenced visual asset. To make it visually less prominent, various transformations related to the size and weight of visual components in the visual asset can be generated. Typically, these transformations can reduce the size of various visual components in the visual asset to be smaller than the size of the referenced visual asset; reduce the visual weight of text displayed in the visual asset (e.g., change from a bolder font version to a thinner font version), and so on.
[0046] In some embodiments, generative artificial intelligence models (e.g., Figure 1 The transformations generated by the generative AI model 122 (shown in the diagram) for generating candidate visual assets can be constrained by various parameters. For example, a visual component containing text data can have a minimum height, which can be predefined, for example, based on the resolution of the display on which the user interface will be rendered, the physical size of the display, the pixel density of the display, etc. Transformations applied to these visual components can modify various parameters associated with these visual components without restriction, as long as these modifications do not affect the height of the visual component (e.g., various transformations related to positioning, font weight, font, etc., can be performed). However, height-related transformations applied to these visual components may be constrained to ensure that the height of these visual components in the candidate visual assets does not fall below a predefined minimum height.
[0047] In some embodiments, visual assets can be modified with user guidance before generating further candidate variants of the visual assets using transformations generated from design goals extracted from natural language prompts based on generative artificial intelligence models. For example, users can add, remove, or modify the appearance of various visual components in the visual assets at any time during the generation of the visual assets using natural language prompts and generative artificial intelligence models. i Modifications to candidate visual assets generated in the inference round can subsequently be used in the next round. i +1 is the basis for generating candidate visual assets in the inference round.
[0048] Figure 5 This is a flowchart based on some embodiments, illustrating example operation 500 for input prompts based on design goals of specified visual assets and for generative artificial intelligence models (e.g., Figure 1 The generative model 122 shown generates visual assets by selecting the visual assets upon which they are based. For example, operation 500 can be performed by a combination of one or more processors (e.g., ...). Figure 1 The computing system of the computing device 100 shown in the figure is executed by the processor 102.
[0049] As shown in the figure, operation 500 can begin at box 510, where the processor receives an instruction for the visual asset to be modified and an input natural language prompt specifying the design goals to be applied to the indicated visual asset. The input natural language prompt can be provided directly to the processor or derived by the processor based on other inputs to the processor. For example, the input natural language prompt can be derived based on intent or other semantic information extracted from user interactions with interactive user interface elements (e.g., sliders, buttons, etc. related to changes in specific attributes of the visual asset), non-textual modal input (e.g., audio and / or visual input), etc.
[0050] At box 520, operation 500 continues, and the processor translates the design goal into multiple transformations to be applied to the indicated visual asset.
[0051] In some embodiments, the multiple transformations include modifications to the vector definition of the indicated visual asset, which enables the creation of a visual asset that conforms to design objectives. In some embodiments, the modifications include one or more of the following: size changes, position changes, or changes in spatial relationships relative to one or more other visual components in the visual asset.
[0052] At box 530, operation 500 continues, generating a candidate visual asset set based on applying the multiple transformations to the indicated visual assets.
[0053] In some embodiments, the candidate visual asset set includes a first visual asset set generated based on applying the plurality of transformations to the indicated visual asset, and one or more other visual asset sets generated based on previous selections of visual assets to which transformations were applied in previous iterations. As described above, the magnitude of the transformations used in different iterations (e.g., inference rounds) may vary depending on the amount of change to be applied to the selected visual asset generated in the previous inference round and the number of inference rounds previously performed. Larger transformations may be applied in earlier inference rounds, or when the user requests a significant rethinking of the visual asset; while smaller transformations may be applied in later inference rounds, or when the user specifies that the next round of inference should generate more refined improvements based on the selected visual asset.
[0054] At box 540, operation 500 continues to output a set of candidate visual assets for selection in the user interface.
[0055] In some embodiments, the indication of visual assets includes selecting visual assets from a candidate set of visual assets, which comprises one or more subsets of visual assets generated based on transformations generated according to design objectives. Visual assets in a second subset of the one or more subsets may be generated based on selected visual assets from the first subset of visual assets.
[0056] In some embodiments, the indicated visual assets include visual assets contained in a first subset of visual assets, which is generated in an inference round prior to the inference round in which the second subset of visual assets is generated. Operation 500 may further include generating a candidate set of visual assets by generating a third subset of visual assets based on the indicated visual assets in the first subset of visual assets and multiple transformations. To output the candidate set of visual assets for selection in a user interface, the second subset of visual assets may be replaced with the third subset of visual assets.
[0057] In some embodiments, operation 500 further includes temporarily modifying the user interface based on a temporary selection of a visual asset from the candidate visual asset set. For example, when outputting a candidate visual asset set for selection in the user interface, each visual asset in the candidate visual asset set may be rendered as a thumbnail in a selectable user interface control (e.g., a button, option button, drop-down menu, etc.). Hovering over the user interface control associated with a candidate visual asset may cause that candidate visual asset to be rendered and displayed in the user interface until the user moves the cursor away from the user interface control associated with that candidate visual asset.
[0058] In some embodiments, operation 500 further includes receiving an instruction from a set of candidate visual assets for implementation in a user interface. Instructions for one or more modifications to be applied to the selected visual asset may be received, and a modified visual asset may be generated based on the indicated modifications and the selected visual asset. In some embodiments, the indicated modifications include replacing placeholder content in the selected visual asset with other content. In some embodiments, the indicated modifications include additional visual components to be added to the selected visual asset, which may be defined at least based on the location where these additional visual components are to be added to the selected visual asset and the size of these additional visual components.
[0059] Example of visual asset modification based on generative artificial intelligence model As mentioned earlier, design goals extracted from input natural language prompts can be used to generate transformations for visual components to be applied to visual assets. Because these transformations apply to specific properties of different visual components within the visual asset, generating transformations for these components can be used to populate one or more control panels, allowing users to make user-guided modifications to the visual asset.
[0060] Figure 6 The illustration shows a control panel, according to some embodiments, for adjusting the appearance of a visual asset based on input prompts that specify the design goals of the visual asset.
[0061] As shown in the figure, visual asset 610 (which can be generated from input natural language prompts as described above, or it can be a predefined visual asset) can be displayed in the user interface. Natural language prompts ( Figure 6 (Not shown in the image) is input into the generative artificial intelligence model to specify the design goals of visual asset 610. Figure 6 In Example 600 shown, the natural language prompt can specify the density and visual weight targets to be modified. Typically, modifications to density design targets can include changing the spacing between various visual components in the visual asset. Modifications to visual weight design targets can include changing text size and / or style, line thickness and / or style, and so on.
[0062] To allow modification of visual assets 610, the graphics design engine (e.g., Figure 1 The graphical design engine 124 shown can be based on generative models (e.g., Figure 1The generative model 122 shown extracts design goals from input natural language prompts and populates various controls in the control panel. In control panel 620, sliders 622 and 624 can be used to apply a set of transformations to visual asset 610 and / or define minimum and maximum values for parameters in control panel 630. Control panel 630 typically allows user-defined modifications to the parameters of visual components associated with the visual asset by manually entering or selecting values for these parameters.
[0063] When a visual asset is selected, the graphics design engine can populate the values of parameters associated with the visual asset components relevant to the design goals specified in the natural language prompts into the corresponding sections of control panel 630. Control panel 630 may contain the relevant parameters of the visual components of the visual asset. In example 600, the visual components displaying their parameters in control panel 630 include a container displaying other visual components of the visual asset, as well as the other visual components of the visual asset. This container may correspond to the "card" component shown in control panel 630, and the other visual components of the visual asset contained in the "card" component may include descriptions, thumbnails, labels, and buttons.
[0064] Figure 7 The illustration shows an example 700 of changes to a visual asset generated by adjusting one or more parameters in a control panel used to adjust the appearance of the visual asset, according to some embodiments.
[0065] Visual assets 710 can be generated by a graphics design engine (e.g., Figure 1 The graphic design engine 124 shown here modifies the position of slider 622 by changing the position of slider 622 to the new position shown in the control panel 720. Figure 6 The visual asset 610 shown is generated by applying changes. As shown, visual asset 710 may be a sparser version of visual asset 610. When the value of the slider changes, the generative model (e.g., Figure 1 The generative model 122 shown can identify transformations that achieve the desired design goals. For example, to generate sparser visual assets, the identified transformations could increase the size of the container associated with the visual assets and increase the spacing between different visual components contained within the container. In some embodiments, when the slider's position changes from... Figure 6 The position shown by slider 622 is changed Figure 7 When the slider 722 is in the position indicated, the graphics design engine can apply visual effects in the control panel 730 to change the parameter values associated with the visual components, thereby enabling the user to easily identify which parameters have been changed.
[0066] Figure 8Example 800 according to some embodiments is illustrated, which shows changes to a visual asset generated by adjusting parameters identified based on input prompts based on design goals of a specified visual asset.
[0067] In Example 800, natural language prompts can be input into a generative artificial intelligence model (e.g., via control panel 812) Figure 1 In the generative model 122 shown, the underlying visual assets 810 are modified. (See also: Regarding...) Figure 2 The control panel 220, as discussed, allows for the specification of natural language prompts that define design goals and the range of variations to be generated by the generative artificial intelligence model. As shown, the natural language prompts specify that the underlying visual asset 810 should be modified to make it "more interesting" and that different cards should be rotated according to their depth in a card stack.
[0068] Candidate visual asset 820 can be generated by the graphics design engine based on transformations identified from the base visual asset 810 and design goals extracted from natural language prompts from the input. (As mentioned above...) Figure 6 and Figure 7 The attributes related to the design goals extracted from the input natural language prompts can be populated into the control panel 830. In Example 800, the attributes identified from the design goals may include depth distances between cards, rotation information of different cards, range of valid rotation values, elevation data, etc. Users can accept attributes generated by transformations generated by a generative artificial intelligence model, or apply various changes to the attributes of the visual component by modifying the values associated with these attributes in the control panel 830. These modifications can be made by directly inputting new values into text boxes associated with different attributes of the visual component, inputting voice and / or image-based input to the generative artificial intelligence model, inputting text prompts to the generative artificial intelligence model, etc. In some embodiments, the attributes of the visual component can be associated with various user interface controls (e.g., sliders, steppers, etc.), and interaction with these user interface controls can modify the values associated with the attributes of the visual component.
[0069] Figure 9 The flowchart illustrates an example operation 900, based on some embodiments, of modifying a visual asset by recognizing changes in parameters from input prompts indicating a design goal for that visual asset using a generative artificial intelligence model. For example, operation 900 may be performed by a processor including one or more processors (e.g., ...). Figure 1 The computing system of the processor 102 of the computing device 100 shown is executed.
[0070] As shown in the figure, operation 900 can begin from box 910, in which the processor receives an instruction for the visual asset to be modified and input natural language prompts specifying the design goals to be applied to the indicated visual asset.
[0071] In some embodiments, the input includes slider inputs associated with design attributes of the visual asset. The slider inputs may contain multiple positions, each associated with a design goal to be applied to the indicated visual asset. For example, a slider for adjusting the density of a visual asset may contain multiple positions corresponding to either a high or low density of the visual asset. In another example, a slider for adjusting the visual weight of a visual asset may contain multiple positions corresponding to either a heavy or light weight for the text, borders, etc., of the visual asset.
[0072] In some embodiments, adjusting the slider input applies global modifications to the identified visual components. Global modifications can alter properties associated with multiple visual components linked to the visual asset.
[0073] In some embodiments, the slider input is displayed in a first control panel. Adjusting the slider input control modifies one or more properties of the identified visual component displayed in a second control panel.
[0074] In some embodiments, the range associated with the slider input may include a minimum value associated with a first set of values for the identified attributes of the identified visual component; and a maximum value associated with a second set of values for the identified attributes of the identified visual component. The first and second sets of values may be identified based on specified design objectives.
[0075] In some embodiments, the input includes adjusting one of the identified attributes associated with one of the identified visual components.
[0076] In some embodiments, the valid range of values for the identified attributes of the identified visual components is associated with the amount of modification to be applied to the definition of the visual asset.
[0077] In some embodiments, the valid value range of the identified attributes of an identified visual component is defined relative to another visual component in the user interface.
[0078] At box 920, operation 900 continues to identify visual components in the indicated visual asset and the attributes of the identified visual components to be modified based on the specified design goals.
[0079] At box 930, operation 900 continues to populate one or more control panels with one or more controls to modify the identified visual components in order to achieve the specified design goals.
[0080] At box 940, operation 900 continues to receive input from at least one of the control panels.
[0081] At box 950, operation 900 continues to modify the indicated visual asset based on the received input.
[0082] By using a generative AI model to identify design goals in natural language prompts and modify visual assets based on those goals, the proposed embodiments accelerate the process of creating visual assets. The generative AI model can generate a variety of candidate visual assets and can generate additional sets of candidate visual assets based on these. Furthermore, a history of candidate visual assets generated by the generative AI model can be maintained, allowing users to move back and forth between visual assets generated in various inference rounds performed by the generative AI model. In doing so, the proposed embodiments enable the generative AI model to perform inference in a non-linear manner, where the output of any previous inference round can be used as context for executing a new inference round. Moreover, identifying design goals and generating transformations applied to visual assets allows for the design of visual assets using both coarse-grained and fine-grained modifications to parameters associated with the visual assets.
[0083] Figure 10 The illustration depicts a network computing system for implementing an interactive application platform on a user computing device, according to some embodiments. For example... Figure 10 The network computing system shown can be implemented using one or more servers that communicate with user computing devices via one or more networks. Figure 10 The network computing system 1050 shown can, for example, correspond to Figure 1 The computing device 100 shown can be used to generate and / or modify visual content based on input prompts that specify design goals for the visual content and a generative artificial intelligence model.
[0084] In some embodiments, the network computing system 1050 performs operations to enable the interactive application platform (“IAP 1000”) to be implemented on the user computing device 10. In some embodiments, a user can implement IAP 1000 by initiating a session (e.g., by visiting a website) and receiving program resources for IAP 1000. A browser component executes these program resources to implement IAP 1000 and has the ability to receive user input and render content based on or in response to user input. As described above, the implementation of IAP 1000 is designed to enable users to create various types of content, such as interactive graphic designs, artwork, whiteboard content, program code rendering, presentations, and / or text content. As further described, IAP 1000 may include logic (“ASL 1016”) for implementing one or more application services, each implemented through IAP 1000 to provide a corresponding feature set and user experience. IAP 1000 also implements application services to share resources such as canvases, workspace files, or design element libraries. In addition, IAP 1000 allows the use of multiple application services during a given online session and / or for a specific application service.
[0085] According to some embodiments, a user of computing device 10 operates a web-based application 80 to access a website, retrieve and execute program resources therefrom to implement IAP 1000. The web-based application 80 can execute scripts, code, and / or other logic (“program components”) to implement the functionality of IAP 1000. In some embodiments, the web-based application 80 may correspond to a commercially available browser, such as Google Chrome (developed by Google Inc.) or Safari (developed by Apple Inc.). In some examples, the process of IAP 1000 may be implemented as scripts and / or other embedded code downloaded by the web-based application 80 from a website. For example, the web-based application 80 may execute code embedded in a web page to implement the process of IAP 1000. The web-based application 80 may also execute scripts to retrieve other scripts and program resources (e.g., libraries) from a website and / or other local or remote locations. For example, the web-based application 80 may execute JavaScript embedded in HTML resources (e.g., web pages built according to HTML 5.0 or other versions, provided according to standards published by the W3C or WHATWG Alliance). In some embodiments, the rendering engine 1020 may utilize graphics processing unit (GPU) acceleration logic, such as acceleration logic provided by a WebGL (Web Graphics Library) program that executes a graphics library shader language (GLSL) program running on the GPU.
[0086] IAP 1000 can be implemented as part of a network service, in which a web-based application 80 communicates with one or more remote computers (e.g., servers for the network service) to execute the processes of IAP 1000. The web-based application 80 retrieves some or all of the program resources used to implement IAP 1000 from a web site. The web-based application 80 can also access various types of datasets to provide IAP 1000. These datasets may correspond to files and design libraries (e.g., pre-designed design elements), which may be stored remotely (e.g., on a server, associated with an account) or locally. In some embodiments, the network computer system 1050 provides a shared design library that the user computing device 10 can use, along with any application services provided through IAP 1000. In this way, a user can initiate a session to execute IAP 1000 to create or edit workspace files rendered on canvas 1022 according to one of the various collaborative application services of IAP 1000.
[0087] In some embodiments, IAP 1000 includes a program interface 1002, an input interface 1018, and a rendering engine 1020. Program interface 1002 may include one or more processes that execute to access and retrieve program resources from local and / or remote sources. In one implementation, program interface 1002 may use programmatic resources (e.g., an HTML 5.0 canvas) associated with web-based application 80 to generate, for example, canvas 1022. As a supplement or variation, program interface 1002 may trigger or otherwise cause the generation of canvas 1022 using programmatic resources and datasets (e.g., canvas parameters) retrieved from a local source (e.g., storage) or a remote source (e.g., from a web service).
[0088] The program interface 1002 can also retrieve program resources, including an application framework for the canvas 1022. This application framework may contain datasets, which define or configure, for example, a set of interactive graphical tools integrated with the canvas 1022. These tools constitute the input interface 1018, enabling users to provide input to generate or update content rendered on the canvas 1022.
[0089] According to some embodiments, the input interface 1018 can be implemented as a functional layer integrated with the canvas 1022 for detecting and interpreting user input. For example, the input interface 1018 can handle user interactions with input mechanisms of the user's computing device (e.g., a pointing device, a keyboard) to detect, for example, cursor positioning / movement relative to the canvas 1022, hover input (e.g., preselected input), selection input (e.g., single or double click), shortcut keys (e.g., keyboard input), and other inputs. When processing user interactions with the pointing device, the input interface 1018 can use a reference to the canvas 1022 to identify the screen position of the user cursor as the user moves or otherwise interacts with the pointing device. Additionally, the input interface 1018 can interpret user input actions based on detected input locations (e.g., whether the input location indicates a selection of a tool, an object rendered on the canvas, or an area of the canvas), the frequency of inputs detected within a given time period (e.g., double-click), and / or the start and end positions of an input or a series of inputs (e.g., the start and end positions of a click and drag), and various other input types specified by the user through one or more input devices (e.g., right-click, screen-tap, etc.). In some embodiments, the input interface 1018 can interpret a series of inputs as a selection of a design tool (e.g., selecting a shape based on the input location), and inputs for defining attributes (e.g., size) of the selected shape. In some embodiments, the input interface 1018 can interpret continuous inputs (corresponding to continuous movement of the user's pointer device) as a selection of a tool (e.g., a shape tool) and the canvas location where the output of the selected tool should appear.
[0090] In some embodiments, IAP 1000 includes application service logic 1016 to enable multiple application services to be used during a given user session, wherein each application service provides specific functionality and / or user experience to the user. As described in some embodiments, IAP 1000 utilizes the corresponding application service logic 1016 to implement each application service to configure interface component 1018, rendering engine 1020, and / or other components of IAP 1000, thereby providing the functionality and user experience of the corresponding application service. In this way, IAP 1000 enables a user to operate multiple application services during a single online session. Furthermore, different application services can share resources, including program resources of IAP 1000, such as canvas 1022. In this way, each application service can feed content to canvas 1022 and / or utilize the functionality and content provided by canvas 1022 during a given session. Furthermore, these application services can also be implemented as alternative modes of IAP 1000, allowing the user to switch between these modes, each providing specific functionality and user experience. In some embodiments, each application service can use a public workspace file associated with the user. By default, a computing device that opens a workspace file can use a default application service to access and / or update that workspace file. Users can also switch the operating mode of the IAP 1000 to use different application services to access, use, and / or update the workspace file.
[0091] The network computing system 1050 may include a site manager 1058 for managing a website that provides a set of web resources 1055 (e.g., web pages) for use by a web-based application 80 of the user computing device 10. Web resources 1055 may include instructions, such as scripts or other logic (“ICAP instructions 1057”), which can be executed by a browser or web component of the user computing device. Web resources 1055 may also include (i) resources shared between application services, provided to the user computing device in conjunction with the user computing device's use of any application service; and (ii) application-specific resources that run on the user computing device for a specific available application service among the available application services. Web resources 1055 may also include a design element library, which is partially or fully shared between application services. This design element library allows the user to select predefined design elements for use on canvas 1022 when the user uses any application service.
[0092] In some variations, once computing device 10 accesses and downloads web resource 1055, web-based application 80 executes IAP instruction 1057 to achieve the aforementioned functionality. For example, IAP instruction 1057 can be executed by web-based application 80 to launch program interface 1002 on user computing device 10. The launch of program interface 1002 can occur simultaneously with, for example, the establishment of a web socket connection between program interface 1002 and service component 1060 of network computing system 1050.
[0093] In some embodiments, Web resource 1055 contains logic executed by web-based application 80 for initiating one or more processes of program interface 1002, thereby enabling IAP 1000 to retrieve additional program resources and datasets to implement the functionality described in the examples. For example, Web resource 1055 may embed logic (e.g., JavaScript code), including GPU-accelerated logic, in an HTML page for download to a user's computing device. Program interface 1002 may be triggered to retrieve additional program resources and datasets from, for example, web service 1052 and / or local resources of computing device 10, thereby implementing each of the multiple application services of IAP 1000. For example, certain components of IAP 1000 may be implemented via web pages that can be downloaded to computing device 10 after authentication and / or after the user performs additional actions (e.g., downloading multiple pages of a workspace associated with an account identifier). Therefore, in the above example, the network computing system 1050 can transmit the IAP instruction 1057 to the computing device 10 through a combination of network communication, including through the download activity of the web-based application 80, wherein the IAP instruction 1057 is received and executed by the web-based application 80.
[0094] Computing device 10 can use web-based application 80 to access a website of network service 1052 to download web pages or web resources. When accessing the website, web-based application 80 can automatically (e.g., via saved credentials) or manually input an account identifier to service component 1060. In some embodiments, web-based application 80 may also transmit one or more additional identifiers associated with the user identifier.
[0095] Furthermore, in some embodiments, service component 1060 may retrieve profile information 1009 from user profile storage 1066 using the user's user identifier or account identifier. Alternatively, the user's profile information 1009 may be determined and stored locally on the user's computing device 10.
[0096] Service component 1060 can also retrieve files (“Active Workspace File 1063”) linked to the active workspace of a user account or identifier from file storage 1064. Configuration file storage 1066 can also identify workspaces identified along with accounts and / or users, and file storage 1064 can store datasets that constitute the workspaces. The datasets stored in file storage 1064 may, for example, include pages of the workspace and one or more data structure representations 1061 of designs that can be rendered from the corresponding active workspace file in an editor.
[0097] As a supplement or variation, each file may be associated with metadata that identifies the application service used to create that specific file. In some embodiments, the metadata identifier is used to view, use, or otherwise update the default application service of the application service.
[0098] Furthermore, in some embodiments, service component 1060 provides a representation 1059 of a workspace associated with a user to web-based application 80, wherein the representation, for example, identifies individual files associated with the user and / or user account. The workspace representation 1059 may also identify a set of files, wherein each file comprises one or more pages, and each page comprises objects that are part of a design interface.
[0099] On user device 10, a user can view a workspace representation through web-based application 80, and the user can choose to open a file in the workspace through web-based application 80. In some embodiments, when the user selects to open a file in the active workspace files 1063, web-based application 80 launches canvas 1022. For example, IAP 100 can launch an HTML 5.0 canvas as a component of web-based application 80, and rendering engine 120 can access one or more data structure representations 1011 of the content rendered on canvas 1022.
[0100] IAP 1000 utilizes application service logic 1016 to implement multiple operating modes, each corresponding to a specific application service. As previously described, the application service logic 1016 associated with each service application can contain instructions and data for configuring IAP 1000 components to include the functionality and features of the corresponding application service. Therefore, the application service logic 1016 can, for example, configure the application framework and / or input interface 1018 to differ in form, function, and / or configuration between the various alternative modes of IAP 1000. Furthermore, the actions and interaction types that the user can perform for registering input can vary depending on the operating mode. Moreover, different operating modes can include different input or user interface features for the user to select and incorporate onto the canvas 1022. For example, when IAP 1000 runs in whiteboard service application mode, program interface 1002 can provide input functionality, allowing the user to select design elements in the form of "sticky notes," whereas in the alternative mode of the interactive graphic design service application, the "sticky note functionality" is unavailable. However, in alternative mode, users can choose from a variety of possible shapes or pre-designed objects and enter text messages to be displayed on canvas 1022.
[0101] Furthermore, the application service logic 1016 can configure the operation of the rendering engine 1020, allowing the functions and behaviors of the rendering engine 1020 to differ between different application services. In this way, the rendering engine 1020 can be used to provide different behaviors for different operating modes to match the specific service application currently active. For example, the configuration of the rendering engine 1020 can affect the appearance of the canvas 1022, the appearance of the content elements rendered on the canvas 1022 (e.g., visual attributes), the behavior or representation of user interactions (e.g., whether the user cursor or pointing device is displayed on the canvas 1022), the type or specific content being rendered, the physics engine used by the rendering engine to represent dynamic events (e.g., objects are moving), which user operations can be performed (e.g., whether the size of the selected object can be adjusted), and so on.
[0102] Furthermore, each application service can use a shared library of content elements (e.g., graphic design elements), as well as core functionalities that enable the sharing and updating of design elements between different application services provided by the platform. Additionally, workspace files created and edited through one application service can be used by other application services. Moreover, the transition between application services can be seamless—for example, a user computing device 10 can open a workspace file using a first application service (e.g., an interactive graphic design application service for UIX design) and then seamlessly switch to using a second application service (e.g., a whiteboard application service) to work on the same file without closing the workspace file. In some embodiments, each application service allows the user to update the workspace file even when it is being used by other computing devices (e.g., in a collaborative environment). In some embodiments, the user can switch modes on the IAP 1000 to switch the application service being used, with each application service using the same workspace file.
[0103] Example Terms Various aspects of this disclosure are described in the following numbered clauses.
[0104] 1. In some embodiments, a processor-implemented method includes: receiving an indication of a visual asset to be modified and an input natural language prompt specifying a design goal to be applied to the indicated visual asset; converting the design goal into a plurality of transformations to be applied to the indicated visual asset; generating a candidate visual asset set based on applying the plurality of transformations to the indicated visual asset; and outputting the candidate visual asset set for selection in a user interface.
[0105] 2: The method according to Clause 1, wherein the candidate visual asset set includes a first visual asset set generated based on applying the plurality of transformations to the indicated visual assets, and one or more other visual asset sets generated based on previous iterations of applying transformations to previous selections of visual assets.
[0106] 3: The method according to Clause 1 or 2, wherein the indication of the visual asset includes selecting the visual asset from a set of candidate visual assets, the candidate set including one or more subsets of visual assets generated based on transformations generated according to the design objective.
[0107] 4: The method according to Clause 3, wherein the visual assets in the second visual asset subset of the one or more subsets are generated based on selected visual assets in the first visual asset subset.
[0108] 5: The method according to Clause 3 or 4, wherein: the indicated visual assets include visual assets contained in a first visual asset subset, the first visual asset subset being generated in an inference round prior to the inference round in which the second visual asset subset is generated; generating the candidate visual asset set includes generating a third visual asset subset based on the indicated visual assets in the first visual asset subset and the plurality of transformations; and outputting the candidate visual asset set for selection in the user interface, including replacing the second visual asset subset with the third visual asset subset.
[0109] 6. The method according to any one of Clauses 1 to 5 further includes: temporarily modifying the user interface based on a temporary selection of visual assets in the candidate visual asset set.
[0110] 7: The method according to Clauses 1 to 6 further includes: receiving an instruction from the candidate visual asset set for selection of a visual asset to be implemented in the user interface; receiving an instruction for one or more modifications to be applied to the selected visual asset; and generating a modified visual asset based on the instruction for the one or more modifications and the selected visual asset.
[0111] 8: The method described in Clause 7, wherein the indicated one or more modifications include replacing the placeholder content in the selected visual asset with other content.
[0112] 9: The method according to Clause 7 or 8, wherein the indicated one or more modifications include an additional visual component to be added to the selected visual asset, the additional visual component being defined at least based on the position to which the additional visual component is to be added to the selected visual asset and the size of the additional visual component.
[0113] 10: The method according to any one of clauses 1 to 9, wherein the plurality of transformations includes a modification of the vector definition of the indicated visual asset, the modification causing the creation of a visual asset that conforms to the design objectives.
[0114] 11: The method according to Clause 10, wherein the modification includes one or more of the following: size change, position change, or spatial relationship change relative to one or more other visual components in the visual asset.
[0115] 12: A processor-implemented method comprising: receiving an instruction to modify a visual asset and an input natural language prompt specifying a design goal to be applied to the indicated visual asset; identifying visual components in the indicated visual asset and attributes of the identified visual components to be modified, based on the specified design goal; populating one or more control panels with one or more controls for modifying the identified visual components to achieve the specified design goal; receiving input from at least one of the one or more control panels; and modifying the indicated visual asset based on the received input.
[0116] 13: The method according to Clause 12, wherein the input includes a slider input associated with the design attributes of the visual asset.
[0117] 14: The method according to Clause 13, wherein the adjustment of the slider input will globally modify the identified visual component.
[0118] 15: The method according to Clause 13 or 14, wherein the slider input is displayed in a first control panel, and wherein adjustments to the slider input are displayed in a second control panel in one or more of the attributes of the identified visual component.
[0119] 16: The method according to any one of Clauses 13 to 15, wherein the range associated with the slider input includes a minimum value associated with a first set of values of the identified attribute of the identified visual component, and a maximum value associated with a second set of values of the identified attribute of the identified visual component, and wherein the first set of values and the second set of values are identified based on the specified design objective.
[0120] 17: The method according to any one of Clauses 12 to 16, wherein the input includes adjustments to attributes of the identified attributes associated with one of the identified visual components.
[0121] 18: The method according to any one of Clauses 12 to 17, wherein the valid range of values of the identified attribute of the identified visual component is associated with the amount of modification to be applied to the definition of the visual asset.
[0122] 19: The method according to any one of Clauses 12 to 18, wherein the valid value range of the identified attribute of the identified visual component is defined relative to another visual component in the user interface.
[0123] 20. A processing system comprising: at least one memory storing executable instructions; and one or more processors configured to execute the executable instructions to cause the processing system to perform the method of any one of clauses 1 to 19.
[0124] 21. A non-transitory computer-readable medium storing executable instructions that, when processed by one or more processors, cause the one or more processors to perform the method described in any one of clauses 1 to 19.
[0125] Any element of any claim referenced in any claim and / or any combination of any element described in this application, in any way, falls within the intended scope and protection of this invention.
[0126] The descriptions of various embodiments are for illustrative purposes only and are not intended to be exhaustive or limiting of the disclosed embodiments. Many modifications and variations will arise to those skilled in the art without departing from the scope and spirit of the described embodiments.
[0127] Various aspects of this embodiment may be embodied as a system, method, or computer program product. Therefore, various aspects of this disclosure may take the form of a completely hardware embodiment, a completely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which are generally referred to herein collectively as a “module,” a “system,” or a “computer.” Furthermore, any hardware and / or software technology, process, function, component, engine, module, or system described in this disclosure may be implemented as a circuit or a set of circuits. Additionally, various aspects of this disclosure may take the form of a computer program product embodied in one or more computer-readable media on which computer-readable program code is stored.
[0128] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include the following: an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0129] The foregoing description of various aspects of this disclosure includes flowcharts and / or block diagrams illustrating methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine. When these instructions are executed by the processor of the computer or other programmable data processing apparatus, the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams are enabled to be implemented. Such processors may include, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.
[0130] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or code fragment containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative embodiments, the functions described in the blocks may occur in a different order than that shown in the figures. For example, two blocks shown consecutively in the figures may actually execute substantially concurrently, or sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware system or a combination of dedicated hardware and computer instructions that performs the specified function or action.
[0131] While the foregoing describes embodiments of this disclosure, other and further embodiments of this disclosure may be devised without departing from its essential scope, the scope of which is defined by the following claims.
Claims
1. A processor-implemented method, comprising: Receive instructions on the visual asset to be modified and input natural language prompts specifying the design goals to be applied to the visual asset to which the instructions are applied; Translate the design goals into multiple transformations to be applied to the indicated visual assets; A candidate visual asset set is generated based on applying the multiple transformations to the indicated visual assets; as well as The candidate visual asset set is output for selection in the user interface.
2. The method according to claim 1, wherein, The candidate visual asset set includes a first visual asset set generated based on applying the plurality of transformations to the indicated visual assets, and one or more other visual asset sets generated based on previous iterations of applying transformations to previous selections of visual assets.
3. The method according to claim 1, wherein, The indication of the visual assets includes selecting the visual assets from a set of candidate visual assets, the candidate set including one or more subsets of visual assets generated based on transformations generated according to the design objectives.
4. The method according to claim 3, wherein, The visual assets in the second visual asset subset of the one or more subsets are generated based on selected visual assets in the first visual asset subset.
5. The method according to claim 3, wherein: The indicated visual assets include visual assets contained in a first subset of visual assets, which is generated in an inference round prior to the inference round in which the second subset of visual assets is generated. Generating the candidate visual asset set includes generating a third visual asset subset based on the visual assets indicated in the first visual asset subset and the multiple transformations. as well as The candidate visual asset set is output for selection in the user interface, including replacing the second visual asset subset with the third visual asset subset.
6. The method according to claim 1, further comprising: The user interface is temporarily modified based on a temporary selection of visual assets from the candidate visual asset set.
7. The method according to claim 1, further comprising: Receive an instruction for a selected visual asset from the candidate visual asset set, to be implemented in the user interface; Receive an instruction to apply one or more modifications to the selected visual asset; as well as Based on the indicated one or more modifications and the selected visual asset, a modified visual asset is generated.
8. The method according to claim 7, wherein, The indicated one or more modifications include replacing the placeholder content in the selected visual asset with other content.
9. The method according to claim 7, wherein, The indicated one or more modifications include an additional visual component to be added to the selected visual asset, the additional visual component being defined at least based on the location where the additional visual component is to be added to the selected visual asset and the size of the additional visual component.
10. The method according to claim 1, wherein, The multiple transformations include modifications to the vector definition of the indicated visual asset, which result in the creation of a visual asset that conforms to the design objectives.
11. The method according to claim 10, wherein, The modifications include one or more of the following: size change, position change, or spatial relationship change relative to one or more other visual components in the visual asset.
12. A processing system, comprising: At least one memory storing executable instructions; as well as One or more processors are configured to execute the executable instructions to enable the processing system to: Receive instructions on the visual asset to be modified and input natural language prompts specifying the design goals to be applied to the visual asset to which the instructions are applied; Translate the design goals into multiple transformations to be applied to the indicated visual assets; A candidate visual asset set is generated based on applying the multiple transformations to the indicated visual assets; as well as The candidate visual asset set is output for selection in the user interface.
13. The processing system according to claim 12, wherein, The candidate visual asset set includes a first visual asset set generated based on applying the plurality of transformations to the indicated visual assets, and one or more other visual asset sets generated based on previous iterations of applying transformations to previous selections of visual assets.
14. The processing system according to claim 12, wherein, The indication of the visual assets includes selecting the visual assets from a set of candidate visual assets, the candidate set including one or more subsets of visual assets generated based on transformations generated according to the design objectives.
15. The processing system according to claim 14, wherein, The visual assets in the second visual asset subset of the one or more subsets are generated based on selected visual assets in the first visual asset subset.
16. The processing system according to claim 14, wherein: The indicated visual assets include visual assets contained in a first subset of visual assets, which is generated in an inference round prior to the inference round in which the second subset of visual assets is generated. In order to generate the candidate visual asset set, the one or more processors are configured to cause the processing system to generate a third visual asset subset based on the visual assets indicated in the first visual asset subset and the plurality of transformations; and In order to output the candidate visual asset set for selection in the user interface, one or more processors are configured to cause the processing system to replace the second visual asset subset with the third visual asset subset.
17. The processing system according to claim 12, wherein, The one or more processors are further configured to cause the processing system to: Receive an instruction for a selected visual asset from the candidate visual asset set, to be implemented in the user interface; Receive an instruction to apply one or more modifications to the selected visual asset; as well as Based on the indicated one or more modifications and the selected visual asset, a modified visual asset is generated.
18. The processing system according to claim 17, wherein, The indicated one or more modifications include one or more of the following: Replace the placeholder content in the selected visual asset with other content, or The additional visual component to be added to the selected visual asset is defined at least based on the position where the additional visual component is to be added to the selected visual asset and the size of the additional visual component.
19. The processing system according to claim 12, wherein, The multiple transformations include modifications to the vector definition of the indicated visual asset, which result in the creation of a visual asset that conforms to the design objectives.
20. A computer-readable medium storing executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: Receive instructions on the visual asset to be modified and input natural language prompts specifying the design goals to be applied to the visual asset to which the instructions are applied; Translate the design goals into multiple transformations to be applied to the indicated visual assets; A candidate visual asset set is generated based on applying the multiple transformations to the indicated visual assets; as well as The candidate visual asset set is output for selection in the user interface.