Customizing applications that use generative models
An interface-guided diffusion model converts user text into UI-compatible images that enhance discoverability and visibility of interface elements, addressing the limitations of conventional customization for visually impaired users.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2024-06-10
- Publication Date
- 2026-06-25
AI Technical Summary
Conventional browser applications face limitations in customization, particularly for visually impaired users, as generated images can blur the boundaries of user interface elements, making them difficult to discover and use.
A generative model, specifically an interface-guided diffusion model, is used to convert user-generated text into UI-compatible output images that consider UI layout information, enhancing the discoverability and visibility of interface elements by selecting appropriate colors, sizes, and shapes that align with the application's layout.
The model generates images that improve the usability of application interfaces for all users, including those with visual impairments, by maintaining or enhancing the visibility of UI elements without interfering with their visibility.
Smart Images

Figure 2026520943000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is a continuation of U.S. Non - Provisional Application No. 18 / 334,936, filed on June 14, 2023, entitled "APPLICATION CUSTOMIZATION USING AN INTERFACE - GUIDED DIFFUSION MODEL", and claims the benefit of its priority, the disclosure of which is incorporated herein by reference in its entirety.
Background Art
[0002] Customization to a conventional browser application can be limited. For example, a user may select a specific image from an image having a predetermined format that conforms to the technical requirements related to visually impaired users for use within the user interface of a browser tab.
Summary of the Invention
[0003] This disclosure uses a generative model (e.g., an interface - guided diffusion model) to convert user - generated text into an output image (e.g., a UI - compatible output image or a UI - influencing output image) for use in an application's interface. The UI - compatible (or UI - influencing) output image can, in some examples, be an image that can be used as the background of an application's interface, and the interface includes one or more UI elements arranged according to a UI layout. The system can use the UI layout information of the interface as an input condition to generate an output image considering the UI elements, thereby improving the discoverability of the UI elements on the interface. For example, the generative model can select a specific color sequence, object size and / or shape, and / or other design elements that can improve the discoverability of the UI elements when the output image is applied to the interface.
[0004] In some embodiments, the techniques described herein relate to methods, the methods comprising generating an output image by a model in response to one or more prompts including user-generated text and input conditional data, wherein the input conditional data includes UI layout information relating to at least one user interface (UI) element included in the application interface, and the methods further comprise providing the output image to the application.
[0005] In some embodiments, the techniques described herein relate to a non-temporary computer-readable medium that, when executed by at least one processor, stores executable instructions causing at least one processor to perform an operation, wherein the operation includes generating input condition data which includes UI layout information relating to at least one user interface (UI) element included in the interface of an application, and providing one or more prompts to a model which include user-generated text and input condition data, the operation further includes receiving an output image from the model which is generated with respect to at least one UI element, and the operation further includes applying the output image to the interface.
[0006] In some embodiments, the technology described herein relates to an apparatus comprising at least one processor and a non-temporary computer-readable medium for storing executable instructions, wherein the executable instructions cause at least one processor to receive one or more prompts including user-generated text and input condition data, the input condition data including UI layout information relating to at least one user interface (UI) element included in the interface of an application, and the executable instructions further cause at least one processor to generate an output image in response to one or more prompts and to provide the output image to an application.
[0007] Details of one or more embodiments are described in the accompanying drawings and the description below. Other features will become apparent from the description and drawings, as well as from the claims. [Brief explanation of the drawing]
[0008] [Figure 1A] This invention describes a system that generates UI-compatible output images using an interface-guided diffusion model, according to one embodiment. [Figure 1B] An example of an application in one aspect is shown. [Figure 1C] An example of a constraint data generator for generating UI layout information, according to one embodiment, is shown. [Figure 1D] An example of an interface-guided diffusion model according to one embodiment is shown. [Figure 1E] An example of a UI-compatible output image generated considering the UI elements of the interface, according to one embodiment, is shown. [Figure 1F] Examples of UI-compatible output images generated considering the UI elements of the interface, according to other embodiments, are shown. [Figure 2] One embodiment of a system for generating training data for an interface-guided diffusion model is presented. [Figure 3] This demonstrates a transition effect in rendering intermediate output images until the final output image is generated, according to one embodiment. [Figure 4] This example demonstrates how to generate and render different UI-compatible output images with different interfaces, according to one embodiment. [Figure 5] This example shows how to display a UI-compatible output image along with the search results. [Figure 6] This flowchart shows an exemplary operation for generating a UI-compatible output image according to one embodiment. [Figure 7] This flowchart shows an exemplary operation for rendering a UI-compatible output image in an application, according to one embodiment. [Modes for carrying out the invention]
[0009] This disclosure relates to a system for converting user-generated text into output images (e.g., UI-compatible output images or UI-influenced output images) for use on an application interface. The application may be a browser application. The interface may be a new tab page within a browser tab. The interface may be an interface for a web application running in a browser tab. In some examples, the application is an operating system or a native application executable by an operating system. For example, several embodiments can generate wallpapers (e.g., background images) for the desktop / home screen of an operating system. The system includes a model (e.g., a generation model) (e.g., an interface-inducing model) (e.g., an interface-inducing diffusion model) capable of receiving one or more prompts containing user-generated text (e.g., a natural language description of the image to be created), and input condition data relating to one or more constraints related to image generation. The input condition data may include user interface (UI) layout information relating to UI elements on the interface. In some examples, the generation model (e.g., an interface-inducing diffusion model) is a diffusion model that generates UI-influenced output images (e.g., UI-compatible output images) based on user-generated queries and UI layout information, and the generation of the output images is influenced (or induced) by existing UI elements on the interface. In some examples, a generative model is a specific diffusion model that generates an output image (e.g., a UI-compatible output image) based on user-generated queries and UI layout information, and the generation of the output image (e.g., a UI-compatible output image) is influenced (or induced) by existing UI elements on the interface.
[0010] According to some conventional approaches, when generated images are incorporated into an application interface, the image data can interfere with the visibility of UI elements. However, a system can use the UI layout information of the interface as input to generate output images that take UI elements into account (e.g., UI-compatible output images), thereby improving the discoverability of UI elements on the interface. Furthermore, visually impaired users may find it relatively difficult to use an application interface when conventional generated ML images are used in the interface. This is because the image data can blur the boundaries of UI elements, making it more difficult to find and identify UI elements on the display. For example, generated ML images may not be designed or formatted to match the application interface, in which case the image data can interfere with the visibility of UI elements. However, generative models (e.g., interface-guided diffusion models) can not only generate output images that avoid visual interference with UI elements, but such models can also generate output images with a monochrome color scheme that takes UI elements into account (e.g., UI-compatible output images), thereby improving the visibility of UI elements for visually impaired users. In some examples, the output image (e.g., a UI-compatible output image) is generated based on a color palette designed to enhance contrast and improve the perception of the output image when viewed by users with color blindness. These and other features are further illustrated with reference to the diagrams.
[0011] Figures 1A to 1F illustrate a system 100 that uses an interface-guided diffusion model 152 to convert user-generated text 128 into a UI-compatible output image 110 for use in the interface 108 of an application 104. The interface-guided diffusion model 152 may be called a model, generative model, machine learning (ML) model, or diffusion model. The UI-compatible output image 110 is, for example, an image that can be used as the background of the application's interface 108, which includes one or more UI elements 112 arranged according to a UI layout. In some examples, the UI-compatible output image 110 may be called an output image that is generated taking into account one or more UI elements in the interface or UI-influenced output image. The UI-compatible output image 110 may maintain or even improve the visibility of the UI elements 112 when arranged according to the layout, and as a result, the overall usability of the interface 108 when displayed with the UI-compatible output image 110 may remain unaffected or even be further improved.
[0012] The system 100 described herein can overcome one or more technical problems associated with the use of generated machine learning (ML) images within an application interface 108 that includes UI elements 112 (e.g., icons, controls, input fields, etc.) positioned in fixed locations on the interface 108. According to some conventional approaches, when a generated image is incorporated into an application interface 108 (e.g., used as a background), the image data may interfere with the visibility of the UI elements 112. However, the system 100 can use UI layout information 118 on the interface 108 as input conditions to generate a UI-compatible output image 110 that takes the UI elements 112 into account, thereby improving the discoverability of the UI elements 112 on the interface 108. For example, an interface-guided diffusion model 152 may select a specific color sequence, object size and / or shape, and / or other design elements that can improve the discoverability of the UI elements 112 when the UI-compatible output image 110 is applied to the interface 108.
[0013] Furthermore, visually impaired users may find it relatively difficult to use the application's interface 108 if conventional generated ML images (even black and white images, for example) are used on the interface 108. This is because the image data can blur the boundaries of UI elements, making it more difficult to find and identify the UI elements 112 on the display 105. For example, generated ML images may not be designed or formatted to match the application's interface 108, in which case the image data may interfere with the visibility of the UI elements 112. However, the interface-guided diffusion model 152 can use the activation of a color blindness setting 124 as a constraint, thereby causing the interface-guided diffusion model 152 to generate a UI-compatible output image 110 with a monochromatic color scheme that takes the UI elements 112 into consideration, thereby improving the visibility of the UI elements 112 for visually impaired users.
[0014] The techniques described herein may enable an application 104 to generate one or more prompts 101, each containing user-generated text 128 and input condition data 116. User-generated text 128 is text provided as user input and is a textual characterization of a UI-compatible output image 110. In some examples, user-generated text 128 includes a natural language description 130 of the image to be created. Input condition data 116 may include UI layout information 118 relating to the application interface 108, display screen information 120 relating to one or more attributes of the device's display 105, and / or activation of a color vision impairment setting 124. In response to a prompt(s) 101, an interface-guided diffusion model 152 may generate a UI-compatible output image 110 that is guided (or influenced) by the shape and position of UI elements 112 on the interface 108, and in some examples may be constrained by other types of input conditions represented by the input condition data 116. The interface-guided diffusion model 152 can provide the application 104 with a UI-compatible output image 110, thereby causing the application 104 to apply the UI-compatible output image 110 to the interface 108 in order to improve the discovery of UI elements 112.
[0015] Application 104 can be any type of application executable by user device 102. In some examples, application 104 is browser application 106. Browser application 106 is a web browser configured to render browser tabs in the context of one or more browser windows. Browser tabs may display web documents (e.g., web pages, PDFs, images, videos, etc.) and / or web applications, applications such as Progressive Web Applications (PWAs), and / or content associated with extensions (e.g., web content). A web application can be an application program stored on a remote server (e.g., a web server) and delivered via browser application 106 through network 150. In some examples, a Progressive Web Application is similar to a web application but is stored (at least partially) on computing device 102 and can be used offline. Extensions add features or functionality to browser application 106. In some examples, extensions can be Hypertext Markup Language (HTML), Cascading Style Sheets (CSS), and / or JavaScript-based (in the case of browser-based extensions). In some examples, application 104 is the operating system of user device 102. In some examples, application 104 is a native application (e.g., a non-browser application) installed on the operating system of user device 102.
[0016] Application 104 may render an interface 108 on the display 105 of the user device 102. Interface 108 is a user interface that allows the user to interact with the functions of Application 104. Interface 108 may be a hypertext markup language (HTML) document. Interface 108 includes one or more UI elements 112. UI elements 112 may be UI elements positioned in a predetermined location on Interface 108. UI elements 112 may include application icons, input fields, user controls, and / or menu items, etc. Some UI elements 112 may have different sizes and / or shapes from other UI elements 112. In some examples, the shape, location, and position of UI elements 112 are defined by the application developer of Application 104. In some examples, UI elements 112 are movable by the user (for example, the user may move a UI element 112 from one location to another on Interface 108).
[0017] In some examples, as shown in Figure 1B, interface 108 is the new tab page 108a of a browser tab. For example, when a user launches a browser application 106 (or creates a new browser tab), the browser application 106 may display the new browser tab along with the new tab page 108a. In some examples, the new tab page 108a is also called the stage page or homepage. The new tab page 108a may contain several UI elements 112 positioned in place. For example, the UI elements 112 may include input fields (e.g., a search field), an application icon (which, when selected, launches the application), user controls, menu items, and / or a user profile icon, etc. However, it should be noted that interface 108 can be any type of user interface of application 104. For example, interface 108 could be an operating system interface, and / or an interface associated with a particular native application, etc.
[0018] Application 104 may also render a query interface 126 on display 105. The query interface 126 may receive user-generated text 128 to generate a UI-compatible output image 110. The user-generated text 128 may include a natural language description 130 regarding the generation of the UI-compatible output image 110. In some examples, the query interface 126 is a UI element 112 of interface 108. In some examples, the query interface 126 is overlaid on interface 108. In some examples, the query interface 126 is a different interface from interface 108. In some examples, the query interface 126 is included in the settings interface of application 104. In some examples, the query interface 126 is an interface of an application (e.g., a language model extension to browser application 106) configured to communicate with application 104 and interface-guided diffusion model 152 (and in some examples, language model 151). When a language model extension is added to browser application 106 (e.g., installed, downloaded, etc.), the language model extension adds a UI-compatible image generation function to browser application 106.
[0019] In some examples, query interface 126 includes a chat interface that displays text-based queries and responses from interface-guided diffusion model 152. In some examples, query interface 126 includes a chat interface associated with a general large language model (e.g., language model 151), and language model 151 communicates with interface-guided diffusion model 152 to obtain UI-compatible output image 110. In some examples, interface-guided diffusion model 152 is a sub-component of language model 151. Application 104 and interface-guided diffusion model 152 (and in some examples, language model 151) are configured to communicate with each other via an application programming interface (API) and / or other communication protocols such as inter-process communication (IPC) or remote procedure call (RPC).
[0020] A user may input user-generated text 128 (e.g., "Create an image of a tree with wet fallen leaves") regarding the generated image via query interface 126. In response to submission of user-generated text 128, application 104 may send prompt 101 to interface-guided diffusion model 152. Prompt 101 may include user-generated text 128. In some examples, prompt 101 includes input condition data 116 generated by application 104. In some examples, application 104 generates a first prompt that includes user-generated text 128 and a second prompt that includes input condition data 116.
[0021] The input condition data 116 may contain information about one or more input conditions (or controls) for generating an image from the interface-guided diffusion model 152. The input conditions(s) are provided to the interface-guided diffusion model 152 as input(s) that influence the generation of the UI-compatible output image 110. In some examples, the input conditions(s) include one or more task-specific input conditions(s) that are learned by the interface-guided diffusion model 152 during the training period. In some examples, the input condition data 116 includes UI layout information 118 about one or more UI elements 112 included in the interface 108 of the application 104. In response to a prompt(s) 101, the interface-guided diffusion model 152 may generate a UI-compatible output image 110 according to user-generated text 128, while being constrained according to the UI layout information 118. By using UI layout information 118 as constraints on the interface-guided diffusion model 152, the interface-guided diffusion model 152 can be made to output a UI-compatible output image 110, thereby improving the discoverability of the UI elements 112 on the interface 108. For example, the interface-guided diffusion model 152 may select a specific color sequence, the size and / or shape of an object, and / or other design elements that can improve the discoverability of the UI elements 112. The interface-guided diffusion model 152 may be provided (e.g., transmitted) to the application 104 for display within the interface 108.
[0022] Application 104 may include a constraint data generator 114 configured to generate UI layout information 118. The UI layout information 118 may include information about the interface 108 of Application 104. In some examples, the UI layout information 118 may include information indicating the size and / or location of UI elements 112 within the interface 108 (e.g., metadata, UI edge map data, and / or image data). In some examples, the UI layout information 118 may include a UI edge map 134. In some examples, the UI edge map 134 may identify a simplified representation (e.g., a box) of the UI element 112 on the interface 108 and / or include information about the position and shape of the UI element 112.
[0023] In some examples, as shown in Figure 1C, the constraint data generator 114 includes a segmentation engine 132 configured to generate UI layout information 118 from a structural description 107 of the interface 108. In some examples, the structural description 107 includes the code of the interface 108 (e.g., Hypertext Markup Language (HTML) code). In some examples, the structural description 107 includes the structure that organizes the UI elements 112 of the interface 108. In some examples, the structural description 107 includes a Document Object Model (DOM) tree. In some examples, the segmentation engine 132 may detect the size and position of the boundaries (e.g., bounding boxes) of the UI elements 112 on the interface 108. In some examples, the UI layout information 118 includes the coordinates (e.g., X coordinate, Y coordinate) and size values (e.g., length and width) of the UI elements 112.
[0024] In some examples, the constraint data generator 114 may generate display screen information 120 relating to the user device 102 or display 105 used by the interface 108. In some examples, the constraint data generator 114 may obtain a device identifier associated with the user device 102 and use the device identifier to identify one or more display screen attributes of the display 105. Display screen attributes may include resolution, color accuracy, contrast ratio, viewing angle, refresh rate, response time, brightness, size and aspect ratio, and / or panel technology (e.g., liquid crystal display (LCD), organic light-emitting diode (OLED), active-matrix organic light-emitting diode (AMOLED), etc.). In some examples, the constraint data generator 114 may generate display screen information 120 including one or more display screen attributes so that the display screen information 120 is included in the input condition data 116. The interface-guided diffusion model may use one or more of the display screen attributes to generate a UI-compatible output image 110 in order to optimize the UI-compatible output image 110 for the display 105.
[0025] In some examples, application 104 may include a color blindness setting 124, which, when activated, restricts the application interface 108's color scheme to one or more colors (e.g., black and white, a color palette accessible to people with visual impairments such as color blindness). In some examples, a visually impaired user may find it relatively difficult to use the application interface if they select their background (even a black and white image). This is because the background image can obscure the boundaries of UI elements, making it more difficult to find and identify UI elements 112 on the display 105. However, the interface-guided diffusion model 152 may use the activation of the color blindness setting 124 as a constraint, thereby causing the interface-guided diffusion model 152 to generate a UI-compatible output image 110 with a monochromatic color scheme. For example, input condition data 116 may also include the activation of the color blindness setting 124. The constraint data generator 114 may determine whether the color vision impairment setting 124 is activated, and if activated, may include information to restrict the color scheme to a monochromatic color scheme. When included in the input condition data 116, the interface-guided diffusion model 152 may generate a UI-compatible output image 110 with a monochromatic color scheme that improves the visibility of the UI element 112.
[0026] In some examples, the interface-guided spread model 152 is a text-to-image transformation ML model capable of receiving one or more prompts 101 containing user-generated text 128 and input conditional data 116. The interface-guided spread model 152 includes a text-to-image transformation machine learning (ML) model. The interface-guided spread model 152 may include one or more neural network blocks, each neural network block may include one or more layers. In some examples, the neural network blocks and layers may be used interchangeably. In some examples, the neural network blocks define a more general architecture, and the neural network blocks include multiple layers. A neural network block (or layer) may refer to a functional unit that performs a specific computation on the input data. A neural network block (or layer) may include a group of interconnected neurons (e.g., nodes) that receive inputs, apply weights to those inputs, and pass the results to an activation function to produce output values. The interface-guided spread model 152 may include a combination of neural network blocks associated with natural language processing and computer vision processing. The interface-guided spread model 152 may include a text embedding layer. The interface-guided spread network model 152 may include one or more convolutional layers. The interface-guided spread network model 152 may include a transposed convolutional layer. The interface-guided spread network model 152 may include a recursive layer. The interface-guided spread network model 152 may include a transformation layer. The interface-guided spread network model 152 may include one or more generative adversarial network (GAN) components.
[0027] In some examples, the interface-guided spread model 152 is a specially configured ML model trained to learn one or more input conditions. In some examples, the interface-guided spread model 152 includes a neural network configured to control the spread model using input condition data 116. In some examples, as shown in Figure 1D, the interface-guided spread model 152 may include a locked neural network block 153 and a trainable neural network block 155, where the weights of the locked neural network block 153 are copied and transferred to the trainable neural network block 155. The locked neural network block 153 may represent a pre-trained large-scale text-to-image transformation ML model. The trainable neural network block 155 and the locked neural network block 153 are connected to convolutional layers 154 and 158, where the weights of the convolutions grow gradually from zero to optimized parameters in a learning manner. In some examples, the convolutional layer 154 includes a 1x1 convolution. In some examples, the convolutional layer 154 includes a 1x1 convolution that is initialized with weights and biases set to zero.
[0028] The trainable neural network block 155 is trained using training input condition data (e.g., training input condition data 216 in Figure 2) to learn input conditions (e.g., limited color scheme indicated by UI layout information 118, display screen information 120, color blindness setting 124, etc.), and the locked neural network block 153 may store the weights of the neural network. At runtime, the interface-guided spread model 152 may receive the input condition data 116. In some examples, the UI layout information 118, display screen information 120, and / or color blindness setting 124 are obtained from prompts 101 received from the application 104. In some examples, the interface-guided spread model 152 may obtain one or more portions of the input condition data 116 from information stored on one or more servers associated with the application 104. In some examples, the interface-guided spread model 152 may obtain UI layout information 118 from resource addresses (e.g., resource locators, universal resource locators (URLs)) associated with the interface 108 contained in the prompt(s)101. The convolutional layer 154 may apply a convolution with learned parameters to the input condition data 116. The input condition data 116 is then combined with user-generated text 128 and input to the trainable neural network block 155. The convolutional layer 158 may apply a convolution to the output of the trainable neural network block 155, thereby generating a UI-compatible output image 110.
[0029] Referring to Figure 1E, the interface-guided diffusion model 152 may receive UI layout information 118-1 relating to a first interface as an input condition and generate a UI-compatible output image 110-1 influenced by the position and size of UI elements in the first interface. As shown in Figure 1E, the UI-compatible output image 110-1 has image data (e.g., windows) having the position and size corresponding to the UI elements in the first interface. Referring to Figure 1F, the interface-guided diffusion model 152 may receive UI layout information 118-2 relating to a second interface as an input condition and generate a UI-compatible output image 110-2 influenced by the position and size of UI elements in the second interface. The interface-guided diffusion model 152 may select a specific color sequence, object size and / or shape, and / or other design elements that can improve the discoverability of UI elements in the first and second interfaces. In some examples, the interface-guided diffusion model 152 may select colors that are compatible with the color scheme defined for interface 108 (for example, colors that are compatible with, inspired by, or based on the CSS associated with interface 108). For example, the interface-guided diffusion model 152 may use CSS (or other color palettes) to generate a UI-compatible output image 110 so that the image includes complementary colors.
[0030] In some examples, System 100 is a web page design system. For instance, a web page designer might want to replace the solid background color of a web page with a suitable background image to improve its usability and appeal. The techniques described herein allow designers to quickly generate background images that are usable and compatible with the existing UI, its layout, and components. In conventional systems, designers may choose a stock image and then have to rearrange the layout and UI components to fit the stock image designated for use as the web page background. It is clear that such conventional procedures involve incorporating the background image, requiring more work from the UI designer, and may not maintain the original web page's design standards and usability.
[0031] The user device 102 may be any type of computing device, including one or more processors 113, one or more memory devices 115, a display 105, and an operating system configured to run (or assist in running) one or more applications. In some examples, the operating system is application 104. In some examples, application 104 is an application that can be run by the operating system. In some examples, the user device 102 is a laptop computer. In some examples, the user device 102 is a desktop computer. In some examples, the user device 102 is a tablet computer. In some examples, the user device 102 is a smartphone. In some examples, the user device 102 is a wearable device. In some examples, display 105 is the display of the user device 102. In some examples, display 105 may also include one or more external monitors connected to the user device 102.
[0032] The processor(s) 113 may be formed on a substrate configured to execute one or more machine-executable instructions, or a portion of software, firmware, or a combination thereof. The processor(s) 113 may be semiconductor-based; that is, the processor may include semiconductor materials capable of performing digital logic. The memory device(s) 115 may include main memory that stores information in a format readable and / or executable by the processor(s) 113. The memory device(s) 115 may store applications 104 and / or browser applications 106 (and, in some examples, language models 151) that, when executed by the processor(s) 113, perform certain operations described herein. In some examples, the memory device(s) 115 may include non-temporary computer-readable media containing executable instructions that cause at least one processor(s) (e.g., processor(s) 113) to perform operations.
[0033] The server computer 160(or more) can be a computing device in a wide variety of device forms, such as a standard server, a group of such servers, or a rack server system. In some examples, the server computer(or more) 160 may be a single system sharing components such as a processor and memory. In some examples, the server computer(or more) 160 may be multiple systems that do not share a processor and memory. The network 150 may include the Internet and / or other types of data networks, such as a local area network (LAN), a wide area network (WAN), a cellular network, a satellite network, or another type of data network. The network 150 may also include any number of computing devices (e.g., computers, servers, routers, network switches, etc.) configured to receive and / or transmit data within the network 150. The network 150 may further include any number of wired and / or wireless connections.
[0034] The interface-guided spread model 152, and in some examples, the language model 151, are executable by one or more server computers 160. The language model 151 may be a large-scale language model configured to answer language queries on a general topic. For example, the language model 151 may be a pre-trained general-purpose neural network-based model configured to understand, summarize, generate, and predict new content on a given user-generated query. In some examples, system 100 does not include a separate language model 151. In some examples, the interface-guided spread model 152 and the language model 151 are separate models. In some examples, the interface-guided spread model 152 and the language model 151 form a single language model, but represent two different subroutines of a common language model.
[0035] A server computer (or more) 160 may include one or more processors 161 formed on a circuit board, an operating system (not shown), and one or more memory devices 163. The memory devices (or more) 163 may represent any (or more) types of memory (e.g., RAM, flash, cache, disk, tape, etc.). In some examples (not shown), the memory devices may include external storage, such as memory that is physically far from the server computer (or more) 160 but accessible from it. The processors (or more) 161 may be formed on a circuit board and configured to execute one or more machine-executable instructions, or a portion of software, firmware, or a combination thereof. The processors (or more) 161 may be semiconductor-based; that is, the processors may include semiconductor materials capable of performing digital logic. The memory devices (or more) 163 may store information in a format that can be read and / or executed by the processors (or more) 161. A memory device(s) 163 may store an interface-guided diffusion model 152 that, when executed by a processor(s) 161(s), performs certain operations described herein, and in some examples, a language model 151. In some examples, the memory device(s) 163 includes a non-temporary computer-readable medium containing executable instructions that cause at least one processor(s) 161 to perform an operation.
[0036] Figure 2 shows a system 200 configured to generate training input conditional data 216 for training an interface-guided diffusion model 252 to generate a UI-compatible output image. The interface-guided diffusion model 252 may be an example of the interface-guided diffusion model 252 and may include any of the features described herein. In some examples, the system 200 may identify a design web page 264, obtain a UI edge map 234 from the design web page 264, and generate training input conditional data 216 from the UI edge map 234. The training input conditional data 216 includes UI layout information 218 for the UI edge map 234. The UI layout information 218 may include any of the information described with reference to the UI layout information 118 in Figures 1A to 1F. In some examples, the UI layout information 218 may include information indicating the size and / or location of UI elements within the UI edge map 234.
[0037] System 200 may include a webpage identifier 262 configured to identify a design webpage 264 from multiple webpages 260 on the internet. For example, the webpage identifier 262 may use a search index to search for and identify a relevant design webpage 264 on the internet. The criteria used by the webpage identifier 262 to identify a design webpage 264 are technical, for example, identifying a webpage that has UI elements and UI layouts that follow guidelines and patterns known to be advantageous in terms of usability. Identification of whether a webpage 260 meets the criteria for a design webpage 264 may be based, for example, on the underlying stylesheet of the webpage 260. System 200 may include a training data generator 214 configured to generate training input condition data 216 from the design webpage 264. In some examples, the training data generator 214 includes a segmentation engine 232 configured to generate a UI edge map 234 from the design webpage 264 identified by the webpage identifier 262. In some examples, the segmentation engine 232 may execute an edge detection algorithm to create a UI edge map 234. In some examples, the UI edge map 234 is a visual representation of the interface, including representations of UI elements (e.g., boxes) located at various locations on the interface. In some examples, the segmentation engine 232 detects the location and size of the boundaries of UI elements from the design web page 264. In some examples, the training data generator 214 may include a visual language model 268 (e.g., a text-image language model) configured to generate caption data 270 from the UI edge map 234. In some examples, the caption data 270 may include location and size information, as well as other information about the UI edge map 234.
[0038] Figure 3 shows a transition effect 380 for displaying a UI-compatible output image 310 generated by the interface-guided diffusion model 352 on the application interface 308. The interface-guided diffusion model 352 may be an example of the interface-guided diffusion model 152 in Figures 1A to 1F and / or the interface-guided diffusion model 252 in Figure 2, and may include any of the details described with reference to those figures. In some examples, generating a UI-compatible output image 310 from the interface-guided diffusion model 352 using input condition data and user-generated text may be relatively fast, and the final output image 310-4 is created relatively quickly. In some examples, generating a UI-compatible output image 310 from the interface-guided diffusion model 352 using input condition data and user-generated text may require a relatively long processing time (e.g., 30 seconds, 1 minute, 5 minutes, etc.). In some examples, the application is configured to receive intermediate output images (e.g., intermediate output image 310-1, intermediate output image 310-2, intermediate output image 310-3) on the interface 308 before the final output image 310-4 is displayed, and to display them as transition effects 380, so that the user can see the evolutionary process of the UI-compatible output image 310.
[0039] The interface-guided diffusion model 352 may include multiple layers 370. These layers may include layers 370-1, 370-2, 370-3, and 370-4. Although four layers are shown in Figure 3, the interface-guided diffusion model 352 may include any number of layers 370. Over time, the layers 370 may successively refine the UI-compatible output image 310 until the final output image 310-4 is generated. Instead of waiting for the process to complete before applying the UI-compatible output image 310 to the interface 308, the application may render intermediate output images while the UI-compatible output image 310 is being generated.
[0040] For example, the interface-guided diffusion model 352 may receive tokens 329 representing user-generated text and input condition data. Layer 370-1 may use tokens 329 to generate an intermediate output image 310-1, which is provided as input to a subsequent layer (e.g., layer 370-2). After the intermediate output image 310-1 is generated, the interface-guided diffusion model 352 may provide the intermediate output image 310-1 to the application so that it is displayed within the interface 308. Layer 370-2 may use the intermediate output image 310-1 to generate an intermediate output image 310-2, which is provided as input to a subsequent layer (e.g., layer 370-3). The intermediate output image 310-2 may be a further improvement of the intermediate output image 310-1. After the intermediate output image 310-2 is generated, the interface-guided diffusion model 352 may provide the intermediate output image 310-2 to the application so that it is displayed within the interface 308 (for example, replacing intermediate output image 310-1).
[0041] Layer 370-3 may generate an intermediate output image 310-3 using the intermediate output image 310-2, which is then provided as input to a subsequent layer (e.g., layer 370-4). The intermediate output image 310-3 may be a further improvement of the intermediate output image 310-2. After the intermediate output image 310-3 is generated, the interface-guided diffusion model 352 may provide the intermediate output image 310-3 to the application, thereby making the intermediate output image 310-3 visible within the interface 308 (e.g., replacing the intermediate output image 310-2). Layer 370-4 may generate a final output image 310-4 using the intermediate output image 310-3. The final output image 310-4 may be a further improvement of the intermediate output image 310-3. After the final output image 310-4 is generated, the interface-guided diffusion model 352 may provide the final output image 310-4 to the application, thereby making the final output image 310-4 visible within the interface 308 (for example, replacing the intermediate output image 310-3).
[0042] Figure 4 shows an example of generating and displaying UI-compatible output images (e.g., UI-compatible output image 410a, UI-compatible output image 410b) associated with different interfaces (e.g., interface 408a, interface 408b) on display 405. For example, when interface 408a is rendered, UI-compatible output image 410a may be displayed within interface 408a, conforming to the UI elements of interface 408a. When interface 408b is rendered, UI-compatible output image 410b may be displayed within interface 408b, conforming to the UI elements of interface 408b. In some examples, UI-compatible output image 410b is different from (but related to) UI-compatible output image 410a.
[0043] An application may generate one or more prompts, each containing user-generated text 428 (for example, "Create an image of a forest with deer") and UI layout information 418a relating to UI elements contained on interface 408a. In response to receiving a prompt(s), the interface-guided diffusion model 452 may perform inference 475-1 to generate a UI-compatible output image 410a that takes into account the UI elements contained on interface 408a. The interface-guided diffusion model 452 may provide the UI-compatible output image 410a to the application displayed on interface 408a. In some examples, the UI-compatible output image 410a is a background image. In some examples, the application updates its settings to indicate that the UI-compatible output image 410a is a background image. In some examples, the UI-compatible output image 410a is a background image for a new tab page in a browser tab.
[0044] In some examples, the application may render interface 408b. In some examples, interface 408b is rendered in response to a user interaction detected on interface 408a. In some examples, interface 408b is a different interface. In some examples, interface 408b contains one or more UI elements that are different from the UI elements on interface 408a. Before interface 408b is rendered, the interface-guided diffusion model 452 may perform inference 475-2 to generate a UI-compatible output image 410b that takes into account the UI elements contained on interface 408b. For example, the interface-guided diffusion model 452 may generate a UI-compatible output image 410b using the UI-compatible output image 410a and UI layout information 418b as inputs. The UI-compatible output image 410b may be related to the UI-compatible output image 410a because the interface-guided diffusion model 452 receives the UI-compatible output image 410a as an input condition. In some examples, the interface-guided diffusion model 452 may receive a seed for a UI-compatible output image 410a and use the seed (along with UI layout information 418b) to generate the UI-compatible output image 410. The seed is a set of digits that informs the interface-guided diffusion model 452 how to generate an image (e.g., a blueprint for a work of art). The seed may guide the interface-guided diffusion model 452 so that it creates a new and unique image. Instead of generating a new image using a random seed, the interface-guided diffusion model 452 may use a seed for a UI-compatible output image 410a, which may form the base image, and the interface-guided diffusion model 452 may use that seed to create a new (but related) image.
[0045] Figure 5 shows a system 100 configured to render search results 540 and UI-compatible output images 510 within a browser tab 508 of a browser application 506. System 500 may be an example of system 100 in Figures 1A to 1F and may include any of the details described with reference to those figures. The browser tab 508 may display a query interface 526 associated with a language model 551. In some examples, the browser tab 508 displays a new tab page, and the query interface 526 is provided on the new tab page.
[0046] The language model 551 could be a large-scale language model configured to answer language queries on a general topic. For example, the language model 551 could be a pre-trained general-purpose neural network-based model configured to understand, summarize, generate, and predict new content related to a given user-generated text 528. In some examples, the language model 551 could work with a search engine 582 to identify search results 540 in response to the user-generated text 528.
[0047] A user can ask the language model 551 various different types of questions (for example, "Can you recommend some good restaurants in Milwaukee?"). The browser application 506 may generate one or more prompts 501 containing input condition data 516 and user-generated text 528. The user-generated text 528 may contain a search query (for example, "Can you recommend some good restaurants in Milwaukee?"). The browser application 506 may send the prompt(s) 501 to the language model 551. The language model 551 may communicate with the search engine 582 to identify search results 540 related to the user-generated text 528. In some examples, the language model 551 may generate a text response 561 in response to the user-generated text 528. The browser application 506 may receive the text response 561 and the search results 540 and display them in the browser tab 508.
[0048] In some examples, the language model 551 may detect an entity 529 (e.g., "Milwaukee") mentioned in user-generated text 528 and communicate with the interface-guided diffusion model 552 to generate a UI-compatible output image 510 related to the entity 529. The entity 529 could be a person, place, item, idea, concept, etc. In some examples, the language model 551 may generate a prompt 511 containing a text description generated by the language model 551 about the image to be created (e.g., "Create an image of Milwaukee," or "A restaurant district in Milwaukee," etc.). In some examples, the prompt 511 may also include input conditional data 516. In response to the prompt 511, the interface-guided diffusion model 552 may generate a UI-compatible output image 510 that takes into account UI elements within a browser tab 508. In some examples, the interface-guided diffusion model 552 may return a UI-compatible output image 510 to the language model 551, which then provides the search results 540, text responses 561, and the UI-compatible output image 510 to the browser application 506. In some examples, the UI-compatible output image 510 may be displayed as part of a browser tab 508.
[0049] In some examples, the search results 540 include selectable UI icons 541 corresponding to web documents 545. When selected, the selectable UI icon 541 displays the corresponding web document 545 within a browser tab 508. Along with each selectable UI icon 541, the search results 540 may include the title of the web document 545, a short snippet from the web document 545, and / or other information related to the web document 545. In some examples, the interface-guided diffusion model 552 may generate graphics 543 for the selectable UI icons 541. For example, with respect to a particular selectable UI icon 541, the interface-guided diffusion model 552 may receive one or more images associated with the web document 545 and create graphics 543 for the selectable UI icon 541. In some examples, the interface-guided diffusion model 552 may return the graphics 543 to a language model 551, which may provide the graphics 543 and the search results 540 to a browser application 506 for display.
[0050] Figure 6 is a flowchart 600 illustrating exemplary operation of a system that uses an interface-guided diffusion model to convert user-generated text into UI-compatible output images for use in an application interface. Flowchart 600 may illustrate how it is performed by a computer. Although flowchart 600 is described in relation to system 100 in Figures 1A–1F, flowchart 600 may be applicable to any of the embodiments described herein. While flowchart 600 in Figure 6 shows the operations in sequence, it should be recognized that this is merely an example and may include additional or alternative operations. Furthermore, the operations in Figure 6 and related operations may be performed in a different order than shown, or in a parallel or overlapping manner. Flowchart 600 may illustrate how it is performed by a computer.
[0051] The flowchart 600 described herein may overcome one or more technical problems associated with the use of generated machine learning (ML) images within the interface of an application that includes UI elements (e.g., icons, controls, input fields, etc.) positioned in fixed locations on the interface. An interface-guided diffusion model may use the UI layout information of the interface as input conditions to generate a UI-compatible output image that takes UI elements into account, thereby improving the discoverability of UI elements on the interface.
[0052] Operation 602 includes generating a UI-compatible output image 110 by an interface-guided diffusion model 152 in response to one or more prompts 101 including user-generated text 128 and input condition data 116, wherein the input condition data 116 includes user interface (UI) layout information 118 relating to at least one UI element 112 included in the interface 108 of the application 104. Operation 604 includes providing the UI-compatible output image 110 to the application 104.
[0053] In some examples, the operation includes generating a UI-compatible output image by an interface-guided diffusion model in response to one or more prompts including user-generated text and input condition data, wherein the input condition data includes user interface (UI) layout information relating to at least one UI element included in the application's interface, and the operation further includes providing the UI-compatible output image to the application. In some examples, the operation includes identifying a design web page from multiple web pages, generating training input condition data from the design web page, and training an interface-guided diffusion model based on the training input condition data. In some examples, generating training input condition data from a design web page includes generating a UI edge map from the design web page, and the training input condition data includes the UI edge map. In some examples, generating training input condition data from a design web page further includes generating caption data from at least one of the UI edge map or the design web page, wherein the caption data includes information about one or more UI elements included on the UI edge map, and the training input condition data includes the UI edge map and the caption data.
[0054] In some examples, the input condition data includes activation of a color blindness setting, and generating a UI-compatible output image includes generating a UI-compatible output image with a monochromatic color scheme using an interface-guided diffusion model. In some examples, the input condition data includes display screen information relating to one or more display attributes of a display screen, and generating a UI-compatible output image includes generating a UI-compatible output image using an interface-guided diffusion model based on the display screen information. In some examples, the operation includes generating an intermediate output image using an interface-guided diffusion model, and providing the intermediate output image for display within the interface using the interface-guided diffusion model until the final output image is generated. In some examples, the interface is a first interface, the UI layout information is first UI layout information, the UI compatible output image is a first UI compatible output image, and the operation includes: receiving second UI layout information about at least one UI element on the application's second interface by the interface-guided diffusion model; generating a second UI compatible output image by the interface-guided diffusion model based on the first UI compatible output image and the second UI layout information; and providing the application with the second UI compatible output image for display on the application's second interface by the interface-guided diffusion model. In some examples, the interface is a new tab page in a browser tab. In some examples, user-generated text includes a natural language description of an image created using the interface-guided diffusion model. In some examples, user-generated text includes a search query related to an entity, and the UI compatible output image includes image data related to the entity. In some examples, the UI layout information is obtained from a resource address provided along with the user-generated text.
[0055] Figure 7 is a flowchart 700 illustrating exemplary operation of a system that uses an interface-guided diffusion model to convert user-generated text into UI-compatible output images for use in an application interface. Flowchart 700 may illustrate how the operation is performed by a computer. Although flowchart 700 is described in relation to system 100 in Figures 1A–1F, flowchart 700 may be applicable to any of the embodiments described herein. While flowchart 700 in Figure 7 shows the operation in sequence, it should be recognized that this is merely an example and may include additional or alternative operations. Furthermore, the operations in Figure 7 and related operations may be performed in a different order than shown, or in a parallel or overlapping manner. Flowchart 700 may illustrate how the operation is performed by a computer.
[0056] The flowchart 700 described herein may overcome one or more technical problems associated with the use of generated machine learning (ML) images within the interface of an application that includes UI elements (e.g., icons, controls, input fields, etc.) located in fixed locations on the interface. An interface-guided diffusion model may use the UI layout information of the interface as input conditions to generate output images (e.g., UI-compatible output images, UI-influenced output images) that take UI elements into account (or are influenced by UI elements), thereby improving the discoverability of UI elements on the interface.
[0057] Operation 702 includes generating input condition data 116 which includes user interface (UI) layout information 118 relating to at least one UI element 112 included in the interface 108 of application 104. Operation 704 includes providing one or more prompts 101 to a model (e.g., a generative model) (e.g., an ML model) (e.g., a diffusion model) (e.g., an interface-guided diffusion model), where one or more prompts 101 include user-generated text 128 and input condition data 116. Operation 706 includes receiving an output image (e.g., a UI-compatible output image 110) from the model (e.g., an interface-guided diffusion model 152). Operation 708 includes applying the output image (e.g., a UI-compatible output image 110) to the interface 108.
[0058] In some examples, the operation includes generating input condition data containing user interface (UI) layout information relating to at least one UI element included in the application's interface, and providing one or more prompts to an interface-guided diffusion model, where the one or more prompts include user-generated text and input condition data; the operation further includes receiving a UI-compatible output image from the interface-guided diffusion model and applying the UI-compatible output image to the interface. In some examples, the operation includes generating UI layout information based on a structure description associated with the interface. In some examples, the operation includes displaying intermediate output images generated by the interface-guided diffusion model until the final output image is displayed. In some examples, the interface is a first interface, the UI layout information is first UI layout information, the UI-compatible output image is a first UI-compatible output image; the operation includes transmitting second UI layout information relating to at least one UI element on a second interface of the application, receiving a second UI-compatible output image generated by the interface-guided diffusion model using the first output image and the second UI layout information, and applying the second output image to the second interface.
[0059] Clause 1. A method comprising, by a model, generating an output image in response to one or more prompts, the prompts comprising user-generated text and input conditional data, wherein the input conditional data comprises user interface (UI) layout information relating to at least one UI element included in the interface of an application, and the method further comprises providing the output image to the application. In some examples, the output image is a UI-compatible output image. In some examples, the output image is a UI-influenced output image. In some examples, the output image is an image generated with consideration (or influence) by the at least one UI element included in the interface.
[0060] Clause 2. The method according to Clause 1, further comprising identifying a design webpage from a plurality of webpages, generating training input condition data from the design webpage, and training the model based on the training input condition data.
[0061] Clause 3. The method according to Clause 2, wherein generating the training input condition data from the design webpage includes generating a UI edge map from the design webpage, and the training input condition data includes the UI edge map.
[0062] Clause 4. The method according to Clause 3, wherein generating the training input condition data from the design web page further comprises generating caption data from at least one of the UI edge map or the design web page, wherein the caption data comprises information about one or more UI elements contained on the UI edge map, and the training input condition data comprises the UI edge map and the caption data.
[0063] Clause 5. The method according to any one of Clauses 1 to 4, wherein the input condition data includes activation of a color blindness setting, and generating the output image includes, by the model, generating the output image with a monochromatic color scheme.
[0064] Clause 6. The method according to any one of Clauses 1 to 5, wherein the input condition data includes display screen information relating to one or more display attributes of a display screen, and generating the output image includes generating the output image by the model based on the display screen information.
[0065] Clause 7. The method according to any one of Clauses 1 to 6, further comprising: generating an intermediate output image by the model; and providing the intermediate output image for display on the interface by the model until a final output image is generated.
[0066] Clause 8. The method according to any one of Clauses 1 to 7, wherein the interface is a first interface, the UI layout information is first UI layout information, the output image is a first output image, and the method further comprises the model receiving second UI layout information relating to at least one UI element on a second interface of the application, the model generating a second output image based on the first output image and the second UI layout information, and the model providing the application with the second output image for display on the second interface of the application.
[0067] Clause 9. The interface is a new tab page in a browser tab, as described in any one of Clauses 1 to 8.
[0068] Clause 10. The user-generated text includes a natural language description of an image created using the model, as described in any one of Clauses 1 to 9.
[0069] Clause 11. The method described in any one of Clauses 1 to 10, wherein the user-generated text includes a search query related to the entity, and the output image includes image data related to the entity.
[0070] Clause 12. The UI layout information is obtained from the resource address provided with the user-generated text, as described in any one of Clauses 1 to 11.
[0071] Clause 13. A non-temporary computer-readable medium that, when executed by at least one processor, stores executable instructions causing the at least one processor to perform an operation, wherein the operation includes generating input condition data including user interface (UI) layout information relating to at least one UI element included in the interface of an application, and providing one or more prompts to a model, the one or more prompts including user-generated text and the input condition data, the operation further includes receiving an output image from the model, the output image being generated with respect to the at least one UI element, and the operation further includes applying the output image to the interface.
[0072] Clause 14. The operation further includes generating the UI layout information based on the structural description associated with the interface, in the non-temporary computer-readable medium described in Clause 13.
[0073] Clause 15. The operation further includes displaying intermediate output images generated by the model until the final output image is displayed, in a non-temporary computer-readable medium as described in Clause 13 or 14.
[0074] Clause 16. A non-temporary computer-readable medium as described in any one of Clauses 13 to 15, wherein the interface is a first interface, the UI layout information is first UI layout information, the output image is a first output image, and the operation further includes transmitting second UI layout information relating to at least one UI element on a second interface of the application, receiving a second output image generated by the model using the first output image and the second UI layout information, and applying the second output image to the second interface.
[0075] Clause 17. A device comprising at least one processor and a non-temporary computer-readable medium for storing executable instructions, wherein the executable instructions include causing the at least one processor to receive one or more prompts including user-generated text and input condition data, the input condition data including user interface (UI) layout information relating to at least one UI element included in the interface of an application, and the executable instructions further include causing the processor to generate an output image in response to the one or more prompts and providing the output image to the application.
[0076] Clause 18. The apparatus according to Clause 17, wherein the executable instructions include a plurality of instructions, the plurality of instructions causing the at least one processor to: identify a design web page from a plurality of web pages; generate training input condition data from the design web page; and train a model based on the training input condition data.
[0077] Clause 19. The apparatus according to Clause 18, wherein the executable instruction includes an instruction causing the at least one processor to generate a UI edge map from a design web page, and the training input condition data includes the UI edge map.
[0078] Clause 20. The apparatus according to Clause 19, wherein the executable instruction comprises a plurality of instructions, the plurality of instructions comprising causing the at least one processor to generate caption data from the UI edge map, the caption data comprising information relating to one or more UI elements contained on the UI edge map, and the training input condition data comprising the UI edge map and the caption data.
[0079] Various embodiments of the systems and technologies described herein can be realized in digital electronic circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include embodiments in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, the at least one programmable processor may be dedicated or general-purpose and coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to them.
[0080] These computer programs (also known as programs, software, software applications, or code) contain machine instructions to a programmable processor and can be implemented in high-level procedural programming languages and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic circuits (PLDs)) used to provide machine instructions and / or data to a programmable processor, including machine-readable medium that receives machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0081] To provide user interaction, the systems and technologies described herein can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) to which the user can provide input to the computer. Other types of devices can similarly be used to provide user interaction. For example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user may be received in any form, including acoustic, verbal, or tactile input.
[0082] The systems and technologies described herein can be implemented within a computing system that includes backend components (e.g., as a data server), middleware components (e.g., an application server), or frontend components (e.g., a client computer having a graphical user interface or web browser that allows a user to interact with embodiments of the systems and technologies described herein), or in any combination of such backend, middleware, or frontend components. The components of the system can be interconnected by digital data communications (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0083] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact via a communication network. The relationship between clients and servers arises from computer programs running on each computer that have a client-server relationship with one another.
[0084] In this specification and the accompanying claims, the singular forms “a,” “an,” and “the” do not exclude plural references unless specifically indicated by the context. Furthermore, conjunctions such as “and,” “or,” and “and / or” are inclusive unless specifically indicated by the context. For example, “A and / or B” includes A alone, B alone, and A and B. Furthermore, the connecting lines or connectors shown in the various figures presented are intended to illustrate exemplary functional relationships and / or physical or logical connections between various elements. Many alternative or additional functional relationships, physical connections, or logical connections may exist in a working device. Furthermore, unless an element is specifically described as “essential” or “important,” an item or component is not essential to the practice of the embodiments disclosed herein.
[0085] While not exclusive, terms such as “approximately,” “substantially,” and “generally” are used herein to indicate that their exact value or range is not required or necessary to be specified. Where used herein, these terms have an immediate and readily apparent meaning to those skilled in the art.
[0086] Furthermore, the use of terms such as up, down, top, bottom, side, end, front, and back in this specification is used in reference to the orientation currently considered or illustrated. It should be understood that if they are considered in relation to other orientations, such terms must be modified accordingly.
[0087] Furthermore, in this specification and the accompanying claims, the singular forms "a," "an," and "the" do not exclude plural references unless specifically indicated by the context. Additionally, conjunctions such as "and," "or," and "and / or" are inclusive unless specifically indicated by the context. For example, "A and / or B" includes A alone, B alone, and A and B.
[0088] While certain exemplary methods, apparatuses, and articles are described herein, the scope of this patent is not limited to them. It should be understood that the terminology used herein is for illustrative purposes only and not intended to limit any particular aspect. Rather, this patent covers all methods, apparatuses, and articles fairly included within the claims of this patent.
Claims
1. It is a method, The method includes generating an output image by a model in response to one or more prompts including user-generated text and input condition data, wherein the input condition data includes UI layout information relating to at least one user interface (UI) element included in the application interface, and the method further includes A method comprising providing the output image to the application.
2. Identifying a design webpage from multiple webpages, The process involves generating training input condition data from the aforementioned design webpage, The method according to claim 1, further comprising training the model based on the training input condition data.
3. Generating the training input condition data from the aforementioned design webpage is, The method according to claim 2, comprising generating a UI edge map from a design web page, wherein the training input condition data includes the UI edge map.
4. Generating the training input condition data from the aforementioned design web page further involves The method according to claim 3, comprising generating caption data from at least one of the UI edge map or the design web page, wherein the caption data includes information about one or more UI elements included on the UI edge map, and the training input condition data includes the UI edge map and the caption data.
5. The aforementioned input condition data includes the activation of the color blindness setting, and the output image is generated as follows: The method according to any one of claims 1 to 4, comprising generating the output image with a monochrome color scheme using the model.
6. The input condition data includes display screen information relating to one or more display attributes of the display screen, and the output image is generated by The method according to any one of claims 1 to 5, comprising generating the output image by the model based on the display screen information.
7. The aforementioned model generates an intermediate output image, The method according to any one of claims 1 to 6, further comprising providing the intermediate output images for display within the interface by the model until the final output image is generated.
8. The interface is a first interface, the UI layout information is first UI layout information, the output image is a first output image, and the method further... The model receives second UI layout information relating to at least one UI element on the second interface of the application, Based on the first output image and the second UI layout information, the model generates a second output image. The method according to any one of claims 1 to 7, comprising providing the application with the second output image for display on the second interface of the application, according to the model.
9. The method according to any one of claims 1 to 8, wherein the interface is a new tab page in a browser tab.
10. The method according to any one of claims 1 to 9, wherein the user-generated text includes a natural language description of an image created using the model.
11. The method according to any one of claims 1 to 10, wherein the user-generated text includes a search query related to the entity, and the output image includes image data related to the entity.
12. The method according to any one of claims 1 to 11, wherein the UI layout information is obtained from the resource address to which the user-generated text is provided.
13. A non-temporary computer-readable medium that stores executable instructions causing the at least one processor to perform an operation, the operation being: To generate input condition data that includes UI layout information for at least one user interface (UI) element included in the application interface, The operation includes providing one or more prompts to the model, wherein the one or more prompts include user-generated text and the input condition data, and the operation further includes: The operation includes receiving an output image from the model, wherein the output image is generated considering the at least one UI element, and the operation further includes: A non-temporary computer-readable medium, which includes applying the output image to the interface.
14. The aforementioned operation further, A non-temporary computer-readable medium according to claim 13, comprising generating the UI layout information based on a structural description associated with the interface.
15. The aforementioned operation further, A non-temporary computer-readable medium according to claim 13 or 14, comprising displaying intermediate output images generated by the model until the final output image is displayed.
16. The interface is a first interface, the UI layout information is first UI layout information, the output image is a first output image, and the operation is further... Transmitting second UI layout information relating to at least one UI element on the second interface of the application, The second output image generated by the model is received using the first output image and the second UI layout information. A non-temporary computer-readable medium according to any one of claims 13 to 15, comprising applying the second output image to the second interface.
17. It is a device, At least one processor, A non-temporary computer-readable medium for storing executable instructions, wherein the executable instructions are transmitted to at least one processor. The executable instruction causes the system to receive one or more prompts, which include user-generated text and input condition data, wherein the input condition data includes UI layout information relating to at least one user interface (UI) element included in the application interface, and the executable instruction further causes the at least one processor to: To generate an output image in response to one or more of the above prompts, A device that causes the aforementioned output image to be provided to the aforementioned application.
18. The executable instruction includes a plurality of instructions, and the plurality of instructions are provided to the at least one processor. Identifying a design webpage from multiple webpages, The process involves generating training input condition data from the aforementioned design webpage, The apparatus according to claim 17, which trains a model based on the aforementioned training input condition data.
19. The executable instruction includes a plurality of instructions, and the plurality of instructions are provided to the at least one processor. The apparatus according to claim 18, wherein the apparatus generates a UI edge map from a design web page, and the training input condition data includes the UI edge map.
20. The executable instruction includes a plurality of instructions, and the plurality of instructions are provided to the at least one processor. The apparatus according to claim 19, wherein the apparatus generates caption data from the UI edge map, the caption data includes information about one or more UI elements included on the UI edge map, and the training input condition data includes the UI edge map and the caption data.