STYLE APPLICATION ENGINE
The document processing device addresses inefficiencies in applying brand elements by automating their application with machine learning, ensuring consistency and flexibility, and improving workflow efficiency on mobile devices.
Patent Information
- Application Number
- DE102025133086
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-16
- Filing Date
- 2025-08-19
- Publication Date
- 2026-04-09
AI Technical Summary
Conventional systems for applying brand elements such as colors, fonts, and image adjustments across digital documents are time-consuming, inconsistent, and impractical on mobile devices due to limited screen space, leading to inefficient workflows and unprofessional outcomes.
A document processing device that automates the application of brand elements with a single click, utilizing machine learning models for precise color matching and image recoloring, ensuring consistency and flexibility across documents, and offering undo/redo functionality for seamless iteration.
Enhances workflow efficiency, improves adherence to brand guidelines, and ensures visual consistency across documents, making it accessible and user-friendly on mobile devices.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED REGISTRATION
[0001] This application claims, pursuant to 35 USC § 119, the benefit of U.S. preliminary application No. 63 / 704,807, filed with the United States Patent and Trademark Office on October 8, 2024, and incorporates the disclosure of that application in its entirety by reference. BACKGROUND
[0002] The following concerns document processing in general, and in particular the application of style effects to documents. Document processing refers to techniques and processes for editing source documents (digital documents such as presentations, flyers, profile covers, etc.). In some cases, modified documents incorporate content from the source documents and may have a different style. Document processing is a combination of natural language processing (NLP) and image processing. For example, image processing is a form of data processing that involves manipulating or generating image data. Recently, machine learning (ML) models have been used in advanced document processing techniques.Among these ML models, transformer networks and generative models such as Generative Adversarial Networks (GANs) were used for various tasks, including recoloring, style transfer, generating images with perceptual metrics, generating images in conditional environments, and image manipulation. SUMMARY
[0003] This disclosure describes systems and methods for document processing. Embodiments of this disclosure include a document processing device that applies a style guide (for example, a brand with style-related elements) to a source document, triggered by receiving a one-click input via a user interface. In some examples, the source document comprises an Entity Component System (ECS) document (documents such as presentations, flyers, Instagram® posts, Stories with text animations and multi-frame edits, etc.). In some cases, a one-click input (“Apply brand” button) from a user triggers a process in which brand-specific colors, fonts, and image recoloring are automatically applied in one step, eliminating the need for manual adjustments.The document processing device improves creative flexibility through mixed variations on subsequent clicks and offers efficient undo and redo functions for quick switching between iterations.
[0004] A method, a device, a non-transitory computer-readable medium, and an image processing system are described. One or more aspects of the method, the device, the non-transitory computer-readable medium, and the system include: obtaining an image and a style guide, wherein the image shows an object with a first color and the style guide contains a second color; identifying a second color from the style guide based on an approximation criterion between the first color and the second color; and generating a modified image using an image generation model based on the image and the second color, wherein the modified image represents the object with the second color.
[0005] A method, a device, a non-transitory computer-readable medium, and an image processing system are described. One or more aspects of the method, the device, the non-transitory computer-readable medium, and the system include: obtaining a document and a style guide, wherein the document contains a text element with a first font and an image showing an object with a first color, and wherein the style guide contains a second font and a second color; applying the second font from the style guide to the text element to obtain a modified text element; applying the second color from the style guide to the image to obtain a modified image, wherein the modified image shows the object with the second color; and generating a modified document containing the modified image and the modified text element.
[0006] An image processing device, system, and method are described. One or more aspects of the device, system, and method include: a storage unit; a processing device coupled to the storage unit and configured to perform operations that include: obtaining an image and a style guide, wherein the image shows an object with a first color and the style guide contains a second color; generating a first color embedding and a second color embedding based on the first color and the second color, respectively; selecting the second color from the style guide by comparing the first color embedding of the first color with the second color embedding of the second color; and generating a modified image based on the image and the second color, wherein the modified image shows the object with the second color. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 shows an example of a document processing system according to aspects of the present disclosure. Fig. Figure 2 shows an example of a one-click trademark application procedure in accordance with aspects of the present disclosure. Fig. Figure 3 shows an example of a user interface for the application of a style guide according to aspects of the present disclosure. Fig. Figure 4 shows an example of the effect of applying a style guide according to aspects of the present disclosure. Fig. Figure 5 shows an example of a style guide with script selection according to aspects of the present revelation. Fig. Figure 6 shows an example of an effect of applying a typeface according to aspects of the present disclosure. Fig. Figure 7 shows an example of a color change effect according to aspects of the present revelation. Fig. Figure 8 shows another example of a color change effect according to aspects of the present revelation. Fig. Figure 9 shows another example of a color change effect according to aspects of the present revelation. Fig. Figure 10 shows an example of a user interface on a mobile device according to aspects of the present disclosure. Fig. Figure 11 shows another example of a user interface on a mobile device according to aspects of the present disclosure. Fig. Figure 12 shows an example of a method for image processing according to aspects of the present disclosure. Fig. Figure 13 shows an example of a document processing device according to aspects of the present disclosure. Fig. Figure 14 shows an example of a transformer network according to aspects of the present revelation. Fig. Figure 15 shows an example of a guided diffusion model according to aspects of the present revelation. Fig. Figure 16 shows an example of a color application tool according to aspects of the present disclosure. Fig. Figure 17 shows an example of a script application tool according to aspects of the present disclosure. Fig. Figure 18 shows an example of a style transformation element and a style guide setting element according to aspects of the present disclosure. Fig. Figure 19 shows an example of a user interface according to aspects of the present disclosure. Fig. Figure 20 shows an example of a style guide according to aspects of the present revelation. Fig. Figure 21 shows an example of a method for image processing according to aspects of the present disclosure. Fig. 22 shows an example of an algorithm according to aspects of the present revelation. Fig. 23 shows another example of an algorithm according to aspects of the present revelation. Fig. Figure 24 shows an example of a textual description of a color according to aspects of the present revelation. Fig. Figure 25 shows another example of a method for image processing according to aspects of the present disclosure. Fig. 26 shows another example of a style guide with scripture selection according to aspects of the present revelation. Fig. 27 shows another example of an effect of applying a typeface according to aspects of the present disclosure. Fig. Figure 28 shows an example of a change-of-state effect according to aspects of the present disclosure. Fig. 29 shows another example of a change-of-state effect according to aspects of the present disclosure. Fig. Figure 30 shows an example of a step-by-step procedure for training a machine learning model according to aspects of the present disclosure. Fig. Figure 31 shows an example of a computer for image processing according to aspects of the present disclosure. Fig. Figure 32 shows an example of a diffusion transformer architecture (DiT architecture) according to aspects of the present disclosure. DETAILED DESCRIPTION
[0007] This disclosure describes systems and methods for document processing. Embodiments of this disclosure include a document processing device that applies a style guide (for example, a brand with style-related elements) to a source document, triggered by receiving a single click input via a user interface. In some examples, the source document comprises an Entity Component System (ECS) document (documents such as presentations, flyers, Instagram® posts, Stories with text animations and edits across multiple image frames, etc.). In some cases, a single click input (“Apply brand” button) from a user triggers a process in which brand-specific colors, fonts, and image recoloring are automatically applied in one step, eliminating the need for manual adjustments.The document processing device improves creative flexibility through mixed variations on subsequent clicks and offers efficient undo and redo functions for quick switching between iterations.
[0008] Conventional systems involve a time-consuming and inconsistent process for applying brand elements (such as colors, fonts, and image adjustments) across digital documents, especially in multi-page or multi-slide projects. These systems cannot be used effectively on mobile devices because limited screen space makes manual editing tedious and inefficient. For example, manually applying branding to every element of a document is labor-intensive and time-consuming, leading to inefficient workflows. Consequently, user satisfaction decreases. Furthermore, the inconsistent application of brand guidelines across different components and pages in conventional systems results in unprofessional and inconsistent outcomes. However, brand identity is essential for businesses.
[0009] Furthermore, mobile devices are increasingly used for professional tasks, but the limited screen space makes document editing difficult. Users are forced to navigate cumbersome interfaces to manually adjust brand elements, making editing on mobile devices impractical. For example, designers often need to try out different variations of brand elements (such as fonts or color schemes), but manually experimenting with these combinations is time-consuming (especially on mobile devices). There is a need for systems and processes that allow for rapid testing of variations without violating brand guidelines.
[0010] Embodiments of the present disclosure provide a document processing device for the automated application of a style guide (for example, brand elements). The document processing device automates the application of fonts, colors, and images across an entire document with a single click. This saves users time and effort, eliminating the need to make manual updates for each element of a source document.
[0011] In some embodiments, the document processing device performs context-aware branding, which includes a process for recognizing font sizes and applying appropriate variations. The document processing device incorporates a machine learning model (for example, a color matching network) that generates color embeddings and calculates the cosine similarity for color matching. The document processing device offers a high degree of precision, ensuring improved adherence to brand guidelines and increasing visual consistency.
[0012] In some implementations, the document processing device performs dynamic image recoloring using a custom generative model or an intelligent recoloring API. The document processing device offers selective recoloring, which preserves image quality while ensuring brand consistency. Additionally, the document processing device performs shuffling and includes synchronization capabilities that involve a process of blending brand variations and synchronizing colors across multiple pages or slides. This enhances creative flexibility (for example, integration and dynamic adaptation) while maintaining brand integrity.
[0013] Embodiments of the present disclosure can be implemented on mobile devices with relatively small screens, making them more accessible and user-friendly for mobile professionals (for example, prioritizing mobile usability and increasing effectiveness in today's multi-device environment).
[0014] Embodiments of the present disclosure offer an adaptive one-click system that applies brand elements (for example, colors, fonts, and image recoloring) to entire multi-page documents with a single action, thus ensuring consistent and context-aware branding. The one-click system includes a shuffle function for quickly generating brand-compliant variations, and its undo / redo capabilities enable seamless iteration (for example, advantageous for mobile users with limited screen space). The combination of automation, flexibility, and mobile optimization improves workflow efficiency compared to existing manual methods.
[0015] The document processing device can be deployed on user devices with varying screen sizes, including mobile devices. By combining multiple manual tasks into a single click and offering undo / redo functionality, the document processing device provides an intuitive and seamless user experience regardless of the device. The document processing device delivers consistent results requiring minimal manual adjustments.
[0016] This disclosure describes systems and methods that improve upon traditional document processing models by increasing the efficiency of applying colors to objects in an input image. For example, a user provides an image containing a target object, selects the "Apply Colors" parameter, and clicks a button to apply the brand to the input image. The Dynamic Brand Identity Color Matching System (DBICMS) uses a machine learning model to calculate candidate color embeddings from a style guide and compares these color embeddings to the color embedding of an object in the input image. This improves the efficiency of applying colors to objects in the input image. Additionally, it enhances contextual compatibility among the objects in the input image because desired colors from the style guide are applied to the objects to ensure brand consistency.
[0017] The term "image" refers to a pixel-based image, a vector image, a media element, or a page of a multimedia document. In some examples, an input document comprises a series of slide sets, with each slide set being considered an image. The image may contain one or more media elements, such as a text element, a picture element, a static element, an animated element, etc. The term "modified image" refers to a modified pixel-based image, a modified vector image, a modified media content element, or a modified page of a multimedia document after a style guide operation has been applied to an original image. A modified image is used to differentiate itself from the original image. Compared to the original image, the modified image may have a different font style, font color, and / or size, corresponding to a text element.Additionally or alternatively, the modified image can have a different graphic color for an image element than the original image.
[0018] The term "style guide" refers to a collection of style-related attributes and assets, including a font, text color, background color, logo, or any combination thereof. A style guide is associated with a predefined theme or brand. The style guide can be modified, for example, by adding or removing a font style from the font pool or by adding or removing a color from the color palette. The style guide can be applied to a single page of an input document (for example, a multi-page flyer) or to all pages of the input document. In some cases, a style guide may refer to an image editing tool or interface where a user applies the style guide to an input image.
[0019] The term "color embedding" refers to the representation of colors in a numerical space, for example, as vectors in a multidimensional embedding space. A machine learning model is trained to encode color information in such a way as to capture relationships and similarities between different colors. In some examples, the machine learning model receives an input prompt containing a color description of an object and generates a color embedding based on this input prompt. Alternatively, the machine learning model receives an input prompt containing an image representing an object and generates a color embedding based on this input prompt. In some cases, colors are embedded in different color spaces, such as RGB, Lab, or a learned color embedding space.A learned color embedding assigns colors to a multidimensional space in which colors that are perceptually similar are located closer together.
[0020] Implementations of the present disclosure are used in document processing, for example, in changing fonts, applying colors, or recoloring graphics of an input document. Examples of applications in the document processing context are given with reference to the Fig. 2-11 explains. Details regarding the architecture of a sample document processing system are explained with reference to the Fig. Sections 1 and 13-20 describe this process. Details on various processes (for example, changing fonts, applying colors, recoloring graphics) are provided with reference to... Fig. 12 and 21-29 are explained. Details of an example procedure for training a machine learning model are given with reference to Fig. 30. Details of a computer system for document processing are given with reference to Fig. 31 is given. Document and image processing
[0021] Fig. Figure 1 shows an example of a document processing system according to aspects of the present disclosure. The example shown includes User 100, User Device 105, Document Processing Device 110, Cloud 115, and Database 120. The Document Processing Device 110 is an example of, or includes, aspects of the corresponding element that relates to Fig. 13 is described.
[0022] In the Fig. In the example shown, user 100 provides an input image. The input image shows a dog wearing a scarf and a hat. The scarf and hat are red. The input image contains text (for example, "happy holidays") in a first font. In some cases, user 100 is provided with a style guide in an image editing user interface. User 100 wants to apply the style guide to the input image by clicking "Apply brand." The input image is transferred via the user device 105 and the cloud 115 to the document processing device 110.
[0023] Document processing device 110 creates a first color embedding based on the scarf's color (i.e., red). It then creates a second color embedding based on a second color from the style guide (for example, a brand-related color like green). The second color (green) is selected from the style guide by comparing the first color embedding of the scarf's color with the second color embedding of the second color. In some examples, a second font is selected from the style guide and applied to the text "happy holidays" based on the font size of "happy holidays" relative to the font size of other text in the input image. Document processing device 110 returns a modified image to user 100 via cloud 115 and user device 105. The modified image shows the dog with the second color (green) and contains the modified text "happy holidays" in the second font.
[0024] The user device 105 can be, for example, a personal computer, laptop computer, mainframe computer, palmtop computer, personal assistant, mobile device, or other suitable processing device. In some examples, the user device 105 includes software that contains an image processing application (for example, an image generator, an image editing tool). In some examples, the image processing application on the user device 105 can include functions of the document processing device 110.
[0025] A user interface can enable the user 100 to interact with the user device 105. In some embodiments, the user interface can include an audio device (for example, an external speaker system), an external display device (for example, a screen), or an input device (for example, a remote control connected directly or via an I / O control module to the user interface). In some cases, a user interface can be a graphical user interface (GUI). In some examples, a user interface can be represented as code that is sent to the user device 105 and rendered locally by a browser.
[0026] The document processing device 110 comprises a computer-implemented network with a user interface, a style guide engine, a language generation model, and an image generation model. The document processing device 110 may also include a processing unit, a storage unit, an I / O module, and a user interface. A training unit may be implemented on a separate device from the document processing device 110. The training unit is used to train a machine learning model. Additionally, the document processing device 110 can communicate with the database 120 via the cloud 115. In some cases, the architecture of the document processing model is also referred to as a network, machine learning model, or mesh model. Further details regarding the architecture of the document processing device 110 are provided in relation to… Fig. 13-20 described. Further details on the operation of the document processing device 110 are described in relation to Fig. Described in sections 2, 12 and 21-29.
[0027] In some cases, the Document Processing Device 110 is implemented on a server. A server provides functionality to one or more users over one or more networks. In some cases, the server comprises a single microprocessor board with a microprocessor responsible for controlling all aspects of the server. In some cases, a server uses microprocessors and protocols to exchange data with other devices / users over one or more networks, for example, via the Hypertext Transfer Protocol (HTTP) and the Simple Mail Transfer Protocol (SMTP), although other protocols such as the File Transfer Protocol (FTP) or the Simple Network Management Protocol (SNMP) may also be used. In some cases, a server is configured to send and receive Hypertext Markup Language (HTML) files (for example, to display web pages).In various embodiments, a server comprises a universal computing unit, a personal computer, a laptop, a mainframe computer, a supercomputer, or another suitable processing device.
[0028] The Document Processing Device 110 can include an artificial neural network (ANN) to apply a style guide to input content (for example, applying or adjusting color, recoloring graphics). An ANN is a hardware or software component with a number of connected nodes (i.e., artificial neurons) that loosely correspond to the neurons of the human brain. Each connection (or edge) transmits a signal from one node to another (similar to the physical synapses in the brain). When a node receives a signal, it processes that signal and then transmits the processed signal to other connected nodes. In some cases, the signals between nodes consist of real numbers, and the output of each node is calculated as a function of the sum of its inputs.In some examples, nodes can determine their output using other mathematical algorithms (for example, by choosing the maximum of the inputs as the output) or with another suitable node activation algorithm. Each node and edge is assigned one or more weighting factors that determine how the signal is processed and transmitted.
[0029] Implementations of the present disclosure are used in document processing, e.g., when changing fonts, applying colors, or recoloring graphics in a source document. Application examples in the context of document processing are given with reference to Fig. 2-11. Details regarding the architecture of an exemplary document processing system are given in relation to Fig. 1 and Fig. 13-20. Details of the various processes (e.g., changing fonts, applying colors, recoloring graphics) are described in relation to Fig. 12 and Fig. 21-29. Details of an example for training a machine learning model are given in relation to Fig. 30. Details of a computer system for document processing are described in relation to Fig. 31 described.
[0030] Fig. Figure 1 shows an example of a document processing system according to aspects of the present disclosure. The example shown includes User 100, User Device 105, Document Processing Device 110, Cloud 115, and Database 120. The Document Processing Device 110 is an example of—or includes aspects of—the corresponding element that relates to Fig. 13 is described.
[0031] In a Fig. In the example shown, an input image is provided by user 100. The input image shows a dog wearing a scarf and a hat. The scarf and hat are red. The input image contains text (e.g., "happy holidays") in a first font. In some cases, a style guideline is provided to user 100 on an image editing user interface. User 100 wants to apply the style guideline to the input image by clicking the "Apply Mark" button. The input image is transferred to the document processing device 110, e.g., via user device 105 and cloud 115.
[0032] Document processing device 110 creates a first color embedding based on the scarf's color (i.e., red). Document processing device 110 creates a second color embedding based on a second color from the style guide (e.g., a brand-related color like green). The second color (green) is selected from the style guide by comparing the first color embedding of the scarf's color with the second color embedding of the second color. In some examples, a second font is selected from the style guide and applied to the text "happy holidays" based on the font size of "happy holidays" relative to the font size of other text in the input image. Document processing device 110 returns a modified image to user 100 via Cloud 115 and user device 105. The modified image shows the dog in the second color (green) and contains the modified text "happy holidays" in the second font.
[0033] The user device 105 can be a personal computer, laptop computer, mainframe, palmtop computer, personal assistant, mobile device, or any other suitable processing device. In some examples, the user device 105 includes software that contains an image processing application (e.g., an image generator, an image editing tool). In some examples, the image processing application on the user device 105 can include functions of the document processing device 110.
[0034] A user interface can enable user 100 to interact with user device 105. In some embodiments, the user interface may include an audio device such as an external speaker system, an external display device such as a screen, or an input device (e.g., a remote control connected directly or via an I / O controller module to the user interface). In some cases, a user interface may be a graphical user interface (GUI). In some examples, a user interface may be represented as code that is sent to user device 105 and rendered locally by a browser.
[0035] The document processing device 110 comprises a computer-implemented network with a user interface, a style guide engine, a language generation model, and an image generation model. The document processing device 110 may also include a processor unit, a memory unit, an I / O module, and a user interface. A training component may be implemented on a separate device from the document processing device 110. The training component is used to train a machine learning model. Additionally, the document processing device 110 can communicate with the database 120 via the cloud 115. In some cases, the architecture of the document processing model is also referred to as a network, a machine learning model, or a mesh model. Further details regarding the architecture of the document processing device 110 are provided in relation to… Fig. 13-20 provided. Further details regarding the operation of the document processing device 110 are provided in relation to Fig. 2, Fig. 12 and Fig. 21-29 provided.
[0036] In some cases, the Document Processing Device 110 is implemented on a server. A server provides one or more functions to one or more users connected over one or more of the various networks. In some cases, the server comprises a single microprocessor board containing a microprocessor responsible for controlling all aspects of the server. In some cases, a server uses microprocessors and protocols to exchange data with other devices / users over one or more of the networks—via Hypertext Transfer Protocol (HTTP) and Simple Mail Transfer Protocol (SMTP), although other protocols such as File Transfer Protocol (FTP) and Simple Network Management Protocol (SNMP) may also be used. In some cases, a server is configured to send and receive Hypertext Markup Language (HTML)-formatted files (e.g., for displaying web pages).In various embodiments, a server comprises a general-purpose computing device, a personal computer, a laptop computer, a mainframe computer, a supercomputer, or any other suitable processing device.
[0037] The Document Processing Device 110 can include an artificial neural network (ANN) for applying a style policy to input content (e.g., applying or adjusting a color, recoloring graphics). An ANN is a hardware or software component that includes a number of connected nodes (i.e., artificial neurons) that roughly correspond to the neurons in the human brain. Each connection, or edge, transmits a signal from one node to another (similar to the physical synapses in a brain). When a node receives a signal, it processes the signal and then transmits the processed signal to other connected nodes. In some cases, the signals between nodes consist of real numbers, and the output of each node is calculated as a function of the sum of its inputs. In some examples, nodes can determine their output using other mathematical algorithms (e.g.,(by selecting the maximum of the inputs as the output) or any other suitable algorithm to activate the node. Each node and edge is associated with one or more node weights that determine how the signal is processed and transmitted.
[0038] Cloud 115 is a computer network configured to provide computing resources, such as data storage and processing power, on demand. In some examples, Cloud 115 provides resources without active user management. The term "cloud" is sometimes used to describe data centers accessible to many users via the internet. Some large cloud networks have functions distributed across multiple locations from central servers. A server is referred to as an edge server if it has a direct or near-direct connection to a user. In some cases, Cloud 115 is limited to a single organization. In other examples, Cloud 115 is available to many organizations. In one example, Cloud 115 comprises a multi-layered communications network with multiple edge routers and core routers.In another example, Cloud 115 is based on a local collection of switches at a single physical location.
[0039] Database 120 is an organized collection of data. For example, Database 120 stores data (such as a dataset for training a machine learning model) in a specified format known as a schema. Database 120 can be structured as a single database, a distributed database, multiple distributed databases, or a disaster recovery backup database. In some cases, a database controller manages data storage and processing within Database 120. In some cases, a user interacts with the database controller. In other cases, database controllers can operate automatically without user intervention.
[0040] Fig. Figure 2 shows an example of a method 200 for applying trademarks with a single click according to aspects of the present disclosure. In some examples, the method 200 describes an operation of the document processing model 1320, which, with respect to Fig. 13 is described. In some examples, these operations are performed by a system that includes a processor which executes a set of codes to control functional elements of a device, such as the one described in Fig. 1. To control the document processing device described.
[0041] Additionally or alternatively, steps of procedure 200 can be performed using special hardware. Generally, these operations are carried out according to the procedures and processes described in accordance with aspects of this disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in conjunction with other operations.
[0042] In step 205, the user provides an image. In some cases, the operations of this step relate to—or can be performed by—a user, as in relation to... Fig. 1. In some cases, the image originates from an input document. The input document may consist of a video with a series of frames, and the image relates to one of the frames of the video.
[0043] In step 210, the user retrieves style guide resources from a database. In some cases, the operations of this step are performed by—or can be performed by—a user, as in the case of... Fig. 1. In some cases, the image shows a primary color, and the style guideline includes at least one color different from that primary color. In some examples, the style guideline includes a font, a text color, a background color, a logo, or any combination thereof.
[0044] In step 215, the user modifies the style guideline. In some examples, the user creates or edits a style guideline by selecting a font from a set of candidate fonts, a color from a set of candidate colors, or a logo from a set of candidate logos. In some cases, the operations of this step relate to—or can be performed by—a user, as in the case of… Fig. 1 and Fig. 3 described.
[0045] In step 220, the system generates a modified image based on the modified style guideline. In some cases, the operations of this step relate to—or can be performed by—a document processing device, as in the case of… Fig. 1 and Fig. 13. In some cases, the modified image displays the object with the second color from the style guide. In some cases, the system receives a single click input via a style transformation element, generating the modified image based on that single click input. In some cases, the system generates a modified document containing the modified image. In some cases, the system applies a first font from the style guide to a first text element in the document and a second font (different from the first font) from the style guide to a second text element in the document.
[0046] Fig. Figure 3 shows an example of a user interface 300 for applying a style guideline according to aspects of the present disclosure. The example shown includes the user interface 300, image 305, style transformation element 310, style guideline setting element 315, candidate logos 320, candidate colors 325, and candidate fonts 330.
[0047] In some embodiments, the user interface 300 retrieves an image 305 and a style guideline, where the image 305 represents an object with a first color and the style guideline contains a second color. In some examples, the user interface 300 provides a style transformation element 310 within a user interface 300. In some examples, the user interface 300 receives a single click input via the style transformation element 310, and the modified image is generated based on the single click input. In some examples, the user interface 300 identifies a color setting parameter, and the second color is selected based on the color setting parameter. In some examples, the user interface 300 receives a page selection input.
[0048] In some embodiments, the user interface 300 retrieves a document and a style guideline, wherein the document comprises a text element with a first font and an image 305 representing an object with a first color, and wherein the style guideline contains a second font and a second color. In some examples, the user interface 300 provides a style transformation element 310 in a user interface 300. In some examples, the user interface 300 receives a single click input via the style transformation element 310, and the modified image is generated based on the single click input. In some examples, the user interface 300 receives a page selection input.
[0049] In a Fig. In the example shown, user interface 300 displays a page of a document before a style guideline is applied (e.g., a brand resource collection).
[0050] According to some embodiments, the user interface 300 receives user input with a request to apply the style policy to the document. In some examples, the document processing model 1320 (as in Fig. (Described in section 13) provides a style transformation element in the user interface 300. The user interface 300 receives a single click input via the style transformation element, and the modified document is generated based on this single click input. In some examples, the user interface 300 retrieves a selection parameter that corresponds to a style attribute from the style policy, and the style attribute is applied to the document based on the selection parameter. In some examples, the user interface 300 retrieves a color palette, where the style policy contains the color palette.
[0051] In some examples, the UI 300 provides a style policy application tool to a user. The UI 300 receives style policy application input via the style policy application tool, and the style policy is based on the style policy application input. In some examples, the UI 300 provides a state change element. The UI 300 receives state change input via the state change element, and the modified document is generated based on the state change input. In some examples, the UI 300 receives a setting.
[0052] The user interface 300 is an example of – or includes aspects of – the corresponding element, which relates to Fig. 4-11, Fig. 13, Fig. 18, Fig. 19 and Fig. 26-29 is described. The style transformation element 310 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 4, Fig. 6-11, Fig. 18, Fig. 19, Fig. 26 and Fig. 27 is described. The style guide setting element 315 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 18 and Fig. 19 is described. Candidate logos 320 are an example of - or include aspects of - the corresponding element, which relates to Fig. 4 is described. Candidate colors 325 are an example of - or encompass aspects of - the corresponding element, which relates to Fig. 4, Fig. 7-11 and Fig. 20 is described. Candidate typefaces 330 are an example of - or include aspects of - the corresponding element that relates to Fig. 4, Fig. 26 and Fig. 27 is described.
[0053] Fig. Figure 4 shows an example of the effect of applying a style guideline according to aspects of the present disclosure. The example shown includes the user interface 400, modified image 405, style transformation element 410, candidate logos 415, candidate colors 420, and candidate fonts 425. The user interface 400 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 3, Fig. 5-11, Fig. 13, Fig. 18, Fig. 19 and Fig. 26-29 is described.
[0054] In a Fig. In example 4, the user interface 400 shows a modified page of the one shown in Fig. 3. The aforementioned document, after a style guideline (e.g., a brand resource collection) has been applied. The colors and fonts are selected from the style guideline (the brand resource collection), which is located in the left pane of user interface 400.
[0055] Style transformation element 410 is an example of—or encompasses aspects of—the corresponding element, which relates to Fig. 3, Fig. 6-11, Fig. 18, Fig. 19, Fig. 26 and Fig. 27 is described. Candidate logos 415 are an example of - or include aspects of - the corresponding element, which relates to Fig. 3 is described. Candidate colors 420 are an example of - or encompass aspects of - the corresponding element, which relates to Fig. 3, Fig. 7-11 and Fig. 20 is described. Candidate typefaces 425 are an example of - or include aspects of - the corresponding element that relates to Fig. 3, Fig. 26 and Fig. 27 is described.
[0056] Fig. Figure 5 shows an example of a style guideline with font selection according to aspects of the present disclosure. The example shown includes the user interface 500, document 505, style transformation element 510, first font 515, second font 520, first text element 525, second text element 530, and third text element 535. The user interface 500 is an example of—or includes aspects of—the corresponding element, which is related to Fig. 3, Fig. 4, Fig. 6-11, Fig. 13, Fig. 18, Fig. 19 and Fig. 26-29 is described.
[0057] Fig. Document 5 shows a page of a document before a font is applied to the document via user interface 500. Document 505 is an example of—or includes aspects of—the corresponding element related to Fig. 7, Fig. 10 and Fig. 26 is described. The first font 515 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 6, Fig. 20, Fig. 26 and Fig. 27 is described. Second font 520 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 6, Fig. 20, Fig. 26 and Fig. 27 is described.
[0058] In some examples, the first font, 515, which is marked as a heading in the style guide (i.e., font for the heading role), differs from the actual font of the second text element, 530 (e.g., a text segment with the largest font size). In some examples, the second font, 520, with the body text role in the style guide, differs from the actual font of the first text element, 525 (e.g., text with the second largest font size). In some cases, the third text element, 535, comprises the remaining text on the page of document 505.
[0059] Fig. Figure 6 shows an example of the effect of applying a typeface according to aspects of the present disclosure. The example shown includes the user interface 600, modified document 605, style transformation element 610, first typeface 615, second typeface 620, first text element 625, second text element 630, and third text element 635. The user interface 600 is an example of—or includes aspects of—the corresponding element, which is related to Fig. 3-5, Fig. 7-11, Fig. 13, Fig. 18, Fig. 19 and Fig. 26-29 is described.
[0060] Fig. Figure 6 shows a modified page of the in Fig. 5 mentioned document, after a font was applied to the document via a single click on the style transformation element 610 (e.g., the "Apply Mark" button) in the user interface 600, to obtain the modified document 605. The document processing model 1320 (as in Fig. (Described in section 13) assigns a font from the style guide (displayed in the left pane of user interface 600) to a corresponding text segment on the document page (i.e., size matching). In some examples, a first font 615, marked as a heading in the style guide (i.e., font for heading role), is applied to the second text element 630 (e.g., a text segment) with the largest font size on the document page. A second font 620, with the body text role in the style guide, is applied to the first text element 625 with the second largest font size. In some cases, a third font, marked as "None," is applied to the third text element 635 (e.g., the remaining text on the document page). In some cases, the third font is a system-defined default font.After clicking the "Apply Marker" button, document processing model 1320 applies the first font 615 and the second font 620 to one text element each. The modified document 605 contains the first text element 625, the second text element 630, and the third text element 635 in their respective new font styles / sizes. A style guideline or marker can contain multiple fonts with the same role. For example, two heading fonts and three body text fonts. Consequently, randomly shuffling a style guideline (or marker) by clicking the style transformation element 610 in the user interface 600 would apply different variations.
[0061] Document processing model 1320 receives a selection parameter that corresponds to a style attribute from the style guideline. The style attribute is then applied to the document based on the selection parameter to produce a modified document. In some examples, the document includes... Fig. 5 and the modified document in Fig. 6 each, one multimedia asset.
[0062] In some examples, the 1320 document processing model offers seamless undo / redo functionality, allowing users to quickly switch between mark variations.
[0063] In some implementations, the Document Processing Model 1320 applies the brand's font variants and categorizes them as heading, body text, or decorative elements based on the brand kit. The Document Processing Model 1320 detects font sizes in an input document and intelligently applies the appropriate font variants in ascending size order to ensure consistency across all text elements. The Document Processing Model 1320 can identify heading, body text, and other fonts in the document and intelligently map them to the correct brand font roll.
[0064] In some implementations, the Document Processing Model 1320 analyzes the colors present in the document. The Document Processing Model 1320 then applies brand colors by using cosine similarity to determine the best match, thus ensuring optimal color alignment within the brand guidelines.
[0065] In some embodiments, the 1320 document processing model selectively recolors chosen elements in images, for example by converting a non-brand color (e.g., an orange cap) into a brand color (e.g., a brand-specific yellow), using an image generation model (or API) for precise recoloring while maintaining image integrity.
[0066] Regarding multi-page synchronization, the Document Processing Model 1320 ensures that colors remain consistent across all pages or slides, giving the document a coherent appearance. In some cases, if there are duplicate slides or pages in the multi-page document, the Document Processing Model 1320 applies the exact same random variations to maintain uniformity throughout the presentation.
[0067] The modified document 605 is an example of—or includes aspects of—the corresponding element relating to Fig. 8, Fig. 9, Fig. 11 and Fig. 27 is described. The style transformation element 610 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 3, Fig. 4, Fig. 7-11, Fig. 18, Fig. 19, Fig. 26 and Fig. 27 is described. The first font, 615, is an example of—or encompasses aspects of—the corresponding element, which relates to Fig. 5, Fig. 6, Fig. 20, Fig. 26 and Fig. 27 is described. The second font, 620, is an example of—or encompasses aspects of—the corresponding element, which relates to Fig. 5, Fig. 6, Fig. 20, Fig. 26 and Fig. 27 is described. The first text element 625 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 7 and Fig. 26 is described. The second text element 630 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 7 and Fig. 26 is described. The third text element 635 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 7 and Fig. 26 is described.
[0068] Fig. Figure 7 shows an example of a recoloring effect for images according to aspects of the present disclosure. The example shown includes the user interface 700, document 705, style transformation element 735, and candidate colors 740. The user interface 700 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 3-6, Fig. 8-11, Fig. 13, Fig. 18, Fig. 19 and Fig. 26-29 is described.
[0069] In one aspect, document 705 comprises first image element 710, second image element 715, first text element 720, second text element 725, and third text element 730. Document 705 is an example of—or comprises—aspects of—the corresponding element that relates to Fig. 5, Fig. 10 and Fig. 26 is described.
[0070] The first image element 710 is an example of – or includes aspects of – the corresponding element, which relates to Fig. 8 is described. The second image element 715 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 8 is described. The first text element 720 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 6 and Fig. 26 is described. The second text element 725 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 6 and Fig. 26 is described. The third text element 730 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 6 and Fig. 26 is described. Style transformation element 735 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 3, Fig. 4, Fig. 6, Fig. 8-11, Fig. 18, Fig. 19, Fig. 26 and Fig. 27 is described. Candidate colors 740 are an example of - or encompass aspects of - the corresponding element, which relates to Fig. 3, Fig. 4, Fig. 8-11 and Fig. 20 is described.
[0071] Fig. Figure 8 shows an example of a recoloring effect for images according to aspects of the present disclosure. The example shown includes the user interface 800, modified document 805, style transformation element 835, and candidate colors 840. The user interface 800 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 3-7, Fig. 9-11, Fig. 13, Fig. 18, Fig. 19 and Fig. 26-29 is described.
[0072] The modified document 805 comprises the first image element 810, the second image element 815, the first modified text element 820, the second modified text element 825, and the third modified text element 830. For example, user interface 800 displays an image in the center of a document. The image contains a dog, a scarf, and a hat. The dog is wearing the scarf and the hat. The scarf and the hat are red.
[0073] The modified document 805 is an example of—or includes aspects of—the corresponding element relating to Fig. 6, Fig. 9, Fig. 11 and Fig. 27 is described. The first image element 810 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 7 is described. The second image element 815 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 7 is described.
[0074] The first modified text element 820 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 9 and Fig. 27 is described. The second modified text element 825 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 9 and Fig. 27 is described. The third modified text element 830 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 9 and Fig. 27 is described. Style transformation element 835 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 3, Fig. 4, Fig. 6, Fig. 7, Fig. 9-11, Fig. 18, Fig. 19, Fig. 26 and Fig. 27 is described. Candidate colors 840 are an example of - or encompass aspects of - the corresponding element, which relates to Fig. 3, Fig. 4, Fig. 7, Fig. 9-11 and Fig. 20 is described.
[0075] Fig. Figure 9 shows an example of a recoloring effect for images according to aspects of the present disclosure. The example shown includes the user interface 900, modified document 905, style transformation element 935, and candidate colors 940. The user interface 900 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 3-8, Fig. 10, Fig. 11, Fig. 13, Fig. 18, Fig. 19 and Fig. 26-29 is described.
[0076] The modified document 905 includes the first modified image element 910, the second modified image element 915, the first modified text element 920, the second modified text element 925, and the third modified text element 930. For example, a user wants to recolor the images in a document. The "Recolor Graphics" setting is enabled via the Style Guides Application Tool in the left pane of user interface 900. After a single click on the "Apply Marker" button is received, user interface 900 displays a modified document in the right pane. The dog in the modified document has the same color as the dog in the input document (see Fig. 8) The color of the scarf and hat was changed to green (as opposed to red in Fig. 8).
[0077] The modified document 905 is an example of—or includes aspects of—the corresponding element relating to Fig. 6, Fig. 8, Fig. 11 and Fig. 27 is described. The first modified text element 920 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 8 and Fig. 27 is described. The second modified text element 925 is an example of – or includes aspects of – the corresponding element, which relates to Fig. 8 and Fig. 27 is described. The third modified text element 930 is an example of - or includes aspects of - the corresponding element, which relates to Fig. 8 and Fig. 27 is described. Style transformation element 935 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 3, Fig. 4, Fig. 6-8, Fig. 10, Fig. 11, Fig. 18, Fig. 19, Fig. 26 and Fig. 27 is described. Candidate colors 940 are an example of - or encompass aspects of - the corresponding element, which relates to Fig. 3, Fig. 4, Fig. 7, Fig. 8, Fig. 10, Fig. 11 and Fig. 20 is described.
[0078] Fig. Figure 10 shows an example of a user interface 1000 on a mobile device according to aspects of the present disclosure. The example shown includes the user interface 1000, document 1005, style transformation element 1010, style guideline 1015, and candidate colors 1020. The user interface 1000 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 3-9, Fig. 11, Fig. 13, Fig. 18, Fig. 19 and Fig. 26-29 is described.
[0079] Fig. Figure 10 shows an example of a style guide application tool and user interface 1000 on a mobile device with a relatively small screen. A document 1005 (e.g., an input document provided by a user) is displayed at the top of the user interface 1000. The document 1005 contains text content such as "product launch party," dates, patterns, artwork, etc. A user can click the style transformation element 1010 (e.g., the "Apply Brand" button) to apply the style guideline 1015 and modify aspects of the document 1005, such as text font, image background color, object color, etc. The user interface 1000 displays candidate colors 1020 at the bottom of the interface. The style guide interface has a vertical layout suitable for mobile electronic devices.
[0080] Document 1005 is an example of – or includes aspects of – the corresponding element relating to Fig. 5, Fig. 7 and Fig. 26 is described. The style transformation element 1010 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 3, Fig. 4, Fig. 6-9, Fig. 11, Fig. 18, Fig. 19, Fig. 26 and Fig. 27 is described. Style guideline 1015 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 11 is described. Candidate colors 1020 are an example of - or encompass aspects of - the corresponding element, which relates to Fig. 3, Fig. 4, Fig. 7-9, Fig. 11 and Fig. 20 is described.
[0081] Fig. Figure 11 shows an example of a user interface 1100 on a mobile device according to aspects of the present disclosure. The example shown includes the user interface 1100, modified document 1105, style transformation element 1110, style guideline 1115, candidate colors 1120, and message 1125. The user interface 1100 is an example of—or includes aspects of—the corresponding element that relates to Fig. 3-10, Fig. 13, Fig. 18, Fig. 19 and Fig. 26-29 is described.
[0082] Fig. Figure 11 shows an example of a style guide application tool and the user interface 1100 on a mobile device with a relatively small screen. After receiving user input (e.g., a single click via the style transformation element 1110 "Apply Mark"), the user interface 1100 displays a modified document in the upper area of the user interface 1100. The color and font of one or more elements in the previous document (see Figure 11) are applied. Fig. 10) are modified based on the style guideline to produce the modified document 1105. For example, the text "product launch party" has a different color and font than in the previous document. Artificial elements (e.g., circles, semicircles) are now orange. The user interface 1100 displays candidate colors 1120 at the bottom of the interface and also displays message 1125. The style guide interface has a vertical layout suitable for mobile electronic devices.
[0083] The modified document 1105 is an example of – or includes aspects of – the corresponding element relating to Fig. 6, Fig. 8, Fig. 9 and Fig. 27 is described. The style transformation element 1110 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 3, Fig. 4, Fig. 6-10, Fig. 18, Fig. 19, Fig. 26 and Fig. 27 is described. Style guideline 1115 is an example of—or includes aspects of—the corresponding element, which relates to Fig. 10 is described. Candidate colors 1120 are an example of - or encompass aspects of - the corresponding element, which relates to Fig. 3, Fig. 4, Fig. 7-10 and Fig. 20 is described.
[0084] Fig. Figure 12 shows an example of a method 1200 for image processing according to aspects of the present disclosure. In some examples, these operations are performed by a system comprising a processor that executes a set of codes to control functional elements of a device. Additionally or alternatively, certain processes are performed using special hardware. In general, these operations are performed according to the methods and procedures described in accordance with aspects of the present disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in conjunction with other operations.
[0085] In step 1205, the system receives an image and a style guideline, where the image represents an object with a first color and the style guideline contains a second color. An example of an image is image 305, which is in Fig. 3 is described. The style guideline is an example of—or encompasses aspects of—the corresponding element, which relates to Fig. 3-8, Fig. 10-11, Fig. 18-19 and Fig. 26-27 is described. In some examples, a style guideline includes a set of fonts, a set of colors, a set of logos, a set of templates, or any combination thereof. Users can create a new style guideline or modify an existing one. The second color from the style guideline is different from the first color. In some cases, the operations of this step relate to—or can be performed by—a user interface, as in relation to Fig. 3-11, Fig. 13, Fig. 18, Fig. 19 and Fig. described on pages 26-29. Network architecture
[0086] Fig. Figure 13 shows an example of an image processing device according to aspects of the present disclosure. The example shown includes the document processing device 1300, the processing unit 1305, the I / O module 1310, the storage unit 1315, the document processing model 1320, and the training unit 1345. The document processing device 1300 is an example of, or includes, aspects of the corresponding element that relates to Fig. 1 is described.
[0087] The document processing device 1300 can include an example of, or aspects of, the guided diffusion model, which relates to Fig. 15. In some embodiments, the document processing device 1300 comprises the processing unit 1305, the I / O module 1310, the user interface 1325, the storage unit 1315, the document processing model 1320, and the training unit 1345. The training unit 1345 updates parameters of the language generation model 1335, which is stored in the storage unit 1315. In some examples, the training unit 1345 is located outside the document processing device 1300.
[0088] The Processing Unit 1305 comprises one or more processors. A processor is an intelligent hardware unit, for example, a general-purpose processing unit, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device (PLD), a logic gate or transistor circuit, or any combination thereof.
[0089] In some cases, the 1305 processing unit is configured to operate a memory matrix via a memory controller. In other cases, a memory controller is integrated into the 1305 processing unit. In some cases, the 1305 processing unit is configured to execute computer-readable instructions stored in the 1315 memory unit to perform various functions described herein. In some aspects, the 1305 processing unit includes specialized components for modem processing, baseband processing, digital signal processing, or transmission processing. According to some aspects, the 1305 processing unit includes one or more processors, as described in relation to Fig. 31 are described.
[0090] The Memory Unit 1315 comprises one or more memory devices. Examples of a memory device include random-accessible transitive memory (RAM), read-only memory (ROM), or a hard disk drive. Examples of memory devices include semiconductor memory and hard disk drives. In some examples, the memory is used to store computer-readable, computer-executable software, including instructions that, when executed, cause at least one processor of the Processing Unit 1305 to perform various functions described herein.
[0091] In some cases, the 1315 memory unit includes a Basic Input / Output System (BIOS) that controls basic hardware or software operations, such as interaction with peripheral devices or components. In some cases, the 1315 memory unit includes a memory controller that operates the memory cells of the 1315 memory unit. For example, the memory controller may include a row decoder, a column decoder, or both. In some cases, memory cells in the 1315 memory unit store information in the form of a logical state. From some perspectives, the 1315 memory unit is an example of the 3110 memory subsystem, which, in terms of Fig. 31 is described.
[0092] According to some aspects, the Document Processing Device 1300 uses one or more processors of the Processing Unit 1305 to execute instructions stored in the Memory Unit 1315 to perform the functions described herein. For example, the Document Processing Device 1300 can receive an image and a style guide, where the image shows an object with a first color and the style guide contains a second color. The Document Processing Device 1300 generates a first color embedding and a second color embedding based on the first color and the second color, respectively. The Document Processing Device 1300 selects the second color from the style guide by comparing the first color embedding of the first color with the second color embedding of the second color.The document processing device 1300 generates a modified image based on the image and the second color, with the modified image showing the object with the second color.
[0093] In some embodiments, the document processing model 1320 is an artificial neural network (ANN) such as the one described in relation to Fig. 15 described a guided diffusion model. An ANN can be a hardware or software component comprising connected nodes (i.e., artificial neurons) that loosely correspond to the neurons in a human brain. Each connection or edge carries a signal from one node to another (similar to the physical synapses in the brain). When a node receives a signal, it processes the signal and transmits it to other connected nodes.
[0094] ANNs possess numerous parameters, including weights and thresholds (biases), assigned to each neuron in the network. These parameters determine the strength of the connections between neurons and influence the neural network's ability to recognize complex patterns in data. Also known as model parameters or model weights, these variables define the behavior and properties of a machine learning model.
[0095] In some cases, the signals between nodes consist of real numbers, and the output of each node is calculated as a function of its inputs. For example, nodes can determine their output using other mathematical algorithms (such as selecting the maximum of the inputs as the output) or any other suitable algorithm. Each node and connection is assigned one or more weights that determine how the signal is processed and forwarded. In some cases, nodes have a threshold below which a signal is not forwarded at all. In some examples, the nodes are grouped into layers.
[0096] The parameters of the Document Processing Model 1320 can be organized into layers. Different layers perform different transformations on their inputs. The first layer is called the input layer, and the last layer is called the output layer. In some cases, signals traverse certain layers multiple times. A hidden (or middle) layer contains hidden nodes and is located between an input layer and an output layer. Hidden layers perform nonlinear transformations on the inputs. Each hidden layer is trained to produce a defined output that contributes to the overall output of the neural network's output layer. Hidden representations are machine-readable data representations of an input, learned from the hidden layers of the neural network and generated by the output layer.As the neural network's understanding of the input improves through training, the hidden representations increasingly differ from previous iterations.
[0097] Training Unit 1345 can train Document Processing Model 1320. For example, parameters of Document Processing Model 1320 can be learned or estimated from training data and then used to make predictions or perform tasks based on learned patterns and relationships in the data. In some examples, the parameters are adjusted during the training process to minimize a loss function or maximize a performance metric (such as in relation to...). Fig. 30). The goal of the training process is to find optimal values for the parameters that enable the machine learning model to make accurate predictions or to perform well in the given task.
[0098] Accordingly, node weights can be adjusted to increase the accuracy of the output (i.e., by minimizing any loss that reflects the difference between the current result and the target value). A link's weight increases or decreases the strength of the transmitted signal. For example, during the training process, an algorithm adjusts the model parameters to minimize the error or loss between the predicted outputs and the actual target values, according to optimization techniques such as gradient descent, stochastic gradient descent, or other optimization algorithms. Once the model parameters have been learned from the training data, the document processing model 1320 can be used to make predictions on new, previously unseen data (i.e., during the inference phase).
[0099] The I / O module 1310 receives inputs for the document processing device 1300 and transmits outputs from the document processing device 1300 to other devices or users. For example, the I / O module 1310 receives inputs for the document processing model 1320 and transmits outputs from the document processing model 1320. In some respects, the I / O module 1310 is an example of the I / O interface 3120, which, in relation to Fig. 31 is described.
[0100] In some embodiments, the document processing model 1320 receives a document containing the image. In some examples, the document processing model 1320 produces a modified document containing the modified image. In some examples, the document processing model 1320 receives a video with a series of frames, where the image comprises one frame of this series of frames.
[0101] According to some embodiments, the document processing model 1320 generates a modified document containing the modified image and the modified text element. In one aspect, the document processing model 1320 includes the user interface 1325, the style guide engine 1330, the language generation model 1335, and the image generation model 1340.
[0102] The user interface 1325 is an example of, or includes, aspects of the corresponding element, which relates to Fig. is described in 3-11, 18, 19 and 26-29.
[0103] In some examples, a style guide includes a font, a text color, a background color, a logo, or any combination thereof. In some examples, the Style Guide Engine 1330 applies the style guide to a group of pages in the document based on a page selection input. In some examples, the Style Guide Engine 1330 applies a first style attribute from the style guide to a first element in the document. In some examples, the Style Guide Engine 1330 applies a second style attribute from the style guide to a second element in the document. In some examples, the Style Guide Engine 1330 applies a font from the style guide to a text element in the document. In some examples, the Style Guide Engine 1330 applies the style guide to a group of frames in a video.
[0104] In some embodiments, the Style Guide Engine 1330 applies the font from the style guide to the text element to produce a modified text element. In some examples, the Style Guide Engine 1330 applies the second color from the style guide to the image to produce a modified image, the modified image showing the object with the second color. In some examples, the Style Guide Engine 1330 applies an additional font, different from the specified font, from the style guide to an additional text element in the document to produce an additional modified text element, the modified document containing the additional modified text element. In some examples, the Style Guide Engine 1330 applies the style guide to a group of pages in the document based on page selection input.In some examples, the Style Guide Engine 1330 applies a first style attribute from the style guide to the first element of the document. In other examples, the Style Guide Engine 1330 applies a second style attribute from the style guide to a second element of the document.
[0105] According to some embodiments, the language generation model 1335 generates a first color embedding and a second color embedding based on the first color and the second color, respectively. In some examples, the language generation model 1335 selects the second color from the style guide by comparing the first color embedding of the first color with the second color embedding of the second color. In some examples, the language generation model 1335 generates a first text description of the first color in the image, with the first color embedding being generated based on the first text description. In some examples, the language generation model 1335 generates a second text description of the second color in the style guide, with the second color embedding being generated based on the second text description.
[0106] According to some embodiments, the language generation model 1335 generates an initial text description of the first color in the image. In some examples, the language generation model 1335 generates an initial color embedding based on the initial text description. In some examples, the language generation model 1335 generates a second text description of the second color in the style guide. In some examples, the language generation model 1335 generates a second color embedding based on the second text description.
[0107] According to some embodiments, the language generation model 1335 generates a first color embedding and a second color embedding based on the first color and the second color, respectively. In some examples, the language generation model 1335 selects the second color from the style guide by comparing the first color embedding of the first color with the second color embedding of the second color.
[0108] According to some embodiments, the image generation model 1340 generates a modified image based on the image and the second color, wherein the modified image represents the object with the second color.
[0109] According to some embodiments, the image generation model 1340 generates a modified image based on the image and the second color, wherein the modified image represents the object with the second color. In some examples, the image generation model 1340 generates the modified image by applying the second color to the object.
[0110] Fig. Figure 14 shows an example of a transformer network according to aspects of the present disclosure. The example shown includes the transformer 1400, the encoder 1405, the decoder 1420, the input 1440, the input embedding 1445, the input position encoding 1450, the previous output 1455, the previous output embedding 1460, the previous output position encoding 1465, and the output 1470.
[0111] In some cases, the encoder 1405 includes a multi-head self-attention sublayer 1410 and a feed-forward network sublayer 1415. In some cases, the decoder 1420 includes a first multi-head self-attention sublayer 1425, a second multi-head self-attention sublayer 1430, and a feed-forward network sublayer 1435.
[0112] According to some aspects, a machine learning model (such as the one in relation to Fig. (Model 13 described) the transformer 1400. In some cases, the encoder 1405 converts the input 1440 (for example, a query or prompt consisting of a sequence of words or tokens) into a sequence of continuous representations, which are fed to the decoder 1420. In some cases, the decoder 1420 generates the output 1470 (for example, a predicted output sequence of words or tokens) based on the output of the encoder 1405 and the previous output 1455 (for example, a previously predicted output sequence), thus enabling autoregressive prediction.
[0113] For example, encoder 1405 decomposes input 1440 into tokens and vectorizes these tokens to obtain input embedding 1445, and adds the input positional encoding 1450 (for example, positional encoding vectors for input 1440 with the same dimension as input embedding 1445) to input embedding 1445. In some cases, the input positional encoding 1450 contains information about the relative positions of words or tokens in input 1440.
[0114] In some cases, the Encoder 1405 comprises one or more coding layers (for example, six layers) that generate contextualized token representations, each representation corresponding to a token and combining information from other tokens using self-attention. In some cases, each coding layer of the Encoder 1405 includes a multi-head self-attention sublayer (for example, sublayer 1410). In some cases, the multi-head self-attention sublayer implements a multi-head self-attention mechanism that receives different linearly projected versions of queries, keys, and values to generate outputs in parallel. In some cases, each coding layer of the Encoder 1405 also includes a fully connected feed-forward network sublayer (for example, sublayer 1415) that includes two linear transformations surrounded by a rectified linear unit (ReLU) activation. FFN(x)=ReLU(W1x+b1)W2+b2
[0115] In some cases, each layer applies different weight parameters (W1, W2) and different bias parameters (b1, b2) to apply an equal linear transformation to each word or token in input 1440.
[0116] In some cases, each sublayer of the encoder 1405 is followed by a normalization layer that normalizes a sum calculated between a sublayer input x and an output sublayer (x) generated by the sublayer: layernorm(x+sublayer(x))
[0117] In some cases, the encoder 1405 is bidirectional, as it pays attention to each word or token in the input 1440, regardless of the position of the word or token in the input 1440.
[0118] In some cases, the Decoder 1420 comprises one or more decoding layers (for example, six layers). In some cases, each decoding layer comprises three sublayers: a first multi-head self-attention sublayer 1425, a second multi-head self-attention sublayer 1430, and a feed-forward network sublayer 1435. In some cases, each sublayer of the Decoder 1420 is followed by a normalization layer that normalizes a sum calculated from the sublayer input x and an output sublayer generated by the sublayer, sublayer(x).
[0119] In some cases, the Decoder 1420 generates a previous output embedding 1460 of the previous output 1455 and adds the previous output position encoding 1465 (for example, position information for words or tokens in the previous output 1455) to the previous output embedding 1460. In some cases, each first multi-head self-attention sublayer of the Decoder 1420 receives the combination of the previous output embedding 1460 and the previous output position encoding 1465 and applies a multi-head self-attention mechanism to this combination. In some cases, each first multi-head self-attention sublayer of the Decoder 1420, for each word in an input sequence, pays attention only to the words that precede that word in the sequence, so that the Transformer 1400's prediction for a word at a given position depends only on the known outputs for the preceding words in the sequence.For example, in some cases each first multi-head self-attention sublayer implements several single attention functions in parallel by placing a mask over the values generated by the scaled multiplication of matrices Q and K, thereby suppressing matrix values that correspond to invalid connections.
[0120] In some cases, every second multi-head self-attention sublayer implements a multi-head self-attention mechanism similar to the multi-head self-attention mechanism implemented in every multi-head self-attention sublayer of the encoder 1405, by receiving a query Q from a previous sublayer of the decoder 1420 and a key K and a value V from the output of the encoder 1405, enabling the decoder 1420 to pay attention to each word in the input 1440.
[0121] In some cases, each feed-forward network sublayer implements a fully connected feed-forward network structure similar to feed-forward network sublayer 1415. In some cases, the feed-forward network sublayers are followed by a linear transform and a softmax function to generate a prediction of output 1470 (for example, a prediction of the next word or token in a sequence). Accordingly, in some cases, transformer 1400 generates a response based on a predicted sequence of words or tokens, as described herein.
[0122] Fig. Figure 15 shows an example of a guided diffusion model according to aspects of the present disclosure. The guided latent diffusion model 1500 in Fig. 15 is an example of, or includes, aspects of the corresponding element (i.e., the language generation model 1335), which relates to Fig. 13 is described.
[0123] Diffusion models are a class of generative neural networks that can be trained to generate new data with features similar to those of the training data. In particular, diffusion models can be used to generate novel images. Diffusion models can be employed for various image generation tasks, including super-resolution image generation, generation of images with perceptual metrics, conditional generation (for example, generation based on text specifications), image inpainting, and image manipulation.
[0124] Types of diffusion models include Denoising Diffusion Probabilistic Models (DDPMs) and Denoising Diffusion Implicit Models (DDIMs). In DDPMs, the generation process involves inverting a stochastic Markov diffusion process. DDIMs, on the other hand, use a deterministic process, so the same input leads to the same output. Diffusion models can also be distinguished by whether noise is added to the image itself or to features of an image generated by an encoder (i.e., latent diffusion).
[0125] Diffusion models work by iteratively adding noise to the data during a forward process and then learning to reconstruct the data by denoising it during a reverse process. For example, the guided latent diffusion model 1500 can take an original image 1505 in a pixel space 1510 as input during training and apply an image encoder 1515 to convert the original image 1505 into original image features 1520 in a latent space 1525. Then, a forward diffusion process 1530 incrementally adds noise to the original image features 1520 to obtain noisy features 1535 (also in latent space 1525) at different noise levels.
[0126] Subsequently, a backward diffusion process removes 1540 (for example, a U-Net ANN or a DiT architecture as in Fig. (described in 32) the noise is gradually removed from the noisy features 1535 at the different noise levels to obtain denoised image features 1545 in the latent space 1525. In some examples, the denoised image features 1545 at each of the different noise levels are compared with the original image features 1520, and parameters of the reverse diffusion process 1540 of the diffusion model are updated based on the comparison. Finally, an image decoder 1550 decodes the denoised image features 1545 to obtain an output image 1555 in pixel space 1510. In some cases, an output image 1555 is generated at each of the different noise levels. The output image 1555 can be compared with the original image 1505 to train the reverse diffusion process 1540.
[0127] In some cases, the image encoder 1515 and the image decoder 1550 are pre-trained before the reverse diffusion process 1540 is trained. In some examples, the image encoder 1515 and the image decoder 1550 are trained together, or the image encoder 1515 and the image decoder 1550 are fine-tuned together with the reverse diffusion process 1540.
[0128] The reverse diffusion process 1540 can also be based on a text prompt 1560 or another control prompt such as an image, a layout, a segmentation map, etc. The text prompt 1560 can be encoded using a text encoder 1565 (for example, a multimodal encoder) to obtain guide features 1570 in a guide space 1575. The guide features 1570 can be combined with the noisy features 1535 on one or more layers of the reverse diffusion process 1540 to ensure that the output image 1555 contains content described by the text prompt 1560. For example, the guide features 1570 can be combined with the noisy features 1535 using a cross-attention block within the reverse diffusion process 1540.
[0129] Fig. Figure 16 shows an example of a color application interface 1600 according to aspects of the present disclosure. Fig. Figure 16 shows that users can add or remove candidate colors from a style guide. In some examples, a style guide includes one or more fonts, one or more text colors, one or more background colors, one or more image colors, one or more images, or any combination thereof. The style guide includes the color palette. The color palette is applied to pages of a document to create a modified document. In some cases, a user selects a group of candidate colors to form a color palette as part of the style guide.
[0130] In the Fig. In the example shown in Figure 16, a user edits a color palette and its associated settings using the Color 1600 application interface. The Color 1600 application interface is a graphical user interface that includes a dialog box labeled "Add Color." The Color 1600 application interface includes the "Swatches" and "Custom" tabs, each representing a method for color selection. For example, the "Swatches" tab displays predefined color options and a "Recommended" area. The "Custom" tab allows for individual color selection. The Color 1600 application interface includes a color swatch selection tool that allows users to select colors from a color swatch (i.e., access to a wide range of colors). Users manage the color selection process using the "Cancel" and "Save" buttons.
[0131] Fig. Figure 17 shows an example of a font application interface 1700 according to aspects of the present disclosure. The font application interface 1700 is a graphical user interface with a dialog box that includes a search bar. The search bar at the top of the font application interface 1700 is used to locate one or more fonts. The font application interface 1700 includes a first section called "Recent" and a second section called "Your fonts." The first section displays recently used fonts. The second section displays fonts specified by the user. For example, the "Recent" section includes fonts such as Anton Regular and PT Serif Regular. Users can click "view more" to display additional fonts. The "Your fonts" section categorizes fonts, for example, Abolition and Abril Display.The Font 1700 application interface provides a font preview for each typeface using the text "The quick brown fox". The Font 1700 application interface includes interactive elements for uploading additional fonts and accessing a wide selection of fonts by clicking the "More fonts" button. This increases efficiency in font selection and customization.
[0132] Fig. Figure 18 shows an example of a style transformation element 1805 and a style guide setting element 1810 according to aspects of the present disclosure. The example shown includes the user interface 1800, the style transformation element 1805, and the style guide setting element 1810. The user interface 1800 is an example of, or includes, aspects of the corresponding element, which relates to Fig. is described in 3-11, 13, 19 and 26-29.
[0133] Fig. Figure 18 shows an enlarged view of a control panel in the left pane of user interface 1800. In some examples, the "recolor graphics" setting applies to a single page of a document. The "recolor graphics" setting can be disabled for a multi-page document. In one embodiment, the style guide setting element 1810 includes a page selector element 1815, a color application parameter 1820, a font application parameter 1825, and an image recoloring parameter 1830.
[0134] In the Fig. In example 18, a style guide application tool in user interface 1800 is used to apply a style guide (for example, a brand or a collection of brand-related assets) to multiple pages of a document. The available settings include "apply colors," "apply fonts," and "apply to all pages." Unlike Fig. In version 19, the "apply to all pages" option is enabled in the Style Guide application tool because the document has multiple pages. In some examples, the "apply to all pages" setting causes the colors and fonts to be applied to all pages of the document (for example, a presentation or an Instagram® Story). The Style Guide application tool in user interface 1800 ensures that the colors and fonts are applied consistently across all pages of the document. For example, if a presentation has multiple slides with a red background and green foreground, the brand colors would be replaced in the same way on all slides (for example, blue brand background, dark red brand foreground).
[0135] The style transformation element 1805 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 3, 4, 6-11, 19, 26 and 27. The Style Guide Setting Element 1810 is an example of, or includes aspects of, the corresponding element, which relates to Fig. 3 and Fig. 19 is described. The page selection element 1815 is an example of, or includes aspects of, the corresponding element, which relates to Fig. 19 is described. The color application parameter 1820 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 19 is described. The font application parameter 1825 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 19 is described. The image recoloring parameter 1830 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 19 is described.
[0136] Fig. Figure 19 shows an example of a user interface 1900 according to aspects of the present disclosure. The example shown includes the user interface 1900, the style transformation element 1905, and the style guide setting element 1910. The user interface 1900 is an example of, or includes, aspects of the corresponding element, which relates to Fig. is described in 3-11, 13, 18 and 26-29.
[0137] In one embodiment, the style guide setting element 1910 comprises a page selection element 1915, a color application parameter 1920, a font application parameter 1925, and an image recoloring parameter 1930.
[0138] In the Fig. The example shown in Figure 19 illustrates the user-selectable fields, or style guide settings, in the 1900 user interface. In some cases, these style guide settings are also referred to as "brand settings." To apply a style guide to a page of a document, users can click a style guide application tool that appears in the 1900 user interface. The available style guide settings include "apply colors," "apply fonts," and "recolor graphics." In this example, the "apply to all pages" option of the style guide application tool is disabled because the document contains only a single page.
[0139] The style transformation element 1905 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 3, 4, 6-11, 18, 26 and 27. The Style Guide Setting Element 1910 is an example of, or includes aspects of, the corresponding element, which relates to Fig. 3 and Fig. 18 is described. The page selection element 1915 is an example of, or includes aspects of, the corresponding element, which relates to Fig. 18 is described. The color application parameter 1920 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 18 is described. The font application parameter 1925 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 18 is described. The image recoloring parameter 1930 is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 18 is described.
[0140] Fig. Figure 20 shows an example of a style guide according to aspects of the present disclosure. The example shown includes candidate colors 2000, a first typeface 2005, a second typeface 2010, font editing tools 2015, and scroll settings 2020.
[0141] Fig. Figure 20 shows an example of login and save functions (for example, saved colors, saved fonts). Any number of colors can be added to a style guide (for example, a collection of branded assets), but companies typically have a maximum of 5 to 8 colors in their brand identity for a marketing campaign. The brand colors are used consistently across all digital media to give customers a clear picture of the company's brand. For example, one company might use green and white everywhere, while another company might use red everywhere. Everything—from their digital application and websites to printed brochures and media—uses the same branded color scheme. In an example in Fig. 20 of the colors are associated with a brand for a mountain sports clothing company.
[0142] A font is assigned the role of "Header," "Body," or "None." A style guide encompasses multiple fonts. With a single click on "Apply brand," the Document Processing Model 1320 applies the style guide as described in [reference to relevant document]. Fig. 13 describes applying a combination of three font types (for example, Header, Body and None) from the style guide associated with a brand to the document.
[0143] Additionally or alternatively, a style guide includes digital assets such as logos, templates, digital images, etc. These brand-related assets can be reused across pages of a target document. These assets are optional for the "Apply brand" function in the user interface.
[0144] Candidate colors 2000 are an example of, or encompass, aspects of the corresponding element that relate to Fig. 3, 4 and 7-11 are described. The first font, from 2005, is an example of, or includes, aspects of the corresponding element, which relates to Fig. 5, Fig. 6, Fig. 26 and Fig. 27 is described. The second font, 2010, is an example of, or includes, aspects of the corresponding element, which relates to Fig. 5, Fig. 6, Fig. 26 and Fig. 27 is described.
[0145] In Fig. Sections 13-20 describe an image processing device, a system, and a method. One or more aspects of the device, system, and method include: a storage unit; a processing device coupled to the storage unit and configured to perform the following: obtaining an image and a style guide, wherein the image shows an object with a first color and the style guide contains a second color; generating a first color embedding and a second color embedding based on the first color and the second color, respectively; selecting the second color from the style guide by comparing the first color embedding of the first color with the second color embedding of the second color; and generating a modified image based on the image and the second color, wherein the modified image shows the object with the second color.
[0146] Some examples of the device, system, and method further include a language generation model configured to generate the first color embedding and the second color embedding. Some examples of the device, system, and method further include an image generation model configured to produce the modified image by applying the second color to the object.
[0147] Some examples of the device, system, and method further include providing a style transformation element in a user interface. Some examples further include receiving a single click input via the style transformation element, whereby the modified image is generated based on the single click input. Style Guide Application
[0148] Fig. Figure 21 shows an example of a method 2100 for image processing according to aspects of the present disclosure. In some examples, these operations are performed by a system comprising a processor that executes a set of code for controlling functional elements of a device. Additionally or alternatively, certain processes are performed using special hardware. In general, these operations are performed according to the methods described herein. In some cases, the operations described herein consist of several sub-steps or are performed in conjunction with other operations.
[0149] At step 2105, the system generates an initial text description of the first color in the image. For example, the initial text description of the first color is "Bright saturated red color in the foreground." This initial text description is also referred to as the color description string. Further examples of color text descriptions are related to... Fig. 22 described. In some cases, the operations of this step refer to a language generation model, as in relation to Fig. 13 described.
[0150] In step 2110, the system generates an initial color embedding based on the first text description. In some examples, the initial color embedding is a representation of the first color in a vector space. In other cases, the operations of this step relate to a language generation model, as in the case of Fig. 13. The process for generating a color embedding is described in relation to Fig. described on pages 22-23.
[0151] In step 2115, the system generates a second text description for the second color in the style guide. For example, the second text description for the second color is "Dark professional blue associated with trust and stability." The second color comes from a pre-configured brand palette. Further examples of brand color descriptions are related to… Fig. 23 described. In some cases, the operations of this step refer to a language generation model, as in relation to Fig. 13 described.
[0152] In step 2120, the system generates a second color embedding based on the second text description. In some examples, the second color embedding is a representation of the second color in a vector space. The second color embedding can also be referred to as the brand color embedding. In some cases, the operations of this step relate to a language generation model, as in the case of Fig. 13. The process for generating a brand color embedding is described in relation to Fig. described on pages 22-23.
[0153] Fig. Figure 22 shows an example of an algorithm 2200 using a Sentence Transformer according to aspects of the present disclosure.
[0154] In some embodiments, a color matching network (also known as language generation model 1335, as in relation to Fig. (described in section 13) is trained to understand the meaning and context of both ECS artwork and brand guidelines, and to provide intelligent color recommendations. The color matching network balances artistic freedom with brand consistency by enabling smoother adjustments rather than rigid transformations.
[0155] In some examples, the color matching network uses text-based embeddings to improve how colors are understood, represented, and matched. Instead of using conventional one-hot encoding or direct numerical representations / clusters of colors (for example, RGB), the color matching network treats colors as semantic concepts and uses LLM-based embeddings to capture the relationships between colors. By creating a textual representation of the color data and using an LLM (for example, HuggingFace's Sentence Transformers library) to generate embeddings, embodiments of the present disclosure can enhance the brand matching process with deeper context and understanding of color relationships.
[0156] In one embodiment, the color matching network treats colors as descriptive features by converting color information into text (rather than treating each color as an isolated data point). The text strings describe not only the raw color values but also attributes such as brightness, hue, emotional context, and spatial hierarchy. The color matching network generates a text string that describes the color information of each object and other relevant properties. This text string is fed to a pre-trained LLM to generate a high-dimensional embedding that captures the relationships between the colors. For example (with regard to a color descriptor string), we suppose an object in the ECS has the RGB color (255, 0, 0) and is located in the foreground (Layer 2) with high brightness and saturation.The color matching network can describe this object as "Bright saturated red color in the foreground." The sentence can include: (1) the color name or a description (for example, "red," "dark blue"); (2) descriptions of brightness and saturation (for example, "bright," "muted"); (3) the location / layer in the visual hierarchy (for example, "foreground," "background"); and (4) emotional tone or implicit mood (for example, "warm," "calm").
[0157] The color matching network generates a dense vector representation (embedding) for each color, capturing more than just the raw RGB values. It captures the relationships and context between the colors, as well as their role in the visual and emotional hierarchy of the artwork.
[0158] Fig. Figure 23 shows an example of an algorithm 2300 for calculating a similarity value for color matching according to aspects of the present disclosure. The similarity value can be used to determine an approximation criterion for selecting a color from a style guide. For example, the approximation criterion may be to determine that a color is below a certain similarity value, or it may be determined by maximizing the similarity value or minimizing a cosine distance.
[0159] The color matching network uses the color embeddings generated by the LLM to enable a more nuanced and context-aware brand matching process. Instead of simply matching raw RGB values with brand colors, the color matching network compares the semantic embeddings of the extracted artwork colors with the embeddings of the brand colors, thus enabling flexible, context-aware matching. First, the color matching network generates a brand palette representation by converting each brand color into a text string that describes not only the color, but also the tone and identity associated with it.
[0160] For example, (brand color descriptor) the color matching network for a brand color "dark blue" generates the description "Dark professional blue associated with trust and stability". Next, the color matching network performs an embedding comparison; that is, using the embeddings for both the extracted artwork colors and the brand colors, the color matching network calculates the cosine similarity between the embeddings to determine how close a given artwork color is to a brand color (not just in RGB space, but in semantic space).
[0161] By calculating the cosine similarity to the color match, the color matching network calculates a similarity value between each artwork color and each brand color, which is used to find the closest match or to suggest slight adjustments to bring the artwork closer to the brand color identity.
[0162] In some examples, where brand color attributes can be adjusted, the color matching network can generate variations of brand color descriptions by changing attributes such as brightness, saturation, or context to see if they result in higher similarity scores. The color matching network then generates embeddings for these variations and includes them in the matching process.
[0163] In some examples, the color matching network sorts the brand colors (or their variations) for each artwork color based on similarity scores from highest to lowest. The color matching network provides a ranking of potential matches, allowing users to select the best one or consider alternative groupings.
[0164] Fig. Figure 24 shows an example of a textual description of a color according to aspects of the present revelation. In some examples, textual descriptions are exemplary inputs for a language generation model, as in relation to Fig. 13 described. Fig. 24 is an example of an approximation criterion for selecting a color from a style guide. Table 1. Example output ArtworkColor Descriptor: Bright red with high saturation and a warm emotional tone. Dominant in the foreground. Top Matcches: Match 1: Brand Color Descriptor: Bright orange with very high brightness and an energeticemotional tone. SimilarityScore: 0.9021 Match 2: Brand Color Descriptor: Energetic orange with high brightness and a vibrantemotional tone. SimilarityScore: 0.8765 Match 3: Brand Color Descriptor: Deep blue with low brightness and a serious emotional tone. SimilarityScore: 0.6543 Table 2. Example output ArtworkColor Descriptor: Soft blue with medium brightness and a calm emotional tone. Used in the background. Top Matches: Match 1: Brand Color Descriptor: Corporate blue with medium brightness and a professional emotional tone. SimilarityScore: 0.9123 Match 2: Brand Color Descriptor: Deep blue with low brightness and a serious emotional tone. SimilarityScore: 0.8567 Match 3: Brand Color Descriptor: Soft green with low saturation and a peaceful emotional tone. SimilarityScore: 0.7345 Table 3. Example output. ArtworkColor Descriptor: Vibrant green with high saturation and an energetic emotional tone. Accents in the midground. Top Matches: Match 1:Brand Color Descriptor: Trustworthy green with medium saturation and a calming emotional tone. SimilarityScore: 0.8789 Match 2:Brand Color Descriptor: Soft green with low saturation and a peaceful emotional tone. SimilarityScore: 0.8123 Match 3:Brand Color Descriptor: Energetic orange with high brightness and a vibrant emotional tone. SimilarityScore: 0.6987
[0165] In one implementation, the color matching network preserves the artistic core by adjusting colors while maintaining the visual essence of the artwork. The network offers contextual awareness, as colors are transformed based on their contextual role within the artwork (for example, logo vs. background). The network generates flexible and consistent results, ensuring brand consistency while providing sufficient flexibility for non-critical elements, thus balancing rigor with creative freedom.
[0166] By treating color transformation similarly to sentence transformation, embodiments of the present disclosure offer a sophisticated and flexible system for matching brand colors in artwork. The color matching network can preserve the meaning or essence of the original colors (just as sentence transformations preserve meaning in text). Through vector color coding, context-aware transformations, and adaptive flexibility, the color matching network improves upon rigid color matching systems.
[0167] Fig. Figure 25 shows an example of a method 2500 for image processing according to aspects of the present disclosure. In some examples, these operations are performed by a system comprising a processor that executes a set of code for controlling functional elements of a device. Additionally or alternatively, certain processes are performed using special hardware. In general, these operations are performed according to the methods described herein. In some cases, the operations described herein consist of several sub-steps or are performed in conjunction with other operations.
[0168] In step 2505, the system receives a document and a style guide, where the document contains a text element with a first font and an image displaying an object with a first color, and where the style guide contains a second font and a second color. An example of a document is document 705, which is in Fig. 7 is described. A text element is an example of, or comprises, aspects of the corresponding element that relate to Fig. 5, Fig. 7, Fig. 26 and Fig. 28 describes, for example the first text element 525 in Fig. 5. An example of an image is in Fig. 3 described, i.e., Figure 305. The style guide is an example of, or includes, aspects of the corresponding element, which relates to Fig. 3-8, 10-11, 18-19, and 26-27 are described. The second color differs from the first color. In some cases, the operations of this step relate to, or can be performed by, a user interface, as described in relation to Fig. described in 3-11, 13, 18-19 and 26-29.
[0169] In step 2510, the system applies the second font from the style guide to the text element to create a modified text element. The modified text element is an example of, or incorporates aspects of, the corresponding element, which is related to... Fig. 6, Fig. 8, Fig. 27 and Fig. 29 is described. An example of the modified text element is in Fig. 6 described, i.e., the first text element 625. In some cases, the operations of this step refer to a style guide engine, as in relation to Fig. 13 described.
[0170] In step 2515, the system uses an image generation model to apply the second color from the style guide to the image, resulting in a modified image where the modified image shows the object with the second color. The modified image is an example of, or includes aspects of, the corresponding element that, in relation to Fig. 4, Fig. 8, Fig. 27 and Fig. 29 describes, for example the modified image 405 in Fig. 4. In some cases, the operations of this step relate to a style guide engine, as in relation to Fig. 13 described.
[0171] In step 2520, the system generates a modified document that includes the modified image and the modified text element. The modified document is an example of, or includes aspects of, the corresponding element, which relates to Fig. 4, Fig. 6, Fig. 8, Fig. 27 and Fig. 29 describes, for example, the modified document 805 in Fig. 8. In some cases, the operations of this step relate to a document processing model, such as in relation to Fig. 13 described.
[0172] Fig. Figure 26 shows an example of a style guide with font selection according to aspects of the present disclosure. The example shown includes the user interface 2600, the document 2605, the style transformation element 2625, and candidate fonts 2630. The user interface 2600 is an example of, or includes, aspects of the corresponding element, which relates to Fig. 3-11, 13, 18, 19 and 27-29 are described.
[0173] Fig. Figure 26 shows a page of a document before a font is applied to the document via the user interface 2600. In some examples, the document 2605 includes a first text element 2610, a second text element 2615, and a third text element 2620. The candidate fonts 2630 include a first font 2635, a second font 2640, a third font 2645, and a fourth font 2650.
[0174] Document 2605 is an example of, or includes, aspects of the corresponding element relating to Fig. 5, Fig. 7 and Fig. 10 is described. The first text element 2610 is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 6 and Fig. 7 is described. The second text element 2615 is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 6 and Fig. 7 is described. The third text element 2620 is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 6 and Fig. 7 is described.
[0175] The style transformation element 2625 is an example of, or includes aspects of, the corresponding element, which relates to Fig. 3-4, 6-11, 18, 19 and 27. The candidate typefaces 2630 are an example of, or comprise, aspects of the corresponding element, which relates to Fig. 3-4 and 27 are described.
[0176] The first font, 2635, is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 5-6, 20 and 27. The second font, 2640, is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 5-6, 20 and 27. The third font, 2645, is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 27 is described. The fourth typeface, 2650, is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 27 is described.
[0177] Fig. Figure 27 shows an example of an effect of applying a typeface according to aspects of the present disclosure. The example shown includes the user interface 2700, the modified document 2705, the style transformation element 2725, and candidate typefaces 2730. The user interface 2700 is an example of, or includes, aspects of the corresponding element, which relates to Fig. 3-11, 13, 18, 19, 26, 28 and 29 are described.
[0178] In some examples, the modified document 2705 includes a first modified text element 2710, a second modified text element 2715, and a third modified text element 2720. The candidate fonts 2730 include a first font 2735, a second font 2740, a third font 2745, and a fourth font 2750.
[0179] Fig. 27 shows a modified side of the in Fig. 26 mentioned documents after applying a font to the document via a single click on the "Apply brand" button in user interface 2700. The document processing model 1320 (as in Fig. (Described in section 13) assigns a font from the style guide (displayed in the left pane of user interface 2700) to a corresponding text segment on the document page (i.e., according to size). In some examples, a first font 2735, marked as "Header" in the style guide (i.e., font for headings), is applied to the largest text on the document page. A second font 2740, with the "Body" role in the style guide, is applied to the second largest text. A third font 2745, marked "None," is applied to the remaining text on the document page. A style guide or mark can include multiple fonts with the same role. For example, a user might select two heading fonts and three body text fonts. The two heading fonts might include, for example, "Header Clean Black" and "Header Clean ExtraBold."The three body text fonts are "Body Clean Italic," "Body Clean Bold," and "Body Clean Regular." The style guide (including the heading and body text fonts) is displayed in the left-hand panel of the 2700 user interface. Clicking the "Apply brand" (shuffle) button again in the 2700 user interface would apply different variations and generate different modified documents.
[0180] Document processing model 1320 receives a selection parameter that corresponds to a style attribute from the style guide. The style attribute is then applied to the document based on the selection parameter to produce a modified document. In some examples, both the document in Fig. 26 as well as the modified document in Fig. 27 each a multimedia asset.
[0181] The modified document 2705 is an example of, or includes, aspects of the corresponding element relating to Fig. 6, Fig. 8, Fig. 9 and Fig. 11 is described. The first modified text element 2710 is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 8 and Fig. 9 is described. The second modified text element 2715 is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 8 and Fig. 9 is described. The third modified text element 2720 is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 8 and Fig. 9 is described.
[0182] The style transformation element 2725 is an example of, or includes aspects of, the corresponding element, which relates to Fig. 3, 4, 6-11, 18, 19 and 26. The candidate typefaces 2730 are an example of, or comprise, aspects of the corresponding element, which relates to Fig. 3, Fig. 4 and Fig. 26 is described.
[0183] The first font, 2735, is an example of, or encompasses, aspects of the corresponding element, which relates to Fig. 5, Fig. 6, Fig. 20 and Fig. 26 is described. The second font, 2740, is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 5, Fig. 6, Fig. 20 and Fig. 26 is described. The third typeface, 2745, is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 26 is described. The fourth typeface, 2750, is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 26 is described.
[0184] Fig. Figure 28 shows an example of a state-change effect according to aspects of the present disclosure. The example shown includes the user interface 2800, the first image 2805, and a state-change element 2810. The user interface 2800 is an example of, or includes, aspects of the corresponding element, which relates to Fig. 3-11, 13, 18, 19, 26, 27 and 29. The state change element 2810 is an example of, or comprises, aspects of the corresponding element, which relates to Fig. 29 is described.
[0185] The 2800 user interface includes an Undo button and a Redo button in the upper right corner. Undo and Redo actions on the workspace allow you to undo or reapply a style guide effect with a single click. In some cases, this prevents Fig. 13 described document processing model 1320 Changes to elements that users do not want to have changed by a brand application. By clicking the "Apply brand" button again in the user interface 2800, different variations of the brand are shuffled.
[0186] The document shown on user interface 2800 (for example, the first image 2805) can represent a modified document after the "Apply brand" button has been clicked on user interface 2800. That is, the style guide is applied to an input document to obtain the modified document.
[0187] In some examples, the UI 2800 includes style guide application settings with logos, colors, and fonts arranged in the left-hand panel of the UI 2800. A style transformation element (for example, the "Apply brand" button) is located in the top left of the UI 2800 to receive one-time user input.
[0188] In some examples, a user can implement style-specific elements (for example, brand elements) throughout the entire document with a single click via the 2800 user interface. In some cases, the 2800 user interface displays a preview image of the modified document. As in the example in Fig. As shown in Figure 16, the modified document contains a coffee cup with latte art foam. The coffee cup is surrounded by circular line decoration. Text content (for example, "BrewSoul") is placed next to it in a bold serif font. The background color of the modified document is pink. The text content is within a light blue area (for example, a light blue semicircle enclosing "BrewSoul").
[0189] Fig. Figure 29 shows an example of a state-change effect according to aspects of the present disclosure. The example shown includes the user interface 2900, the second image 2905, and a state-change element 2910. The user interface 2900 is an example of, or includes, aspects of the corresponding element, which relates to Fig. 3-11, 13, 18, 19 and 26-28. The state change element 2910 is an example of, or includes aspects of, the corresponding element, which relates to Fig. 28 is described.
[0190] As in the example in Fig. As shown in Figure 29, a user clicks the "Undo" button in the upper right corner of the user interface (Figure 2900). A document is displayed in the user interface (Figure 2900) showing the effect after the "Undo" action. The Undo and Redo buttons are located in the upper right corner of the user interface (Figure 2900).
[0191] The 2900 user interface displays a document after clicking the "Undo" button; that is, the document shows its state (for example, fonts, background colors, graphic elements) before receiving the one-click input to apply the style guide. Clicking the Undo button in the upper right corner of the 2900 user interface allows the 1320 document processing model (as described in...) to be used. Fig. 13) undo the previous style guide application and return to the document's previous style (for example, the input document before the style guide was applied).
[0192] In Fig. Page 29 of the document shows a logo with a stylized coffee cup featuring latte art, surrounded by a partial outline. The text "BrewSoul" is written in a rounded sans-serif typeface (in Fig. 28 (a bold serif font). The document has a beige background (in Fig. 28 pink background). The light blue semicircle around "BrewSoul" is not present due to the undo action.
[0193] In Fig. References 21-29 describe a method, an apparatus, a non-transitory computer-readable medium, and an image processing system. One or more aspects of the method, the apparatus, the non-transitory computer-readable medium, and the system include: obtaining a document and a style guide, wherein the document contains a text element and an image representing an object with a first color, and wherein the style guide contains a font and a second color; applying the font from the style guide to the text element to obtain a modified text element; applying the second color from the style guide to the image to obtain a modified image, wherein the modified image represents the object with the second color; and generating a modified document containing the modified image and the modified text element.
[0194] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include generating an initial text description of the first color in the image. Some examples further include generating an initial color embedding based on the initial text description. Some examples further include generating a second text description of the second color in the style guide. Some examples further include generating a second color embedding based on the second text description.
[0195] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include providing a style transformation element in a user interface. Some examples further include receiving a single click input via the style transformation element, with the modified image being generated based on the single click input.
[0196] Some examples of the method, apparatus, non-transitory computer-readable medium and system further include applying an additional typeface different from the typeface from the style guide to an additional text element of the document to obtain an additional modified text element, wherein the modified document includes the additional modified text element.
[0197] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include receiving a page selection input. Some examples further include applying the style guide to a large number of pages of the document based on the page selection input.
[0198] Some examples of the method, device, non-transitory computer-readable medium, and system further include applying a first style attribute from the style guide to a first element of the document. Some examples further include applying a second style attribute from the style guide to a second element of the document.
[0199] Fig. Figure 30 shows an example of a step-by-step procedure for training a machine learning model according to aspects of the present disclosure. Fig. Figure 30 shows a flowchart that represents an algorithm as a step-by-step procedure 3000 in an exemplary implementation of feasible operations for training a machine learning model. In some embodiments, the procedure 3000 describes a process of the training component 1345, which is associated with Fig. Procedure 3000 provides one or more examples of generating training data, using the training data to train a machine learning model, and using the trained machine learning model to perform a task.
[0200] To begin with this example, a machine learning system collects training data (Block 3002) to serve as the basis for training a machine learning model—that is, to define what will be modeled. The training data can be collected by the machine learning system from a variety of sources. Examples of training data sources include public datasets, system platforms of service providers that offer programming interfaces (e.g., social media platforms), systems for collecting user data (e.g., digital surveys and online crowdsourcing systems), and so on. Collecting training data can also involve data augmentation and synthetic data generation techniques to expand and diversify the available training data, balancing techniques to even out the number of positive and negative examples, and so forth.
[0201] The machine learning system can also be configured to identify features relevant (Block 3004) to a specific task for which the machine learning model is to be trained. Examples of such tasks include classification, natural language processing, generative artificial intelligence, recommendation systems, reinforcement learning, clustering, and so on. To this end, the machine learning system collects training data based on the identified features and / or filters the training data after collection based on these features. The training data is then used to train a machine learning model.
[0202] To train the machine learning model in the example shown, the model is first initialized (Block 3006). Initializing the machine learning model involves selecting a model architecture (Block 3008) to be trained. Examples of model architectures include neural networks, convolutional neural networks (CNNs), long short-term memory (LSTM) neural networks, generative adversarial networks (GANs), decision trees, support vector machines, linear regression, logistic regression, Bayesian networks, random forest learning, dimensionality reduction algorithms, boosting algorithms, deep learning networks, etc.
[0203] Furthermore, a loss function is selected (Block 3010). The loss function is used to measure the difference between an output of the machine learning model (i.e., predictions) and target values (e.g., as expressed by the training data), which is used to train the machine learning model. Additionally, an optimization algorithm is selected (Block 3012), which, in conjunction with the loss function, is used to optimize the parameters of the machine learning model during training; examples include gradient descent, stochastic gradient descent (SGD), etc.
[0204] The initialization of the machine learning model also includes setting initial values for the machine learning model (Block 3014); examples include initializing node weights and biases to increase training efficiency and reduce computational resource consumption. Hyperparameters are also set, which are used to control the training of the machine learning model; examples include regularization parameters, model parameters (e.g., the number of layers in a neural network), learning rate, batch sizes selected from the training data, and so on. The hyperparameters are set using various techniques, including randomization, heuristics learned from other training scenarios, and so forth.
[0205] The machine learning model is then trained by the machine learning system using the training data (Block 3018). A machine learning model is a computer-based representation that can be adapted (e.g., trained and retrained) based on inputs of the training data to approximate unknown functions. In particular, the term "machine learning model" can encompass a model that uses algorithms (e.g., using the model architectures described above) to learn from known data and to learn and relearn from the analysis of training data to generate outputs that reflect the patterns and attributes expressed by the training data.
[0206] Examples of training methods include supervised learning, which uses labeled data; unsupervised learning, which finds underlying structures or patterns in the training data; reinforcement learning based on optimization functions (e.g., rewards and / or punishments); the use of nodes in deep learning; and so on. The machine learning model, for example, can be configured to include a multitude of nodes that together form a multitude of layers. The layers might include, for example, an input layer, an output layer, and one or more hidden layers. Computations are performed by the nodes within the layers through the hidden states using a system of weighted connections that are "learned" during training.by using the selected loss function and backpropagation to optimize the performance of the machine learning model when performing an associated task.
[0207] During the training of the machine learning model, it is determined whether a termination criterion (decision block 3020) is met. This criterion is used to validate the machine learning model. The termination criterion can be used to reduce overfitting of the machine learning model, decrease the consumption of computational resources, and promote the machine learning model's ability to handle previously unseen data—that is, data not explicitly included as examples in the training data. Examples of a termination criterion include a predefined number of epochs, stabilization of validation loss, reaching a performance improvement threshold, achieving a certain level of accuracy, or the use of performance metrics such as precision and recall.If the termination criterion is not met (“no” from decision block 3020), procedure 3000 in this example continues to further train the machine learning model with the training data (block 3018).
[0208] If the termination criterion is met (“yes” from decision block 3020), the trained machine learning model is used to generate an output based on subsequent data (block 3022). The trained machine learning model is then, for example, trained to perform a task described above and is therefore, once trained, configured to perform this task based on subsequently received input data that is processed by the machine learning model.
[0209] Fig. Figure 31 shows an example of a computing device 3100 for document processing according to aspects of the present disclosure. The computing device 3100 can be an example of the one associated with Fig. The document processing device 1300 described in section 13 is to be. In one aspect, the computing device 3100 comprises processor(s) 3105, a storage subsystem 3110, a communication interface 3115, an I / O interface 3120, user interface component(s) 3125 and a channel 3130.
[0210] In some embodiments, the 3100 computing unit is an example of, or includes, aspects of the document processing model of Fig. 13. In some embodiments, the computing unit 3100 comprises one or more processors 3105 that can execute instructions stored in the memory subsystem 3110 to perform media generation.
[0211] According to some aspects, the Computing Unit 3100 comprises one or more Processors 3105. In some cases, a processor is an intelligent hardware component (e.g., a general-purpose processing unit, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, or a combination thereof). In some cases, a processor is configured to access a memory array using a memory controller. In other cases, a memory controller is integrated into a processor.In some cases, a processor is configured to execute computer-readable instructions stored in memory to perform various functions described herein. In some embodiments, a processor includes specialized components for modem processing, baseband processing, digital signal processing, or transmission processing.
[0212] According to some aspects, the 3110 memory subsystem comprises one or more memory devices. Examples of a memory device include random-access memory (RAM), read-only memory (ROM), or a hard disk. Examples of memory devices include semiconductor memory and hard disk drives. In some examples, memory is used to store computer-readable, computer-executable software, including instructions that, when executed, cause a processor to perform various functions described herein. In some cases, memory includes, among other things, a basic input / output system (BIOS) that controls basic hardware or software operations, such as interaction with peripheral devices or facilities. In some cases, a memory controller controls memory cells. For example, the memory controller may include a row decoder, a column decoder, or both.In some cases, memory cells in a memory store information in the form of a logical state.
[0213] According to some aspects, the 3115 communication interface operates at the boundary between communicating entities (such as the 3100 computing device, one or more user devices, a cloud, and one or more databases) and the 3130 channel, and can record and process communications. In some cases, the 3115 communication interface is provided to enable a processing system coupled with a transceiver (e.g., a transmitter and / or receiver). In some examples, the transceiver is designed to transmit (send) and receive signals for a communication device via an antenna.
[0214] In some aspects, the 3120 I / O interface is controlled by an I / O controller to manage input and output signals for the 3100 compute unit. In some cases, the 3120 I / O interface manages peripheral devices that are not integrated into the 3100 compute unit. In some cases, the 3120 I / O interface represents a physical connection or port to an external peripheral device. In some cases, the I / O controller uses an operating system such as iOS®, Android®, MS-DOS®, MS-WINDOWS®, OS / 2®, UNIX®, LINUX®, or another well-known operating system. In some cases, the I / O controller represents or interacts with a modem, keyboard, mouse, touchscreen, or similar device. In some cases, the I / O controller is implemented as a component of a processor.In some cases, a user interacts with a device via the I / O interface 3120 or via hardware components controlled by the I / O controller.
[0215] According to some aspects, the user interface component(s) 3125 enable a user to interact with the computing device 3100. In some cases, the user interface component(s) 3125 include an audio device, such as an external speaker system, an external display device, such as a screen, an input device (such as a remote control connected directly or via the I / O controller to a user interface), or a combination thereof. In some cases, the user interface component(s) 3125 include a graphical user interface (GUI).
[0216] Fig.Figure 32 shows an example of a diffusion transformer (DiT) architecture according to aspects of the present disclosure. The example shown includes: predicted noise 3205, predicted covariance 3210, linear and transformation layers 3215, normalization layer 3220, DiT block(s) 3225, patchify operation 3230, embedding 3235, noisy latent representation 3240, timestamp information 3245, label information 3250, and an implementation of a block of the DiT blocks 3225 by means of a DiT block 3296.The DiT block 3296 comprises: second residual connection 3260, second scaling operations 3262, feed-forward network 3264, second scaling and shift operation after normalization 3266, second normalization 3268, first residual connection 3270, first scaling operations 3272, self-attention 3274, first scaling and shift operation after normalization 3276, first normalization 3278, input token 3280, conditioning token 3282, multilayer perceptron (MLP) 3284, parameter of the first scaling and shift operation after normalization 3286, first scaling parameter 3288, parameter of the second scaling and shift operation after normalization 3290, and second scaling parameter 3292. In some embodiments, the architecture uses a latent diffusion transformer 3294. In some embodiments, the DiT block 3296 uses an “adaLN-Zero” technique.
[0217] Diffusion transformers (DiTs) are a popular architecture for diffusion models and were designed to remain structurally faithful to the standard transformer architecture. DiT leverages the scaling properties of the transformer architecture. For training denoising diffusion probabilistic models (DDPMs) of images (e.g., spatial representations of images), DiT is based on a vision transformer (ViT) architecture that operates on sequences of image patches. DiT processes images by dividing them into patches, converting these patches into tokens, and applying attention mechanisms to model relationships between different areas of the image. This approach allows the model to capture both local and long-range dependencies in the image generation process.
[0218] In some cases, the input for DiT is a spatial representation X. For 256 × 256 × 3 images, z has dimensions of 32 × 32 × 4. The first layer of a DiT performs the patchify operation, in which the DiT divides an input image into patches and converts the patches (a type of spatial input) into a sequence of T tokens by linearly embedding each patch in the input, with each token having dimension d. Following the patchify process, ViT frequency-based position embeddings are applied to all input tokens. In some cases, the number of tokens T generated by patchify is determined by a patch size hyperparameter p. In some cases, T = (I / p) 2, where I is another shape parameter; halving p thus quadruples T, which in some cases at least quadruples the total number of giga-floating-point operations (GFLOPS) of the transformer. In some examples, changing p has no effect on the number of parameters in downstream layers of the DiT, i.e., the number of parameters in subsequent layers of the DiT is independent of p. In some examples, p = 2, 4, or 8. Various patch sizes, transformer block architectures, and model sizes are implemented.
[0219] Following the patchify operation, attention mechanisms are applied to model relationships between different areas of the image in one or more DiT blocks. In addition to noisy image inputs, diffusion models sometimes process additional conditional information such as noise timestamps t, class labels c, natural language information, etc. Four variants of transformer blocks for processing conditional inputs, encompassing both input information and conditional information, are described below.
[0220] In some cases, DiT blocks in the DiT network are implemented using Adaptive Layer Norm (adaLN) blocks. Based on adaptive normalization layers in Generative Adversarial Networks (GANs) and conventional diffusion models with U-Net backbones, some examples replace standard normalization layers in transformer blocks with adaptive layer norm (adaLN). Instead of directly learning dimension-wise scaling parameters β and shift parameters γ, adaLN regresses β and γ from the sum of the embedding vectors of the noise timestamps t and the class labels c. An adaLN adds only a relatively small number of GFlops to the model and is more efficient. Additionally, adaLN is a conditioning mechanism that applies the same function to all tokens.
[0221] In some cases, DiT blocks in the DiT network are implemented using adaLN zero blocks, which utilize zero-initialization techniques. In residual networks (ResNets), it is advantageous to initialize each residual block as an identity function x ↦ x. In some examples, zero-initializing the final batch norm scaling factor γ to 0 in each block accelerates large-scale training in supervised learning scenarios. Diffusion models based on U-Nets use a similar initialization strategy by initializing the final convolution layer to 0 in each block before the residual connections. An adaLN zero block is derived from an adaLN block using similar zero-initialization techniques.In addition to regressing the dimension-wise scaling β and the shift parameters γ, the system also regresses dimension-wise scaling parameters applied immediately before residual connections within the DiT block. The network initializes a multilayer perceptron (MLP) to output a zero vector for all αs; this initializes the entire DiT block as an identity function. As with the adaLN block, adaLN-Zero adds a negligible number of G-flops to the model.
[0222] In some cases, DiT blocks in the DiT network are implemented with in-context conditioning, where vector embeddings of the noise timestamps and the class labels are appended to the input sequence as two additional tokens, and after a final block, the network removes the two conditioning tokens from the sequence.
[0223] In some cases, DiT blocks in the DiT network include cross-attention blocks. The DiT network concatenates the embeddings t and c into a sequence of length two, separate from the image-token sequence. The transformer block is modified to include an additional multi-head cross-attention layer after the multi-head self-attention layer.
[0224] In some cases, the DiT network comprises a sequence of N DiT blocks, each operating on a hidden dimension size d. Similar to ViT, the DiT network uses standard transformer configurations that scale N, d, and attention heads together. In some examples, Small (S), Base (B), Large (L), and XLarge (XL) variants of model sizes are implemented. Small or Base model sizes have N = 12 layers of DiT blocks, Large model sizes have 24 layers of DiT blocks, and XLarge model sizes have 28 layers of DiT blocks.
[0225] After the final DiT block, the DiT network decodes the sequence of image tokens into an output noise prediction and an output diagonal covariance prediction. Both outputs have dimensions corresponding to the original spatial input. A standard linear decoder is used for decoding, with a final normalization layer (or adaptive normalization layer if the DiT block is an adaLN block) linearly decoding each token into a p × p × 2C tensor, where c is the number of channels of spatial input to the DiT network and p is the patch-size hyperparameter. Finally, the decoded tokens are transformed back into their original spatial arrangement to obtain the predicted noise and covariance.
[0226] The DiT architecture uses a latent diffusion transformer 3294 in some cases. The DiT architecture processes the noisy latent representation 3240, which may be a noisy version of an input image encoded in a latent space. The patchify operation 3230 divides the noisy latent representation into a sequence of patches, which are processed as tokens. The tokens are vector representations of each patch of the image in the latent space and are adjusted by attention processes. Each of these tokens also receives timestamp information 3245 and label information 3250, and corresponding embedding 3235, which encodes the current denoising timestamp and the class labels as conditional information. In some cases, the embedding 3235 is referred to as conditional embedding or conditional information embedding.In some cases, a position embedding is added to the patchified input tokens during Patchify operation 3230, encoding the spatial position of each token in the image. In some examples, the position embedding is a ViT frequency-based position embedding. The input tokens 3280 generated by Patchify operation 3230 and the conditioning tokens 3282 generated by embedding 3235 are processed by N DiT blocks 3225, where N can be 12, 24, or 28. Other values of N can be used. In some cases, conditional tokens refer to tokens generated based on embedding 135, which encodes timestamp information 3245 and label information 3250.
[0227] Each of the DiT blocks 3225 comprises several processing stages. DiT block 3296 illustrates one embodiment of a block of the DiT blocks 3225. In some embodiments, DiT block 3296 is an example of, or comprises, aspects of, the adaLN zero block. In some cases, input tokens 3280 interact with conditioning tokens 3282 through several attention mechanisms. In particular, after applying the first normalization 3278 to the input tokens and the application of MLP 3284 to the conditioning tokens, MLP 3284 generates or updates the parameters of the first scaling and shifting operation after normalization 3286, designated β1, to scale and shift the output of the first normalization 3278 accordingly.Since the normalized input tokens from the first normalization 3278 are scaled and shifted during the first scaling and shifting operation after normalization 3276 using the conditional information contained at least in β1, the input information and the conditional information can interact with each other. Self-attention 3274 allows the scaled and shifted normalized input tokens—namely, the output of the first scaling and shifting operation after normalization 3276—to pay attention to each other. MLP 3284 also creates or updates the first scaling parameter 3288, designated δ1, for the first scaling operations 3272 to scale the output of self-attention 3274 (e.g., multi-head self-attention), thus further enhancing the interaction between the input information and the conditional information.The input tokens 3280 are then summed at the first residual connection 3270 with the output of the first scaling operations 3272. In some examples, δ1 has the initial value 0, and the DiT block 3296 is initialized as an identity function.
[0228] A similar process is performed in the second half of DiT block 3296. MLP 3284 generates or updates the parameters of the second scaling and shifting operation after normalization 3290, designated β2, for the second scaling and shifting operation after normalization 3266, in order to scale and shift the output of the second normalization 3268 accordingly. Since the output of the second normalization 3268 is scaled and shifted using the conditional information contained at least in β2, the input information and the conditional information can continue to interact. The feed-forward network 3264 then processes the scaled and shifted output of the second scaling and shifting operation after normalization 3266.MLP 3284 also generates or updates the second scaling parameter 3292, designated δ2, for the second scaling operations 3262 to scale the output of the feed-forward network 3264, further enhancing the interaction between the input and conditional information. In some cases, the feed-forward network 3264 is a pointwise feed-forward network. The output from the first residual connection 3270 is then summed at the second residual connection 3260 with the output of the second scaling operations 3262, and the result is the final output of the DiT block 3296. In some examples, δ2 has an initial value of 0, and the DiT block 3296 is initialized as an identity function. This process is repeated for each DiT block in the sequence.
[0229] After processing through all DiT blocks 3225, the outputs undergo a normalization layer 3220 and subsequently linearization and transformation layers 3215. The final output is the predicted noise 3205, which represents the model's prediction for the noise originally added to generate the noisy latent representation 3240, and the predicted covariance 3210, which represents the model's prediction for the covariance. The predicted noise 3205 is removed from the noisy latent representation 3240 at each diffusion time point, and the predicted covariance can influence how noise is removed or reinjected in the reverse, or denoising, process. At the end of the denoising schedule, the latent sample is decoded to generate the synthetic image in pixel space.
[0230] The performance of the present devices, systems, and methods has been evaluated, and the results show that embodiments of the present disclosure achieve improved performance compared to conventional techniques. Exemplary experiments demonstrate that the document processing device and machine learning model described in the embodiments of the present disclosure outperform conventional systems.
[0231] The present description and drawings represent example configurations and do not represent all implementations covered by the claims. For example, the described operations and steps can be rearranged, combined, or otherwise modified. Furthermore, structures and devices can be represented schematically in the form of block diagrams to clarify the relationships between components and to avoid obscuring the described concepts. Similar components or features may have the same name but different reference numerals assigned to different figures.
[0232] Some modifications of the disclosure are obvious to those skilled in the art, and the principles defined herein can be applied to other variants without departing from the scope of protection of the disclosure. The disclosure is therefore not limited to the examples and embodiments described herein, but should be interpreted to the broadest extent consistent with the principles and novel features of the present disclosure.
[0233] The described methods can be implemented or performed by devices comprising a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device, a logic circuit consisting of discrete gates or transistors, discrete hardware components, or any combination thereof. A general-purpose processor can be a microprocessor, a conventional processor, a controller, a microcontroller, or a state machine. A processor can also be implemented as a combination of several computing devices (for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration).The functions described herein can therefore be implemented in hardware or software and executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions can be stored as instructions or code on a computer-readable medium.
[0234] Computer-readable media include both non-transitory storage media and communication media, including any media that enables the transfer of code or data. A non-transitory storage medium can be any available medium that a computer can access. For example, a non-transitory computer-readable medium can include transitory memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), a compact disc (CD) or other optical storage medium, a magnetic storage medium, or any other non-transitory medium for carrying or storing data or code.
[0235] Furthermore, connecting components can be considered computer-readable media. For example, if code or data is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology such as infrared, radio, or microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology is included in the definition of a medium. Combinations of different media are also included within the scope of computer-readable media.
[0236] In this disclosure and the following claims, the word "or" indicates an inclusive list, so that, for example, the list "X, Y, or Z" means X or Y or Z or XY or XZ or YZ or XYZ. Furthermore, the expression "based on" is not used to represent a closed set of conditions. For example, a step described as "based on condition A" may be based on condition A and condition B. In other words, the expression "based on" is to be understood as "based at least partially on." Additionally, the words "a" or "an" in this disclosure and the following claims mean "at least one."
[0237] The following section describes further features, properties and advantages of the invention using a point-by-point approach: 1. A procedure, comprehensive: Receipt of a document and a style guide; Receiving user input that includes a request to apply the style guide to the document; and Generating a modified document based on the document and style guide in response to user input. 2. The procedure according to point 1, wherein receiving the user input comprises the following: Providing a style transformation element in a user interface; and Receiving a one-click input via the style transformation element, where the modified document is generated based on the one-click input. 3. The procedure according to point 1 or 2, furthermore including: Initiating a style transformation mode based on user input, generating the modified document based on the style transformation mode. 4. The procedure according to one of points 1 to 3, furthermore including: Obtaining a selection parameter that corresponds to a style attribute from the style guide, where the style attribute is applied to the document based on the selection parameter. 5. The procedure according to one of points 1 to 4, wherein: The style guide includes one or more fonts, one or more text colors, one or more background colors, one or more image colors, one or more images, or any combination thereof. 6. The procedure according to any one of points 1 to 5, furthermore including: Obtaining a color palette, with the style guide including the color palette. 7. The procedure according to point 6, wherein generating the modified document comprises the following: Identifying an image of the document; and Applying the color palette to the image to obtain a modified image, where the modified document includes the modified image. 8. The procedure according to one of points 1 to 7, wherein: The style guide is applied to the first page and the second page of the document. 9. The procedure according to point 8, wherein: The style of the first page is consistent with the style of the second page. 10. The procedure according to one of points 1 to 9, wherein: a first style attribute of the style guide is applied to a first element of the document, and a second style attribute of the style guide is applied to a second element of the document. 11. The procedure according to one of points 1 to 10, wherein: Both the original document and the modified document each contain a multimedia asset. 12. The procedure according to any one of points 1 to 11, wherein receiving the user input includes the following: Providing a style guide application tool for a user; and Receiving a style guide application input via the style guide application tool, where the style guide is based on the style guide application input. 13. The procedure according to one of points 1 to 12, wherein generating the modified document includes the following: Identify a first typeface and a second typeface that differs from the first typeface, from the style guide; Applying the first font to the first element of the document to obtain a first modified element; and Applying the second font to a second element of the document to obtain a second modified element, where the modified document includes the first modified element and the second modified element. 14. The procedure according to any one of points 1 to 13, furthermore including: Providing a state change element in a user interface; and Receiving a state change input via the state change element, where the modified document is generated based on the state change input. 15. The procedure according to any one of points 1 to 14, furthermore including: Receiving a setting in a user interface; and Determine, based on the setting, whether an element of the style guide should be applied to the document, and generate the modified document based on the determination. 16. A non-transitory, computer-readable medium that stores code for document processing, the code comprising instructions which, when read by at least one processor will be executed, causing at least one processor to perform operations that include the following: Receipt of a document and a style guide; Identifying an element of the document that matches a style from the style guide; and Generating a modified document based on the document and the style guide by applying the style from the style guide to the element of the document. 17. One comprehensive system: a storage unit; and a processing device coupled to the storage unit and configured to perform the following: Receipt of a document and a style guide; Receiving user input that includes a request to apply the style guide to the document; and Generating a modified document based on the document and style guide in response to user input. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 63 / 704,807
[0001]
Claims
[1] Procedure, encompassing: Receiving an image and a style guide, where the image shows an object with a first color and the style guide contains a second color; Identifying a second color from the style guide based on an approximation criterion between the first color and the second color; and Generate, using an image generation model, a modified image based on the image and the second color, where the modified image represents the object with the second color. [2] The method of claim 1, further comprising: Generating an initial text description of the first color in the image; and Generating a second text description of the second color in the style guide, where the approximation criterion is based on the first text description and the second text description. [3] Method according to claim 1 or 2, further comprising: Providing a style transformation element in a user interface; and Receiving a one-click input via the style transformation element, where the modified image is generated based on the one-click input. [4] Method according to any one of claims 1 to 3, further comprising: Identifying a color application parameter, where the second color is selected based on the color application parameter. [5] Method according to any one of claims 1 to 4, wherein: The style guide includes a font, a text color, a background color, a logo, or any combination thereof. [6] Method according to any one of claims 1 to 5, further comprising: Receiving a document containing the image; and Generating a modified document containing the modified image. [7] Method according to claim 6, further comprising: Receiving a page selection input; and Applying the style guide to a large number of pages of the document based on page selection input. [8] Method according to claim 6 or 7, wherein generating the modified document comprises: Applying the first style attribute from the style guide to the first element of the document; and Applying a second style attribute from the style guide to a second element of the document. [9] Method according to any one of claims 6 to 8, wherein generating the modified document comprises: Applying a font from the style guide to a text element of the document. [10] Method according to any one of claims 1 to 9, further comprising: Generating an initial color embedding that represents the first color based on the image; and Generating a second color embedding that represents the second color from the style guide, where the approximation criterion is based on a distance between the first color embedding and the second color embedding. [11] Non-transitory computer-readable medium storing code for document processing, the code comprising instructions which, when executed by at least one processor, cause the at least one processor to perform operations comprising: Receiving a document and a style guide, wherein the document contains a text element with a first font and an image showing an object with a first color, and wherein the style guide contains a second font and a second color; Apply the second font from the style guide to the text element to obtain a modified text element; Applying the second color from the style guide to the image using an image generation model to obtain a modified image, where the modified image shows the object with the second color; and Generating a modified document containing the modified image and the modified text element. [12] Non-transitory computer-readable medium according to claim 11, wherein the code further comprises instructions executable by the at least one processor to perform operations comprising: Generating an initial text description of the first color in the image; Generating an initial color embedding based on the first text description; Generating a second text description for the second color in the style guide; and Generating a second color embedding based on the second text description. [13] Non-transitory computer-readable medium according to claim 11 or 12, wherein the code further comprises instructions executable by the at least one processor to perform operations comprising: Providing a style transformation element in a user interface; and Receiving a one-click input via the style transformation element, where the modified image is generated based on the one-click input. [14] Non-transitory computer-readable medium according to any one of claims 11 to 13, wherein the code further comprises instructions executable by the at least one processor to perform operations comprising: Applying a third font, different from the second font, from the style guide to an additional text element of the document to obtain an additional modified text element, with the modified document containing the additional modified text element. [15] Non-transitory computer-readable medium according to any one of claims 11 to 14, wherein the code further comprises instructions executable by the at least one processor to perform operations comprising: Receiving a page selection input; and Applying the style guide to a large number of pages of the document based on page selection input. [16] Non-transitory computer-readable medium according to any one of claims 11 to 15, wherein generating the modified document comprises: Applying the first style attribute from the style guide to the first element of the document; and Applying a second style attribute from the style guide to a second element of the document. [17] System, encompassing: a storage unit; and a processing device coupled to the storage unit, wherein the processing device is configured to perform operations that include the following: Receiving an image and a style guide, where the image shows an object with a first color and the style guide includes a second color; Identifying a second color from the style guide based on an approximation criterion between the first color and the second color; and Generate, using an image generation model, a modified image based on the image and the second color, where the modified image shows the object with the second color. [18] System according to claim 17, further comprising: a language generation model configured to generate a first color embedding based on the first color and a second color embedding based on the second color. [19] System according to claim 17 or 18, wherein: The image generation model is configured to produce the modified image as a synthetic image by applying the second color to the object. [20] System according to any one of claims 17 to 19, wherein the processing device is further configured to perform operations comprising: Providing a style transformation element in a user interface; and Receiving a one-click input via the style transformation element, where the modified image is generated based on the one-click input.
Citation Information
Patent Citations
US63704807B2
US-ANMELDUNGNR.63/704,807