Content generation using existing media assets with generative machine learning models
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2026-04-08
AI Technical Summary
Existing systems fail to efficiently break down and recombine media assets into dynamic components suitable for various display formats and aspect ratios, often requiring redundant processing and resource waste due to inability to distinguish between overlaid and integrated text, and struggle with coordinating generative models for high-quality asset generation.
A system utilizing generative machine learning models to extract and generate media asset components, including background images, text, and logos, while distinguishing between overlaid and integrated text, and incorporating feedback loops for improved image quality, allowing for dynamic asset recombination and resource optimization.
Enables high-quality, adaptable media asset generation that conserves resources by distinguishing text types and ensures image quality, supporting various display formats and aspect ratios, and reduces redundant processing.
Smart Images

Figure 2026510503000001_ABST
Abstract
Description
Technical Field
[0001] Priority This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 501,191, filed May 10, 2023, the disclosure of which is incorporated herein by reference.
[0002] The present disclosure generally relates to automatically generating content items or media assets using a generative machine-learned model based on a user's profile or preferences and existing media assets.
Background Art
[0003] Communication campaigns can utilize multi-modal, multi-platform delivery systems to deliver content items to various endpoints for various audiences. Content items can include data or other information or messages. Content items can be media assets or can include media assets. A user can create a communication campaign by providing a set of content items for delivery to a multi-modal, multi-platform delivery system.
Summary of the Invention
Means for Solving the Problems
[0004] Aspects and advantages of embodiments of the present disclosure are described in part in the following description, or can be learned from the description, or can be learned through practice of the embodiments.
[0005] One exemplary aspect of this disclosure relates to a computing system for generating content items. The computing system may include one or more processors and one or more non-temporary computer-readable media. The computer-readable media may store together machine-learned generative models, machine-learned selection models, and instructions. The machine-learned generative models may be configured to generate multiple content items. The machine-learned selection models may be configured to select selected content items from multiple content items. When the instructions are executed by one or more processors, they cause the computing system to perform an action. The action may include receiving data indicating a request for multiple media assets with multiple media modalities generated based on existing media assets.The operation involves generating multiple media asset components, (i) extracting one or more signals from an existing media asset comprising at least one of background image data, color palette data, text data, or logo image data, (ii) generating multiple background image assets using a first generative machine learning model, which generates a mask image containing one or more bounding boxes based on at least one of background image data, text data, or logo image data, and repairing the mask image by generating pixels to fill one or more bounding boxes of the mask image using the machine learning model, and generating multiple images, each image having a distinct aspect ratio. The operation may include generating multiple background image assets, including a ratio; (iii) generating multiple text assets using a second generative machine learning model, which involves obtaining text data associated with one or more bounding boxes, inputting the text data into the machine learning model, and having the machine learning model generate multiple text assets as output; and (iv) generating multiple media asset components using a third generative machine learning model, which involves generating unique profile data based on at least one of color palette data or logo image data. The operation may also include automatically sending multiple media assets, including one or more image assets, text assets, and unique profile data, to a content item generation pipeline to generate multiple candidate content items.
[0006] In some examples, generating a mask image containing one or more bounding boxes includes identifying one or more text objects and determining, for each of the one or more text objects, that a first text object among the one or more text objects is an overlay text object. In some examples, generating a mask image containing one or more bounding boxes includes generating a first bounding box for the first text object based on the determination that the first object is an overlay text object.
[0007] In some examples, generating multiple media asset components includes extracting a logo image. In some examples, generating multiple media asset components includes selecting a dominant color from an image color palette, enlarging the logo image, sharpening the logo image, and upscaling the logo image. In some examples, generating multiple media asset components includes generatively expanding the logo image to generate one or more images, each image including a set of different aspect ratios.
[0008] In some examples, generatively expanding a logo image involves generating pixels to blend with the existing pixels of the logo image to produce an image larger than the original logo image.
[0009] In some examples, generating multiple background image assets involves adjusting at least one of the following for each background image asset: brightness, saturation, or contrast.
[0010] In some examples, the data for a unique profile includes at least one of the following: logo data, color palette data, font data, or image styling data.
[0011] In some examples, the operation involves inputting multiple output assets into the content creation pipeline. In some examples, the operation involves the content creation pipeline generating multiple content items, each of which content items includes a unique combination of content assets and aspect ratios.
[0012] In some examples, generating multiple text assets involves extracting signals containing text data. In some examples, generating multiple text assets involves compiling text data. In some examples, generating multiple text assets involves inputting the compiled text data into a second generative machine learning model. In some examples, generating multiple text assets involves obtaining at least one short or long heading as output from the second generative machine learning model. In some examples, generating multiple text assets involves inputting at least one short or long heading into a second generative machine learning model. In some examples, generating multiple text assets involves obtaining a description as output from the machine learning model.
[0013] In some examples, one or more bounding boxes include the position and size of the bounding boxes.
[0014] In some examples, the content item generation pipeline includes prompt generation components and content item generation components.
[0015] In some examples, the operation includes a prompt generation component generating input prompt data based on a background image, text assets, and unique profile data. In some examples, the operation includes providing the input prompt data to a content item generation component. In some examples, the operation includes a content item generation component generating one or more candidate content items based on the input prompt data.
[0016] In some examples, content item generation components include generative machine learning models.
[0017] In some examples, the operation involves training a generative machine learning model, which includes generating a training dataset based on comparing generated content item components with existing media assets, and automatically tuning one or more parameters of the generative machine learning model to reduce the differences between the generated content items and existing media assets based on comparing the generated content item components with existing media assets.
[0018] In one exemplary embodiment, the Disclosure provides an exemplary computer implementation. The exemplary computer implementation includes receiving data indicating a request for multiple media assets having multiple media modalities generated based on existing media assets. The exemplary method generates multiple media asset components by (i) extracting one or more signals comprising at least one of background image data, color palette data, text data, or logo image data; (ii) generating multiple background image assets using a first generative machine learning model based on at least one of background image data, text data, logo image data, or color palette data; (iii) generating multiple text assets using a second generative machine learning model based on at least one of text data or background image data; and (iv) generating unique profile data using a third generative machine learning model based on at least one of color palette data, text data, or logo image data. The exemplary computer implementation includes automatically sending the multiple media assets, comprising one or more image assets, text assets, and unique profile data, to a content item generation pipeline to generate multiple candidate content items.
[0019] In some examples, the method involves receiving data that indicates existing media assets being uploaded.
[0020] In some examples, the method involves parsing web resources associated with unique profile data to retrieve existing media assets.
[0021] In some examples, existing media assets include flattened image data.
[0022] In some examples, requests are associated with a client account, and the client account is associated with an account profile that stores inputs to a machine learning-based media asset generation pipeline.
[0023] In some cases, profiles were retrieved from a database and pre-generated prior to the request.
[0024] In some examples, the Disclosure provides exemplary temporary or non-temporary computer-readable media embodied in a computer-readable storage device and storing instructions that, when executed by a processor, cause the processor to perform an action. In exemplary temporary or non-temporary computer-readable media, an action includes receiving data indicating a request for multiple media assets with multiple media modalities generated based on existing media assets. The action includes generating multiple media asset components by (i) extracting one or more signals comprising at least one of background image data, color palette data, text data, or logo image data; (ii) generating multiple background image assets using a first generative machine learning model based on at least one of background image data, text data, logo image data, or color palette data; (iii) generating multiple text assets using a first generative machine learning model based on at least one of text data or background image data; and (iv) generating unique profile data using a third generative machine learning model based on at least one of color palette data, text data, or logo image data. The operation includes automatically submitting multiple media assets, including one or more image assets, text assets, and unique profile data, to the content item generation pipeline in order to generate multiple candidate content items.
[0025] Other aspects of the present disclosure are directed to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.
[0026] These and other features, aspects, and advantages of the various embodiments of the present disclosure will become better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the relevant principles.
[0027] A detailed description of embodiments directed to those skilled in the art is set forth in this specification with reference to the accompanying drawings.
Brief Description of the Drawings
[0028] [Figure 1] A diagram showing a block diagram of an exemplary system according to an exemplary embodiment of the present disclosure. [Figure 2] A diagram showing a block diagram of an exemplary system according to an exemplary embodiment of the present disclosure. [Figure 3] A diagram showing a block diagram of an exemplary system according to an exemplary embodiment of the present disclosure. [Figure 4] A diagram showing a flowchart of an exemplary method for generating a media asset according to an exemplary embodiment of the present disclosure. [Figure 5] A diagram showing a flowchart of an exemplary data flow for generating a media asset according to an exemplary embodiment of the present disclosure. [Figure 6] A diagram showing a block diagram of an exemplary method for generating a media asset according to an exemplary embodiment of the present disclosure. [Figure 7] A diagram showing an exemplary graphic representation of steps for generating a media asset from an existing media asset according to an embodiment of the present disclosure. [Figure 8]This figure shows an exemplary graphic representation of the steps for generating a media asset from an existing media asset according to an embodiment of the present disclosure. [Figure 9] This figure shows an exemplary graphic representation of the steps for generating a media asset from an existing media asset according to an embodiment of the present disclosure. [Figure 10] This figure shows an exemplary graphic representation of the steps for generating a media asset from an existing media asset according to an embodiment of the present disclosure. [Figure 11] This figure shows an exemplary graphic representation of the steps for generating a media asset from an existing media asset according to an embodiment of the present disclosure. [Figure 12] This figure shows an exemplary graphic representation of an asset feedback layer according to an exemplary embodiment of the present disclosure. [Figure 13] This figure shows an exemplary graphic representation of an asset feedback layer according to an exemplary embodiment of the present disclosure. [Figure 14] This figure shows an exemplary graphic representation of an asset feedback layer according to an exemplary embodiment of the present disclosure. [Figure 15] This figure shows an exemplary graphic representation of the steps for generating a media asset from an existing media asset according to an embodiment of the present disclosure. [Figure 16] This flowchart illustrates an exemplary method for training a machine learning model according to embodiments described in the present disclosure. [Figure 17] This is a block diagram of an exemplary processing flow for using a machine learning model(s) to process input(s) to generate output(s) according to an exemplary embodiment of the present disclosure. [Figure 18] This is a block diagram of an exemplary sequencing model according to an exemplary embodiment of the present disclosure. [Figure 19]This is an exemplary block diagram of the technique for arranging an exemplary input sequence for processing by a sequence processing model according to exemplary embodiments of the aspects of the present disclosure. [Figure 20] This is a block diagram of an exemplary model development platform according to an exemplary embodiment of the aspects of the present disclosure. [Figure 21] This is a block diagram of an exemplary training workflow for training a machine learning model according to exemplary embodiments of the aspects of the present disclosure. [Figure 22] This is a block diagram of an estimation system for running one or more machine learning models for performing estimations, according to an exemplary embodiment of the aspects of the present disclosure. [Figure 23] This is a block diagram of an exemplary networked computing system according to an exemplary embodiment of an aspect of the present disclosure. [Figure 24] This is a block diagram of an exemplary computing device according to an exemplary embodiment of the aspects of the present disclosure. [Figure 25] This is a block diagram of an exemplary computing device according to an exemplary embodiment of the aspects of the present disclosure. [Modes for carrying out the invention]
[0029] The repeated reference numbers across multiple drawings are intended to identify the same features in various embodiments.
[0030] This disclosure generally relates to providing the generation of media asset components based on existing media content items received. For example, the systems and methods described herein can decompose existing media content items into components. The components can be generated using a generative machine learning model that can be trained and tuned based on a profile or requirements. The systems and methods may additionally or alternatively include rearranging the media asset components so that the generated content items are visually similar (e.g., the vector difference between two images is within a threshold). This disclosure may include the use of one or more generative machine learning models, such as a large language model that can perform optical character recognition, generate bounding boxes for restoration, and generate components and assets.
[0031] Existing methods do not provide a way to acquire initial media assets, such as flat images, and break them down into various components. Existing media channels require dynamic asset components for maximum performance across various display formats. For example, the dimensions or aspect ratios required for rendering via a mobile device interface may differ from those rendered on a desktop computer. Additionally, or alternatively, different applications within the same device may require various aspect ratios, abbreviated descriptions, or other distinguishing features of the content. Therefore, it is crucial to be able to acquire individual media asset components that can be combined on demand to meet these dynamic requirements. In some examples, existing media content items may be flattened images that cannot be resized without distorting the image or otherwise degrading image quality. Therefore, a solution is needed that can acquire these existing media content items and break them down into individual media assets that can be recombined.
[0032] In some examples, the initial media content item may include a flattened image where pixels between the overlaid text are aligned with the underlying image. Therefore, the overlaid text or other media elements cannot be simply removed. Furthermore, some underlying background images may contain text that needs to be distinguished from the overlaid text (e.g., the SPF marking on a sunscreen bottle compared to a graphic that reads "For Sale"). Therefore, this disclosure provides a technical advantage over existing systems by distinguishing between overlaid text and text integrated within a background image, determining whether to initiate the use of other generative models in the workflow. For example, additional models may include models capable of filling background images based on generated bounding boxes, generating text assets based on detected overlaid text, or utilizing data extracted in other ways. By determining whether text is overlaid text or text integrated within a background image, the system can conserve resources by preventing the use of additional computational resources to fill areas of the background image that should not be removed. This prevents redundant processing and wasted resources.
[0033] While methods exist for performing the individual actions described herein, complexity arises when using the output of one model as input to another model to obtain high-quality media asset components. For example, coordinating the generation of a mask to determine which part of an image to fill may involve iterations of training a model to distinguish between text to be restored and text that is not (e.g., words on a bottle in an image compared to text overlaid as part of an existing media asset). Furthermore, ensuring that the generated media asset is similar to the existing asset to comply with requirements may require coordinating and training a generative model to produce a media asset that satisfies image quality metrics such as clarity, or other metrics. This disclosure enables improved asset generation by incorporating existing assets, extracting text, performing painting, augmenting a base image, generating new text, and overlaying the text and augmented base image to generate a new asset.
[0034] This system and method provides improvements to the use, training, and tuning of machine learning models to perform each step of the process of generating new assets. By breaking down assets into sub-components, assets are no longer limited to the initial two-dimensional set aspect ratio. Rather, the sub-components can be used to generate assets of various sizes and extend to additional asset types such as text-only assets, image-only assets, and video assets.
[0035] This disclosure may include post-processing checks that provide feedback data to a model that generates output assets to determine image quality (e.g., based on distortion or blur) and automatically adjust system parameters to provide improved image quality (smaller distortion or blur metric than the previously generated image).
[0036] In some embodiments, this disclosure relates to automatically generating content items or media assets based on a client's profile or preferences. An exemplary embodiment provides generating multiple media assets based on a media asset profile using a machine learning-based media asset generation pipeline by instructing a machine learning-based asset generation model to generate media assets that align with a client's media asset preferences. Furthermore, exemplary embodiments of this disclosure generally relate to generating content based on information extracted from a website using a machine learning-based model. For example, a user can enter a web address into the system, and the system can generate content for the user based on data extracted using the web address. Exemplary techniques include automatically generating multiple content items for a communication campaign and selecting content items from the multiple content items based on predictions of how well selected content items will perform in the communication campaign.
[0037] For example, a client can be a user associated with a user account. Users can interact with the campaign generation system to create new communication campaigns. Users can interact with the campaign generation system using their user account. The campaign generation system can associate a user account with a set of campaign preferences, which may include a set of media asset preferences. If a user account is associated with other communication campaigns, the media asset preferences may include preferences obtained based on those other campaigns (e.g., preferences received directly through input from the user account, preferences implicitly learned from the user account's actions, etc.).
[0038] A new communication campaign can communicate data that directs the audience to data resources related to the campaign. This data can be or contain resource locators (e.g., URIs, URLs, deep links, app links). Data resources can be web resources (e.g., web pages, web applications) or local resources (e.g., native applications running on a client device).
[0039] Users can provide the campaign generation system with resource locators for data resources. This can be provided early in the campaign generation process, such as when initiating the generation of a new campaign. The campaign generation system can process the resource locators to identify data resources and retrieve data from and about them. For example, in the case of a campaign pointing to a webpage, the campaign generation system can use the provided URL to load the target webpage, parse or crawl the page to extract existing media assets (e.g., images, text, video, color palette, typography, etc.) and learn about themes, styles, and entity branding associated with the target webpage. Other relevant resources can be parsed. A sitemap can be used to parse resources on the document tree where the data resource resides. The system can also parse resources linked to the data resource, such as based on a relevance metric for linked resources.
[0040] The campaign generation system can predict resource locators to prefetch content in order to improve latency. For example, for a user account with known associations to known resource locators, the campaign generation system can start parsing data resources using known resource locators even before the user has confirmed the resource locators.
[0041] The campaign generation system can prefetch content for improved latency by having the resource locator begin parsing data resources as soon as an input field is populated, even if the user has not yet completed other input fields on the same interface screen. In this way, for example, by the time the user moves on to the next input screen, the parsing process is either well underway or already completed, thereby reducing user latency.
[0042] Media asset preferences can be updated using data analyzed from data resources. The campaign generation system can use media asset preferences to form or update account profiles describing the account's communicative personality (e.g., brand personality), account assets, performance data from any past campaigns, and learned characteristics of relevant audiences and learned trends or characteristics for the entire group of communicators. This account profile can be dynamically maintained as campaigns are delivered and updated, as campaign communications are received and used by recipient endpoints, etc. The account profile can also be dynamically updated as users interact with the machine learning-based media asset generation pipeline to save current progress, current preferences, selections, inputs, signals, etc.
[0043] The campaign generation system may collect additional input signals from users. These input signals can refine predicted or pre-configured features of account profiles or media asset preferences. For example, based on data analyzed from data resources, the campaign generation system can predict the initial goals, general themes and styles of a communication campaign, as well as other data resources related to the communication campaign. Users can refine, update, approve, or reject these predictions by providing additional input signals.
[0044] Additional input signals may include product / service names, product / service descriptions (e.g., free-form, multi-line input where the user specifies details about those products / services; suggestions / pre-placed), brand characteristics (e.g., adjectives describing the brand; may be suggested in relation to a point-and-click interface), and social media opt-in (e.g., permission to retrieve assets from social media platforms associated with the user account). A machine learning model can provide dictation for one or more of the input fields for the signals based on the account profile or analyzed data resources. Thresholds may be used to dictate when a confidence level is exceeded (e.g., corresponding to the quality of the dictation).
[0045] Additional input signals can persist in relation to a user account. Additional input signals can persist in relation to assets generated based on those additional input signals. Additional input signals may include metadata indicating whether a particular signal was manually modified by the user. This persisted signal data can be resurfaced for the user if they create another relevant campaign in the future. For example, if a user creates a campaign on the same or similar data resource, the signal data can be resurfaced without having to re-analyze the data resource first. This can improve latency and reduce processing requirements. The signal data can be used as input to a machine learning-based asset generation pipeline when analyzing data resources. A subset of the signal data, such as only manually confirmed / modified signals, can be used as input to a machine learning-based asset generation pipeline. In this way, for example, the machine learning-based asset generation pipeline can learn from user input / modifications and avoid making the same errors with respect to future campaigns.
[0046] The campaign generation system can process data analyzed from data resources, account profile data, and additional input signals to obtain media assets for use in communication campaigns. The campaign generation system can implement a machine learning-based media asset generation pipeline to retrieve or modify existing media assets, generate new media assets, or retrieve new media assets from a database, guided by account profile data and additional input signals. For example, the machine learning-based media asset generation pipeline can generate images (e.g., background images), headlines, descriptions, videos, logos, color palettes, site links, and visual styles and themes. The machine learning-based media asset generation pipeline can retrieve or modify existing images, headlines, descriptions, videos, logos, color palettes, site links, and visual styles and themes. The machine learning-based media asset generation pipeline can query relevant databases (e.g., stock media asset databases) to obtain new images, headlines, descriptions, videos, logos, color palettes, site links, and visual styles and themes.
[0047] The machine learning-based media asset generation pipeline can extract or modify existing media assets or existing media content items. It can analyze data resources and extract any content from them. Content from data resources can be modified or optimized. For example, images or videos can be resized, text overlays on images or videos can be removed and filled (e.g., using a machine learning-based restoration model), and images or videos can be edited (e.g., exposure, color, sharpness). Text media assets can be paraphrased or edited for clarity. Logos can be identified, rescaled, optimized for overlays (e.g., background removal, alpha channel generation), recolored, etc. Other existing assets can be retrieved from the media library associated with the account. The media library may include assets used in past campaigns, uploaded or generated but not yet used, etc.
[0048] A machine learning-based media asset generation pipeline can generate media assets using one or more machine learning-based models. It can use a machine learning-based natural language understanding model to parse text on data resources to understand the content of those resources and learn the context in which the content is presented (e.g., the style or subject of the data resource). For example, a machine learning-based natural language understanding model can use additional data, such as image data or other contextual data, to determine whether text on a data resource is overlay text or text embedded within an image. The model can then determine whether further steps should be taken in relation to the location of the text content. For example, the model can determine whether to generate a bounding box to create a mask for an image that should be filled in or repaired. Additionally or alternatively, the model can determine whether detected text should be ignored or utilized in a text generation process (e.g., generating text assets such as headlines and descriptions). A machine learning-based media asset generation pipeline can obtain a set of asset generation instructions that may be based on, or may include, one or more of the following: the representation of content and its context, account profiles, media asset preferences, or additional new input signals.
[0049] The campaign generation system can refer to an allowlist to determine whether a user account is authorized to use the machine learning-based media asset generation pipeline. For example, campaigns related to products in sensitive verticals may bypass automated asset generation and require manual control by the user. For instance, a user may be queryed to provide manual input and control in order to generate assets using the machine learning-based media asset generation pipeline. The user may be completely locked out of the machine learning-based media asset generation pipeline.
[0050] A machine learning-based media asset generation pipeline can use a machine learning-based image generation model to process asset generation instructions and generate images that are both based on and consistent with those instructions. Various image generation architectures can be used, including convolutional neural networks, transformers, generative adversarial networks, and diffusion models. The image generation model can process images from a data resource as exemplary input, prompting the model to generate similar images, textual descriptions of the desired image, and other signals or instructions, or trained soft prompts. For example, product images from a data resource can be provided to the image generation model(s), prompting the model(s) to include the product in the generated image, or to outpaint around the product in a new environment (e.g., generatively fill). This is an example of a technique for contextualizing or recontextualizing product images while improving the faithful reproduction of product attributes. Other exemplary techniques for image asset generation include processing assets from data resources to extract attributes (subject, color, mood), using a machine learning-based language model to generate asset generation commands and prompts based on the extracted attributes, and inputting the prompts or asset generation commands and prompts into the image generation model.
[0051] A machine learning-based media asset generation pipeline can use a text generation model to process asset generation instructions and generate text that is both based on and aligned with those instructions. Various text generation architectures can be used, including convolutional neural networks, converters, generative adversarial networks, and diffusion models. Exemplary architectures include encoder-only, encoder-decoder, or decoder-only converter-based models trained on large text corpora. The text generation model can process images from a data resource as exemplary input to prompt relevant descriptions, desired output text, and other signals or instructions, as well as trained soft prompts.
[0052] In some examples, a text generation model may process resource locators, text from data resources, free-form text provided by users or generated by the asset generation pipeline (e.g., using a prompt generator), existing text assets associated with user accounts, tone and brand indicators (e.g., brand-related adjectives or other descriptors, which may be obtained from additional signal inputs), etc. A text generation model(s) may be configured to categorize text assets (e.g., as call to action, promotional phrases, descriptions, etc.). The quality of the generated assets can be evaluated (e.g., by the generation model itself, by a quality control model, etc.). This can be used for subsequent ranking / selection of text assets. For example, a quality metric may include evaluating relevance or foundation with respect to the data resource (e.g., evaluating whether "contactless delivery" is a phrase that accurately describes the content of the data resource). In the case of existing text assets, the campaign generation system may process the existing text assets along with any of the above inputs to rewrite the assets (e.g., change the tone, etc.).
[0053] A machine learning-based media asset generation pipeline can use video generation models to process asset generation instructions and generate videos that are both based on and aligned with those instructions. Various video generation architectures can be used, including convolutional neural networks, converters, generative adversarial networks, diffusion models, and continuous or discrete-time cascaded diffusion models.
[0054] A machine learning-based media asset generation pipeline can use audio generation models to process asset generation instructions and generate audio that is both based on and aligned with those instructions. Various audio generation architectures can be used, including convolutional neural networks (e.g., for processing spectrograms), converters (e.g., for processing sequences or embeddings of audio data), generative adversarial networks, spread models, and continuous or discrete-time cascaded spread models.
[0055] A machine learning-based media asset generation pipeline can use a machine learning-based prompt generator model to generate prompts for input to other generative models in the pipeline. The machine learning-based prompt generator model can be trained end-to-end with one or more of the other generative models to improve performance. The prompt generator model may include language generation models (e.g., a "large language model"). The machine learning-based media asset generation pipeline can facilitate generative models with different prompts to obtain a variety of different outputs. For example, the output layer of a prompt generator model can be sampled (e.g., randomly, sampling the top K outputs, etc.) to obtain a classification of the prompt outputs. This classification can then be input to the corresponding generative model to generate a variety of outputs related to the instruction. The prompt generator can receive prompts provided by the user and rewrite the prompts based on representation symbols or images (e.g., "progress" → "person climbing a mountain"). For example, a prompt can be rewritten by inputting the original prompt and instruction into a language generation model (e.g., prompting the model to "suggest an image associated with 'progress'"). The prompt generator model can receive user-provided prompts (e.g., obtained through a feedback loop as described below) or system-rewritten prompts, and extend the prompts to be more performance-oriented (e.g., "mountain climber" → "mountain climber. photo, detail, HDR, high resolution, 4K").
[0056] Generated assets can be associated with metadata. For example, image assets created or enhanced using a machine learning-based media asset generation pipeline can store metadata containing information about which tools / pipelines (and which versions) were used to create or enhance the asset. This can flow into assets derived from assets created / enhanced by the machine learning-based media asset generation pipeline. "Enhanced" may include optimized / optimized features. In this way, the campaign generation system can perform an analysis on how well the enhancements are implemented (and possibly test against unenhanced versions). Furthermore, the campaign generation system can facilitate the recall ("takedown") of generated / enhanced assets (or derived assets) as needed. This can be limited to net-newly generated content or generated content covering the main portion (20%+) of the image. For generated images where prompts are used, prompts can be saved. Prompts entered by any user, as well as any prompts generated by a prompt generator, can be saved.
[0057] A machine learning-based media asset generation pipeline can query relevant databases for assets. For example, a stock photo or video database can be queried for content similar to the asset, either extracted from or generated based on the data resource.
[0058] A machine learning-based media asset generation pipeline can retrieve assets (e.g., generate, modify, and query databases) based on learned attribute insights. For example, a learned attribute insight model can map subjects (e.g., products, topics, etc.) to additional content, keywords, or features (e.g., attributes) that are more relevant for higher performance. For instance, an image asset of "dog toy" could be mapped to outdoor and sunlight depictions based on learned relationships, leading to higher performance. Such insights can be used to generate / modify assets of any type. These insights can be used to broaden or narrow search queries for relevant assets from the asset database. Such insights can also be surfaced to the user during the generation workflow for additional information. Furthermore, such insights can be provided in prompts (e.g., passed directly to the generative model, passed to a prompt generator, etc.) to improve asset generation.
[0059] A machine learning-based media asset generation pipeline can optimize acquired media assets. Optimization may include cropping, repair, outpainting, upscaling, recoloring, sharpening, or other editing or modifications. Optimization can be performed by one or more machine learning-based optimization models (e.g., image editing models, video editing models, audio editing models, etc.). Optimization can be logged in metadata. Optimization steps can be rolled back by reloading the state of the saved asset from the metadata.
[0060] A machine learning-based media asset generation pipeline can rank acquired media assets. For example, a machine learning-based ranking model can rank acquired media assets based on the potential performance of the media asset in a communication campaign (e.g., the predicted likelihood that a user will interact with the corresponding content item and execute a hyperlink embedded in the content item). A machine learning-based ranking model can rank acquired media items based on their relevance to data resources. The machine learning-based asset generation pipeline can generate embedded representations of data resources and compare them to the embedded representations of acquired media items to determine relevance. Ranking can be based on the source of the image (e.g., system-generated, crawled from a data resource, user-uploaded, etc.). Ranking can be based on image recognition results (e.g., an image recognized as belonging to a product described on the data resource). Ranking can be based on alignment with additional signals input by the user.
[0061] Ranking can also be performed based on best practices. Machine learning-based ranking models can be trained to identify best practices for media assets. Heuristic-based best practices can also be checked. A best practice score may be provided. The score can be based on estimated performance lift (e.g., for a specific audience). For example, it may be determined that positioning a product at the center of a media asset tends to result in a measurable increase in website visits.
[0062] Based on ranking, a machine learning-based media asset generation pipeline can select and present to the user assets generated from multiple assets in the content database (e.g., new and / or modified assets). For example, the top-ranked assets can be selected for presentation. A set of the top K assets can be selected. Asset sampling can be selected from different ranking positions (e.g., to be more robust to ranking errors).
[0063] Acquired assets may be presented differently based on their ranking. For example, up to a threshold ranking, some assets may be recorded or pre-selected, allowing the user to simply confirm the pre-selection and proceed. Below the threshold ranking, assets may be offered as suggestions for manual selection by the user. Acquired assets may also be presented differently based on their asset type. For example, text assets may be recorded as described above. In some situations, image assets may not be recorded.
[0064] The campaign generation system can request user feedback on acquired media assets. The campaign generation system can provide a user interface that presents acquired media assets, along with interactive input elements provided for editing the acquired media assets. The campaign generation system can provide a user interface that presents input fields for providing natural language commands for changes to be made to acquired media assets (e.g., "make the flowers brighter"), or further commands for generating new media assets based on presented candidates (e.g., "generate more assets like this asset"). For example, to generate more assets like a presented asset, the campaign generation system can input the existing asset into the corresponding generative model as part of a prompt to generate similar assets. When generating more assets like a previously generated asset, the prompt used to generate the previously generated asset can be reused. One modification option involves inputting an existing asset (generated or otherwise) into the model along with a prompt, and then outputting multiple options for the modified asset based on the prompt.
[0065] User feedback can be fed back into a machine learning-based asset generation pipeline, allowing media assets to be regenerated or modified according to the feedback signals. This process can be repeated until the user approves the media asset.
[0066] User feedback can be obtained using an input-response interface. For example, a natural language input / output interface, such as voice or text, can be provided to receive user input in natural language and implement the requested changes. The system can also generate output in natural language to describe the updates that have been implemented.
[0067] The campaign generation system can output media assets to the content item generation system. The content item generation system can use the media assets to generate content items. For example, the content item generation system can combine text assets (e.g., headlines, taglines, descriptions) with image assets (e.g., product images, background images) to create content items for distribution. The content item generation system can generate content items based on their potential use. For example, the use of a content item may include interacting with the content item to execute a hyperlink embedded in it. For example, a hyperlink may use a resource locator to direct an endpoint device to a data resource.
[0068] User feedback and choices can provide training data to improve one or more components of a machine learning-based asset generation pipeline. For example, losses, rewards, or penalties may be based on user feedback and choices. The campaign generation system can train one or more components of the machine learning-based asset generation pipeline to reduce losses, increase rewards, or reduce penalties. Training techniques may include supervised training (e.g., supervised by user input), unsupervised training (e.g., learning patterns of account behavior and optimizing output based on those patterns), and reinforcement learning (e.g., the asset generation pipeline as an agent pursuing rewards).
[0069] Other model-fitting techniques, such as soft prompts, may be used. For example, one or more soft prompts can be trained for input to any one or more of the generative models. The soft prompts may be associated with a specific user account or campaign. In this way, for example, the asset generation pipeline can be customized to improve performance for individual user accounts, and optionally, the entire pipeline will not be retrained to do so.
[0070] Early ranking in a machine learning-based asset generation pipeline can be used to prioritize the use of available processing bandwidth. For example, before generating assets, the instructions for generating assets can be ranked (e.g., by processing asset generation instructions and any other inputs with a machine learning-based ranking model), and the media asset generation pipeline can generate a set of instructions ranked to the top or top K. This pre-generation ranker can be trained based on the final output of the machine learning-based media asset generation pipeline. In this way, for example, fewer, lower-ranked media assets are generated first by ranking them before instruction generation. In this way, the processing used to generate media assets can also be allocated to higher-priority (e.g., higher-ranked) generation tasks.
[0071] The generated assets can be processed by policy checks. For example, a policy check system can evaluate the generated output for any sensitive material (e.g., material that violates platform policies). Generated assets that violate policies can be excluded and not presented to the user.
[0072] A policy check system can be applied to inputs to a campaign generation system (e.g., user-provided inputs, data parsed from data resources, etc.). The policy check system can screen for personally identifiable information (PII), obscene language, sensitive subjects, or other policy-based screening rules. The policy check system can screen any user-provided or data-parsed input and exclude it from further processing in any other model component.
[0073] In some embodiments, the system can perform product studio functionality. Product studio functionality can be a customer-facing product that enables product packaging photography. Product studio functionality can upscale images, remove backgrounds, and place products within scenes (e.g., transforming packaging images into more lifestyle images). For example, product studio functionality can place clothing products on various models. Furthermore, product studio functionality can implement automated product variants. Automated product variants can use aggregated data to determine the type of scene that is appealing to the customer's target audience, or alternatively, allow the customer to specify the scene; automatically identify product photos from a source (e.g., a URL); remove backgrounds (if any); generate multiple variations of the product image (placing the product in a scene or background, or alternatively, on a model in the case of clothing); and / or test all scenes on traffic to see which performs best. For example, in a lotion campaign, the system could place the product among trees, such as with palm trees, if the system determines that this placement will lead to more user clicks or conversions. Product studio functionality can also upscale images.
[0074] The product studio function can generate scenes by accepting image and text inputs using ML models. Furthermore, the system can generate several text inputs using formulas for optimal results. <background>Before <surroundings>Surrounded by<product type> <platform>For example, "a skincare jar on a light stone platform surrounded by almonds or in front of green plants." The system can automatically generate this text. Furthermore, the original product image can be layered back into the ML output to ensure product integrity. In some examples, the ML model can be modified more than the background. The system can apply edge smoothing to ensure the product blends into the scene. The system can move the product lower than the center so that the ML model can position the product more appropriately on the surface.
[0075] The embodiments of this disclosure provide several technical effects, advantages, and / or improvements in computing and artificial intelligence technologies that involve using machine learning algorithms to generate new data such as images, audio, text, video, or other types of media. The technologies described herein improve the use of generative models by improving the quality of the generated content. The quality of the generated content is tailored specifically to entities (e.g., companies, users) by using data extracted from the entities' web resources. For example, by using more content-relevant data, the system improves the performance of generative models. Furthermore, the system leverages better training techniques by developing more efficient and effective training techniques specific to entities (e.g., based on data extracted from the entities' web resources) to reduce the time and resources required to train the model. Furthermore, the system can incorporate user feedback and provide feedback to generative models that can help the model learn from user preferences and improve over time, via reinforcement learning or active learning. Furthermore, the disclosure can reduce processing by reducing the amount of manual input provided by the user and by reducing the number of interface screens that need to be retrieved, loaded, interacted with, and updated. For example, a user only needs to enter a website's web address, and the system can automatically extract content from the website and automatically generate content items for the user.
[0076] Exemplary embodiments of this disclosure are discussed in further detail here with reference to the drawings.
[0077] Figure 1 shows an exemplary system for implementing a machine learning-based media asset generation pipeline 100. The machine learning-based media asset generation pipeline 100 may include a machine learning-based text generator 101. The machine learning-based media asset generation pipeline 100 may include a machine learning-based image generator 102. The machine learning-based media asset generation pipeline 100 may include a machine learning-based audio generator 103. The machine learning-based media asset generation pipeline 100 may include a machine learning-based video generator 104. The machine learning-based media asset generation pipeline 100 may include one or more optimizers 105 that apply one or more optimization algorithms to one or more outputs of any one of the machine learning-based generator models 101-104. The machine learning-based media asset generation pipeline 100 may include one or more rankers 106 to rank one or more outputs of any one of the machine learning-based generator models 101-104.
[0078] The machine learning-based media asset generation pipeline 100 can ingest data from data resources 110 and account profiles 120. Account profiles 120 can include media asset preferences. Account profiles 120 can include media libraries 122. Account profiles 120 can include social media accounts 124. Account profiles 120 can include historical signals / controls 126 input to the machine learning-based media asset generation pipeline 100. The machine learning-based media asset generation pipeline 100 can process the data extracted from data resources 110 and account profiles 120 according to new signals / controls 130. New signals / controls 130 can include user inputs to customize media asset generation.
[0079] The machine learning-based media asset generation pipeline 100 may include an asset feedback layer 140. The asset feedback layer 140 facilitates the input of user feedback on the generated assets and can initiate the generation of updated or different assets. After selection, confirmation, or approval using the asset feedback layer 140 (for example, shown in Figures 12, 13, and 14), the machine learning-based media asset generation pipeline 100 may output a media asset 150. The media asset 150 may include any type of media asset output. The media asset output may include, for example, text assets, image assets, audio assets, video assets, and / or unique profile data (e.g., brand profile data, color palette, logo).
[0080] Figure 2 shows a flowchart of an exemplary machine learning-based media asset generation pipeline 200 according to an exemplary embodiment of the present disclosure. In some examples, in 202, the system may receive a website and / or asset library. In 204, the system may determine an understanding of products and brands based on the information received and / or acquired in 202. In 206, the system may identify existing assets based on the information received and / or acquired in 202. In 208, the system may customize products and / or brands based on the determination in 204. In 210, the system may modify (e.g., update) existing assets identified in 206. In 212, the system may determine logos and colors based on the information derived in 208 and / or 210. In 214, the system may determine insights about companies and / or products based on the information derived in 208 and / or 210. In 214, the system may also perform gap analysis for prediction or auto-generated missing information based on the information derived in 208 and / or 210.
[0081] Furthermore, in 216, the system can generate a new asset based on the information derived in 214. In 218, the system can modify the new asset generated in 216 by adding (e.g., modifying) text, images, videos, and / or site links. The text, images, videos, and / or site links selected in 218 can be determined or generated based on the information derived in 212 and 214. In 220, the system can receive user input to customize the new asset generated in 216 and modified in 218. In 222, the system can provide (e.g., present) the customized asset 220 using an AI-powered format.
[0082] The machine learning-based media asset generation pipeline 200 may include an overall model. The overall model may be a machine learning-based generative model configured to generate multiple content items. Additionally or alternatively, the overall model may be a machine learning-based selection model configured to select a content item from among multiple content items. In some embodiments, the overall model is trained to receive a set of input data 204 describing a web resource and, as a result of receiving the input data 204, automatically generate output data 206 of new media assets and content items. For example, the system may receive user input related to a web resource from a user's user device. The system may extract multiple assets (e.g., images, words, videos, or audio files) from the web resource. Furthermore, the system may use the overall model (e.g., a machine learning-based generative model) to process the multiple assets and generate multiple content items. Furthermore, the system may use the overall model (e.g., a machine learning-based selection model) to determine a selected content item from among multiple content items. The system may then trigger the presentation of the selected content item on a graphical user interface displayed on the user's device.
[0083] In another embodiment, the system may receive data indicating a request for multiple media assets, including multiple media modalities. Furthermore, the system may obtain a media asset profile for the client account associated with the request. The media asset profile may include data indicating media asset preferences for the client account, and the media asset profile may be generated by processing existing media assets associated with the client account. The system may generate multiple media assets using a machine learning-based media asset generation pipeline 200 by instructing the overall model (e.g., a machine learning-based asset generation model) to generate media assets that are consistent with the media asset preferences, based on the media asset profile. The system may then send one or more of the multiple media assets to a content item generation system for generating a content item using one or more of the multiple media assets, based on the data received indicating one or more selections of the multiple media assets.
[0084] According to several embodiments, the system can work with clients to create high-quality media assets, automatically drawing in curated and all kinds of media assets for the client's business. Any business, large or small, can start advertising using the system in seconds, even if they don't yet have any assets. This system can lower barriers and popularize creative development for advertising for everyone, enabling all businesses to reach customers in a personalized and engaging way.
[0085] The system combines best machine learning models, including generative AI, with deep insights to help automatically embed entire asset groups for most new campaigns in real time. With a single click, clients can instantly start building asset group sets to deliver results for their specific goals, and then modify content items and / or media assets based on suggestions received from the system.
[0086] For example, a client can input just the right amount or less of information to generate content items, and when the client generates these content items, in some embodiments, the client may be able to review the system's assumptions, have the opportunity to make adjustments, and accept media assets (e.g., content items) that the client desires. The client can either directly publish the recommended media assets or use them only as a starting point to customize or build their own.
[0087] The system may include a user interface framework for collecting inputs for the creation, collection, and combination of intelligent assets. The system can then surface these assets and system assumptions and return them to the client (e.g., customer). The system can enable the refinement of media assets based on user input, all within the media asset construction process or onboarding flow process.
[0088] Figure 3 shows a block diagram 300 of an exemplary system according to an exemplary embodiment of the present disclosure. The system can receive a URL 302 from a user. For example, the system can receive user input related to the URL from the user's user device. The system can extract multiple assets 304 from a data resource 110 associated with the URL 302. The multiple assets 304 may include brand understanding, product and service large language models (LLMs), images, sitemaps, log understanding, social accounts, business LLMs, asset libraries, performance data, and historical campaign data. Furthermore, the system, a machine learning-based media asset generation pipeline 100, can process the multiple assets 304 to generate multiple content items 308. The overall model 306 can perform ranking and insight determination, text and / or image-generating artificial intelligence, automated asset generation, stock lockup, product generation, and video creation. The multiple content items 308 may include images, headlines, descriptions, videos, logos, colors, site links, personalities, and visual styles. The system can use a machine learning-based content item generation pipeline 310 to determine selected media assets from multiple media assets and generate content items 312. The system can then trigger the presentation of the new content items on a graphical user interface displayed on the user's device.
[0089] Figure 4 shows a flowchart of an exemplary method 400 for generating media assets by decomposing existing content items, according to several embodiments of the present disclosure. Method 400 can be carried out by processing logic that may include hardware (e.g., processing devices, circuits, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions implemented or executed on the processing device), or a combination thereof. In some embodiments, Method 400 is carried out by a server computing system (e.g., server computing system 60) or a client computing system (e.g., computing device 50). Although shown in a specific sequence or order, the order of processes can be changed unless otherwise specified. Thus, the illustrated embodiments should be understood as examples only, and the processes shown can be carried out in a different order, and some processes can be carried out in parallel. Furthermore, in various embodiments, one or more processes can be omitted. Thus, not all processes are required in all embodiments. Other process flows are possible.
[0090] In operation 402, the processing logic may receive data indicating a request for multiple media assets, including multiple media modalities generated based on existing media assets. For example, in some embodiments, the processing logic may receive data indicating existing media assets being uploaded. For example, a user may provide existing media assets to the system via an online portal or other interface. In some embodiments, existing media assets may be retrieved by parsing a web resource associated with unique profile data to retrieve the existing media assets. For example, a user may have an account associated with unique profile data. A URL may be provided to the web resource that can be parsed for the existing media assets.
[0091] Existing assets can include flattened image data. For example, an existing asset may be a jpg or other image file containing a single layer of pixels. This is distinguishable from a multi-layer image file, which can contain various layers that may have overlapping or different pixel values based on the layer. While layered image files can be more easily divided into components, in a flattened image, multiple layers (such as a background image and overlay text) are combined or merged into a single layer. Consequently, pixels in the top layer overwrite pixels in lower layers, rendering portions of the original background image unrecoverable.
[0092] In some embodiments, requests are associated with a client account, which in turn is associated with an account profile that stores inputs to a machine learning-based media asset generation pipeline. The profile may be retrieved from a database, or it may be pre-generated prior to the request. For example, a user identifier may be associated with a profile or unique profile data.
[0093] In operation 404, the processing logic can generate multiple media components by executing operations 406, 408, 410, and 412. For example, referring to Figure 6, the processing logic can retrieve an existing media asset 605. The operations performed can generate multiple media assets. Media assets may include text assets, image assets 615, and unique profile assets such as logo assets 620, color assets 625, font assets 630, and image styling 635.
[0094] Media assets can be acquired by the machine learning content item generation pipeline 640 to generate assets for various media channels, as will be further described in relation to operation 414.
[0095] Generating multiple media asset components may include extracting a logo image. Generating multiple media asset components may include selecting a dominant color from an image color palette. Generating multiple media asset components may include upscaling a logo image by enlarging it and / or sharpening it.
[0096] Generating multiple media asset components may include generatively expanding a logo image to generate one or more images. Each image may have a different set of aspect ratios. Generatively expanding a logo image may include generating pixels to blend with the existing pixels of the logo image to generate an image larger than the original logo image.
[0097] For example, referring to Figure 8, a graphic display 800 is shown that extracts and generates unique profile data (e.g., brand profile) such as dominant colors and graphics. For example, an existing asset 805 can be used as input to a machine learning-based asset generation pipeline. The pipeline can extract or otherwise determine dominant colors 810. Additionally or alternatively, the pipeline can extract a logo 815 from an image. The extracted logo 815 may be low resolution or otherwise unsuitable for use when generating future assets. Thus, the pipeline can upscale the logo 820 to create a larger version while maintaining clear visual resolution. Additionally or alternatively, the pipeline can uncrop the logo 825 or otherwise generatively outpaint the logo to create a larger visual buffer around the logo.
[0098] In operation 406, the processing logic can extract one or more signals from an existing media asset, each containing at least one of the following: background image data, color palette data, text data, or logo image data.
[0099] In operation 408, the processing logic can generate multiple background image assets using a first generative machine learning model. In some examples, generating multiple background images may include generating a mask image containing one or more bounding boxes based on at least one of the background image data, text data, or logo image data. The one or more bounding boxes may include the position and size of the bounding boxes.
[0100] In some examples, generating multiple background images may involve using a machine learning model to repair a mask image by generating pixels to fill one or more bounding boxes of the mask image. In some examples, generating multiple background images may involve generating multiple images, each having a distinct aspect ratio.
[0101] Generating multiple background image assets may involve adjusting at least one of the following for each background image asset: brightness, saturation, or contrast.
[0102] Generating a mask image containing one or more bounding boxes based on at least one of background image data, text data, or logo image data may include identifying one or more text objects. Generating a mask image for each of the one or more text objects may include determining that a first text object among the one or more text objects is an overlay text object. Generating a mask image for each of the one or more text objects may include generating a first bounding box for the first text object based on the determination that the first object is an overlay text object.
[0103] Figure 7 shows an exemplary graphic display 700 that generates several background image media assets based on an existing media asset 705 (e.g., an existing content item). For example, the existing media asset 705 can be input into a model that can generate several bounding boxes (e.g., bounding box 715 and bounding box 720) to create a mask 710. The system can determine which bounding boxes should be filled by determining which are associated with overlay text, such as bounding box 715, and which are associated with underlying image text, such as bounding box 720.
[0104] For example, the model can be trained to distinguish between bounding boxes that should be repaired and those that should be ignored and not included in the mask. The output image 725 can be the image generated by repairing the bounding box 715.
[0105] In some embodiments, the background image generation model can enhance the image by adjusting the image's saturation, sharpness, brightness, or other features to generate an enhanced image 730.
[0106] In some embodiments, the background image generation model can provide background images with various aspect ratios by outpainting the image or otherwise generatively expanding it. For example, background image 735 may be more suitable for full-screen mobile application advertisements, background image 740 may be more suitable for search or email advertisements, and background image 745 may be more suitable for advertisements rendered on desktops.
[0107] In operation 410, the processing logic can generate multiple text assets using a second generative machine learning model. Generating multiple text assets may include obtaining text data associated with one or more bounding boxes. Generating multiple text assets may include inputting the text data into the machine learning model. Generating multiple text assets may include the machine learning model generating multiple text assets as output.
[0108] Generating multiple text assets can include extracting signals containing text data. Generating multiple text assets can include compiling text data. Generating multiple text assets can include inputting the compiled text data into a second generative machine learning model. Generating multiple text assets can include obtaining at least one short or long heading as output from the second generative machine learning model. Generating multiple text assets can include inputting at least one short or long heading into a second generative machine learning model. Generating multiple text assets can include obtaining a description as output from the machine learning model.
[0109] Referring to Figure 9, an exemplary graphic representation 900 can be seen that extracts and generates text data, including a headline and description, from an existing media content item 905. The processing logic can extract text from the existing media content item 905 to generate a text data signal 910. The text data signal 910 can be cleaned up to generate cleaned-up text 915. The cleaned-up text 915 can be used as input to a second generative machine learning model to generate additional headlines 920 and / or descriptions 925. The headlines 920 and / or descriptions 925 can be used by a content item generation pipeline to generate one or more media content items for various media channels.
[0110] In operation 412, the processing logic may use a third generative machine learning model to generate unique profile data based on at least one of color palette data or logo image data. The unique profile data may include at least one of logo data, color palette data, font data, or image styling data. The unique profile data may include brand data associated with the profile or other identifiers. For example, the unique profile data may be associated with the "look and feel" of typical content items associated with a profile (e.g., business).
[0111] In operation 414, the processing logic can automatically send multiple media assets, including one or more image assets, text assets, and unique profile data, to the content item generation pipeline in order to generate multiple candidate content items. The content item generation pipeline includes a prompt generation component and a content item generation component.
[0112] In some examples, the processing logic can generate input prompt data based on background images, text assets, and unique profile data using a prompt generation component. The processing logic can provide the input prompt data to a content item generation component. The processing logic can then generate one or more candidate content items based on the input prompt data using the content item generation component. The content item generation component includes a generative machine learning model.
[0113] In some embodiments, the processing logic can train a generative machine learning model. For example, training a generative machine learning model may include generating a training dataset based on comparing generated content item components with existing media assets. Training a generative machine learning model may also include automatically tuning one or more parameters of the generative machine learning model to reduce the differences between generated content items and existing media assets based on comparing generated content item components with existing media assets.
[0114] In some embodiments, the processing logic can input multiple output assets into the content creation pipeline. The processing logic can generate multiple content items through the content creation pipeline, where each of the multiple content items includes a unique combination of content assets and aspect ratios.
[0115] Returning to Figure 6, the machine learning-based content item generation pipeline 640 can generate content items that are displayed through various media channels that require different formats. For example, media channels may include a video channel 645, an email channel 650, a display channel 655, a search channel 660, a discovery channel 665, and / or a map channel 670. Different channels may require different formats; for example, some may contain images, text, and audio, while others may contain only text. An exemplary discovery channel asset 675 is shown, which demonstrates a combination of image, text, logo, and embedded (e.g., selectable component linking to a URL) assets.
[0116] Figure 5 shows a flowchart of an exemplary data flow 500 for generating a media asset according to an exemplary embodiment of the present disclosure. For example, a media asset generation pipeline(s) 510 may take an existing media asset 505 (e.g., a previously existing content item) as input. The media asset generation pipeline(s) 510 may include an image processing component(s) 515, a mask generation component(s) 520, and / or a generative machine learning model(s) 525. The image processing component(s) 515 may process the existing media asset 505 to extract background image data, text data, color palette data, or other relevant data. The mask generation component(s) 520 may determine which parts of the background image need to be generatively filled (e.g., to remove overlaid text). The generative machine learning model(s) 525 may generatively fill based on the generated mask and may also generate text or other assets based on any of the extracted signals associated with the existing media asset 505.
[0117] A media asset generation pipeline(s) 510 can output media assets 530. The media assets 530 may include text assets 535, image assets 540, and / or unique profile data 545. As described herein, text assets 535 may include words, sentences, paragraphs, or other phrases. Image assets 540 may include background images, logo images, or other images. Unique profile data 545 may include brand information or other information related to a user profile.
[0118] The media asset 530 can be taken in by the content item generation pipeline 550 to generate candidate content items 560. The content item generation pipeline 550 may include content item generation models 555 that can take the media asset 530 and / or other signals as input to generate candidate content items 560. The generated candidate content items 560 may be content items generated in a new way that they look visually different from the original existing media asset.
[0119] The systems described herein can perform several methods. For example, exemplary methods can be performed by one or more computing systems (e.g., one or more computing systems discussed with respect to Figures 1-25). Various steps of the methods can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of this disclosure.
[0120] The computing system can receive user input associated with web resources from the user's user device, and the web resources are associated with the user's account.
[0121] The computing system can extract multiple assets from web resources, each of which can be an image, word, video, or audio file.
[0122] The computing system can use machine learning-based generative models to process multiple assets and generate multiple content items.
[0123] The computing system can use a machine learning-based selection model to determine which content item is selected from among multiple content items.
[0124] A computing system can trigger the presentation of selected content items on a graphical user interface displayed on a user device.
[0125] In some examples, the operation may further include receiving user interaction on a graphical user interface, which modifies a selected content item. Furthermore, the operation may include using a machine learning generative model to process the user interaction and the selected content item to generate a modified content item. Additionally, the operation may include triggering the presentation of the modified content item on a graphical user interface displayed on the user's device. Furthermore, one or more parameters of the machine learning generative model may be updated based on the user interaction.
[0126] In some examples, the behavior may further include receiving user interaction on a graphical user interface. This user interaction may be associated with rejecting a selected content item. Furthermore, the behavior may include using a machine learning-based selection model to process multiple content items and user interactions to generate new content items. Additionally, the behavior may include triggering the presentation of new content items on a graphical user interface displayed on the user's device.
[0127] In some examples, the operation may further include receiving user interaction on a graphical user interface, where the user interaction receives selected content items. Furthermore, the operation may include using a machine learning model to determine advertising campaigns based on the selected content items. Additionally, the operation may include triggering the presentation of advertising campaigns on a graphical user interface displayed on the user's device.
[0128] In some examples, a web resource can be a website, and user input is the website's uniform resource locator (URL).
[0129] In some examples, multiple content items may contain a first content item, which can be generated by modifying an image asset from among multiple assets. Furthermore, multiple content items may contain a second content item, which is a generative image generated by a machine learning-prepared generative model using the image asset.
[0130] In some examples, the operation may further involve using a machine learning-based selection model to calculate a conversion score for each content item among multiple content items, where the conversion score indicates the likelihood of a user interacting with each content item. For example, a selected content item might be the one with the highest conversion score among multiple content items.
[0131] In some embodiments, the computing system may receive data indicating requests for multiple media assets, including multiple media modalities.
[0132] The computing system can retrieve the media asset profile of the client account associated with the request. The media asset profile contains data indicating the client account's media asset preferences, and the media asset profile was generated by processing existing media assets associated with the client account.
[0133] The computing system can generate multiple media assets by using a machine learning-based media asset generation pipeline to instruct a machine learning-based asset generation model to generate media assets that are consistent with media asset preferences, based on the media asset profile.
[0134] Based on receiving data indicating one or more selections of multiple media assets, the computing system may send one or more of the multiple media assets to a content item generation system in order to generate a content item using one or more of the multiple media assets.
[0135] In some examples, multiple media modalities include two or more modalities selected from text, images, or audio.
[0136] In some examples, the operation could further include generating media asset profile data by parsing web resources associated with the client account.
[0137] In some examples, the operation may further involve parsing web resources and extracting existing media assets from them.
[0138] In some examples, the operation may further include parsing web resources to extract visual style data associated with the client account. For example, visual style could include color information, layout information, or typography information.
[0139] In some examples, the operation may further include parsing web resources and extracting text style data associated with the client account. Text style data may include intonation or inflection of the copy on the web resource.
[0140] In some examples, the operation may further include parsing web resources to extract landing page data associated with a client account. Landing page data may include URLs to web pages associated with multiple media assets.
[0141] In some cases, media asset profiles were retrieved from a database and pre-generated prior to the request.
[0142] In some examples, the operation may further involve generating at least one of multiple media assets by editing an existing image asset using at least one of the following editing operations: cropping, rotating, filling, recoloring, defocusing, deblurring, noise reduction, and re-brightening. The editing operations are performed optionally using machine learning-based image editing tools. Furthermore, the existing image asset may be edited based on historical performance data associated with the image asset. Furthermore, the existing image asset may be edited based on a set of content item guidelines for generating content items using the existing image asset.
[0143] In some examples, the operation may further include inputting data from media asset profiles and requests for generated assets that match the data from media asset profiles into a machine learning-based media asset generation model.
[0144] In some examples, the operation may further include using a machine learning-based performance estimation model to determine one or more generated assets, the machine learning-based performance estimation model configured to identify asset characteristics associated with historical performance data. Furthermore, the operation may further include using the machine learning-based performance estimation model to generate augmented inputs for input to a machine learning-based media asset generation model in order to induce asset characteristics associated with historical performance data. Furthermore, the operation may include using the machine learning-based performance estimation model to rank the assets generated from the machine learning-based media asset generation model.
[0145] In some examples, the operation may further include presenting one or more generated media assets for review on a user interface accessible by the client account. Furthermore, the operation may include receiving inputs via the user interface that provide corrections to one or more generated media assets. Furthermore, the operation may include regenerating one or more generated media assets based on the received inputs using a machine learning-based media asset generation pipeline. Furthermore, the user interface may include one or more selectable input elements that relate to one or more generated media assets and indicate corresponding correction actions to be performed with respect to one or more generated media assets. Selectable input elements may be configured to provide received inputs when selected. The user interface may include a natural language input element for receiving correction inputs in natural language format, and the natural language input element may be configured to provide received inputs.
[0146] In some examples, a media asset profile can be based on one or more features of a machine learning model, images, sitemaps, logos, social media accounts, asset libraries, performance data, a historical set of media assets, or a historical set of generated media assets, with one or more features associated with the client account.
[0147] In some examples, multiple media assets can include two or more categories from the following categories: images, headlines, descriptions, videos, logos, colors, sitelinks, calls to action, and audio.
[0148] Figure 10 shows an exemplary graphical representation of extracting existing media assets from source 1005 associated with an account, and using the extracted assets and media assets generated from other sources as input to a machine learning-based media asset generation pipeline 1030. For example, source 1005 may include display network channels, uploads, social media, or imports. Access to source 1005 can be used to obtain image ads 1010. Image ads 1010 may be existing content items such as flattened image ads.
[0149] The machine learning-based media asset generation pipeline 1030 can generate multiple media assets using various inputs. These various inputs may include image advertisements 1010, URLs 1015, business profiles 1020, and / or mobile inputs 1025. Using these inputs, background image assets, text assets, and / or unique profile data can be generated, as described herein.
[0150] Figure 11 shows an example of extracting media assets from an existing content item. For example, the content item may include content item 1105A, content item 1105B, and content item 1105C. The system can generate multiple candidate background images, such as background image 1110A, background image 1110B, and background image 1110C. The system can determine a primary color and / or color palette, such as color 1115A, color 1115B, or color 1115C. The system can generate text assets, such as text asset 1120A, text asset 1120B, and text asset 1120C. Background images 1110A, 1110B, and / or 1110C may be generated using methods described herein, including mask generation, fill / repair, and generative expansion (e.g., uncropping).
[0151] Figure 12 shows an exemplary user interface for interacting with the asset feedback layer 140. The asset feedback layer 140 can display acquired media assets along with a source indicator (e.g., "From URL"). In some embodiments, the asset feedback layer 140 can display a loading indicator (e.g., a solid bar instead of text that has not yet been loaded) while receiving assets generated from the machine learning-prepared media asset generation pipeline 100. The asset feedback layer 140 can pre-position the asset fields in the generated assets.
[0152] The asset feedback layer 140 can display a loading indicator (e.g., a solid area instead of an image that has not yet been generated) while receiving assets generated from the machine learning-trained media asset generation pipeline 100. In some examples, it may display different loading status messages (e.g., "#Generating images with AI", "#Searching for the best matching stock image").
[0153] Figure 13 shows an exemplary user interface for interacting with the asset feedback layer 140. The asset feedback layer 140 can provide menu options related to the acquired asset for performing actions in relation to the asset.
[0154] Figure 14 shows an exemplary user interface for interacting with the asset feedback layer 140. The asset feedback layer 140 can provide menu options for removing assets. The asset feedback layer 140 can provide an interface for providing feedback related to asset removal. This feedback can be used to train one or more components of the machine learning-trained media asset generation pipeline 100.
[0155] Additionally, or alternatively, the asset feedback layer 140 may provide an interface for generating additional assets. The asset feedback layer 140 may provide a proposed asset generation prompt, which may be configured to be selectable for initiating processing of the proposed prompt. The asset feedback layer 140 may provide exemplary generated assets. The asset feedback layer 140 may provide an interface for viewing the output associated with a given prompt in a browser. The asset feedback layer 140 may provide an interface for inputting natural language prompts.
[0156] Additionally, or alternatively, the asset feedback layer 140 may provide an interface for viewing other extracted and suggested media assets, such as colors, text assets, and site links.
[0157] Figure 15 shows an exemplary data flow for generating various content items for various distribution mechanisms (e.g., media channels). Content items can be configured to trigger the loading of data resources (e.g., data resource 110 shown in Figure 1) in response to interactions. For example, a machine learning-based content item generation pipeline 310 can generate content items for multiple media channels, such as video, email, display, search, discover, or maps. Two exemplary formats are shown in Figure 15.
[0158] Figure 16 shows a flowchart of a method 1600 for training one or more machine learning models according to an aspect of this disclosure. For example, exemplary machine learning models may include a machine learning media asset generation pipeline, a machine learning content item generation pipeline, a machine learning text generator, a machine learning image generator, a machine learning audio generator, and a machine learning video generator.
[0159] One or more parts of Exemplary Method 1600 may be implemented by a computing system including one or more computing devices, such as the computing system described with reference to other figures. Each of the parts of Exemplary Method 1600 may be implemented by any (or any combination) of one or more computing devices. Furthermore, one or more parts of Exemplary Method 1600 may be implemented, for example, on the hardware components of the devices described herein to train one or more systems or models. Figure 16 shows the elements to be performed in a particular order for illustrative and explanatory purposes. Those skilled in the art will understand that, using the disclosures provided herein, any element of any of the methods discussed herein can be adapted, rearranged, extended, omitted, combined, or modified in various ways without departing from the scope of this disclosure. Figure 16 is described for the purposes described exemplary and is not intended to limit to the elements / terms described in relation to other systems and figures. One or more parts of Exemplary Method 1600 may be implemented by other systems, additionally or alternatively.
[0160] In 1602, exemplary method 1600 may include obtaining training instances. The set of training data may include multiple training instances, which are split across multiple datasets (e.g., a training dataset, a validation dataset, or a test dataset). Training instances may be labeled or unlabeled. Although referred to as “training” instances in exemplary method 1600, it should be understood that runtime estimates can form training instances when a model is trained using an evaluation of the model’s performance against its runtime instances (e.g., online training / learning). Exemplary data types for training instances and various tasks associated therewith are described throughout this disclosure.
[0161] In 1604, exemplary method 1600 may include processing a training instance using one or more machine-trained models to produce an output. The output may be obtained directly from one or more machine-trained models or as a downstream result of a chain of processing operations that includes the output of one or more machine-trained models.
[0162] In 1606, exemplary method 1600 may include receiving an evaluation signal associated with the output. The evaluation signal may be obtained using a loss function. Various loss determinations can be used, such as mean squared error, probability loss, cross-entropy loss, hinge loss, contrast loss, or various other loss functions. The evaluation signal can be computed using known ground truth labels (e.g., supervised learning), predicted or estimated labels (e.g., semi-supervised or self-supervised learning), or unlabeled (e.g., unsupervised learning). The evaluation signal may be a reward (e.g., for reinforcement learning). The reward can be computed using a machine-learned reward model configured to generate a reward based on the received output(s). The reward can be computed using feedback data describing human feedback on the output(s).
[0163] In 1608, exemplary method 1600 may include updating a machine-trained model using an evaluation signal. For example, the parameter values of a machine-trained model(s) may be learned using various training or learning techniques, such as backpropagation, in some embodiments. For example, an evaluation signal may be backpropagated from the output (or another source of the evaluation signal) through the machine-trained model(s) to update one or more parameters of the model(s) (for example, based on the gradient of the evaluation signal with respect to the parameter(s)). For example, a system(s) containing one or more machine-trained models may be trained in an end-to-end manner. Gradient descent can be used to iteratively update parameters over several training iterations. In some embodiments, performing error backpropagation may include performing backpropagation with truncation over time. Exemplary method 1600 may perform several generalization techniques (e.g., load decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0164] In some embodiments, exemplary method 1600 can be performed to train a machine learning model from an initialized state to a fully trained state (for example, if the model exhibits a desired performance profile based on accuracy, precision, recall, etc.).
[0165] In some embodiments, exemplary method 1600 can be performed at a specific stage of the training procedure. For example, in some embodiments, exemplary method 1600 can be performed to pre-train a machine learning model. Pre-training may include, for example, large-scale training on potentially noisy data to achieve a broad base of performance levels across various task / data types. In some embodiments, exemplary method 1600 can be performed to fine-tune a machine learning model. Fine-tuning may include, for example, smaller-scale training on higher-quality data (e.g., labeled, curated, etc.). Fine-tuning may affect all or some of the parameters of the machine learning model. For example, different parts of the machine learning model may be “frozen” during a particular training stage. For example, parameters related to the embedding space may be “frozen” during fine-tuning (e.g., to retain information learned from a broader domain than present in the fine-tuning dataset). Exemplary fine-tuning approaches include reinforcement learning. Reinforcement learning may be based on user feedback on the model's performance in use.
[0166] Figure 17 is a block diagram of an exemplary processing flow for processing input(s) 1702 using machine learning model(s) 1701 701 and generating output(s) 1703.
[0167] A machine learning model(s) 1701 may be one or more machine learning models or model components, or may include one or more. For example, a machine learning model(s) 1701 may include a machine learning media asset generation model 1701A and / or a machine learning content item generation model 1701B. An exemplary machine learning model may include a neural network (e.g., a deep neural network). An exemplary machine learning model may include a nonlinear or linear model. An exemplary machine learning model may use other architectures instead of, or in addition to, a neural network. An exemplary machine learning model may include a decision tree-based model, a support vector machine, a hidden Markov model, a Bayesian network, a linear regression model, a k-means clustering model, and so on.
[0168] Exemplary neural networks can include feedforward neural networks, recurrent neural networks (RNNs) including long-short-term memory (LSTM) based recurrent neural networks, convolutional neural networks (CNNs), distributed models, generative adversarial networks, or other forms of neural networks. Exemplary neural networks can also be deep neural networks. Some exemplary machine-learned models can leverage attention mechanisms such as self-attention. For example, some exemplary machine-learned models can include multi-head self-attention models.
[0169] A machine learning model(s) 1701 can include one or more instances of the same model configured to work with data from input(s) 1702. A machine learning model(s) 1701 can include an ensemble of different models that can interact cooperatively to process data from input(s) 1702. For example, a machine learning model(s) 1701 can employ a mixed structure of specialties. See, for example, Zhou et al., Mixture-of-Experts with Expert Choice Routing, ARXIV:2202.09368v2 (October 14, 2022).
[0170] Input(s) 1702 can generally contain or represent various types of data. Input(s) 1702 can contain one type or many different types of data. For example, the input could contain existing media assets(s) 1702A (e.g., existing content items) and / or data resources 1702B. Output(s) 1703 can contain data of the same type(s) or different types of data compared to Input(s) 1702. Output(s) 1703 can contain one type or many different types of data. For example, Output(s) 1703 could contain media assets(s) 1704 and / or content items(s) 1705. Media assets(s) 1704 could contain, for example, text assets(s) 1704A, image assets(s) 1704B, and / or unique profile data(s) 1704C.
[0171] Exemplary data types for input(s) 1702 or output(s) 1703 include natural language text data, software code data (e.g., source code, object code, machine code, or any other form of computer-readable instructions or programming language), machine code data (e.g., binary code, assembly code, or any other form of machine-readable instructions that can be directly executed by a computer's central processing unit), assembly code data (e.g., low-level programming languages that use symbolic representations of machine code instructions to program processing units), genetic data or other chemical or biochemical data, image data, audio data, audiovisual data, tactile data, biometric data, medical data, financial data, statistical data, geographic data, astronomical data, historical data, sensor data in general (e.g., digital or analog values such as voltage or other absolute or relative level measurements from real or artificial inputs such as audio sensors, light sensors, displacement sensors, etc.). Data may be raw or processed, and may be in any format or schema.
[0172] In a multimodal input 1702 or output 1703, exemplary combinations of data types include image data and audio data, image data and natural language data, natural language data and software code data, image data and biometric data, sensor data and medical data, etc. It should be understood that any combination of data types in input 1702 or output 1703 is possible.
[0173] An exemplary input 1702 may include one or more data types, such as the exemplary data types described above. An exemplary output 1703 may include one or more data types, such as the exemplary data types described above. The data types of input 1702 may be the same as or different from the data types of output 1703. It should be understood that the exemplary data types described above are provided for illustrative purposes only. The data types contemplated within the scope of this disclosure are not limited to the examples given above.
[0174] Figure 18 is a block diagram of an exemplary embodiment of an exemplary machine learning model configured to process a sequence of information. For example, an exemplary embodiment of machine learning model(s) 1701 may include machine learning sequence processing model(s) 4. The exemplary system can pass input(s) 1702 to sequence processing model(s) 4. Sequence processing model(s) 4 may include one or more machine learning components. Sequence processing model(s) 4 can process data from input(s) 1702 to obtain an input sequence 5. Input sequence 5 may include one or more input elements 5-1, 5-2, ..., 5-M, etc., obtained from input(s) 1702. Sequence processing model(s) 4 can use prediction layers(s) 6 to process input sequence 5 and generate an output sequence 7. The output sequence 7 may include one or more output elements 7-1, 7-2, ..., 7-N, etc., generated based on the input sequence 5. The system can generate output(s) 1703 based on the output sequence 7.
[0175] A sequence processing model (or multiple models) 4 may include one or more machine learning model components configured to collect, generate, or otherwise infer sequences of information. For example, some exemplary sequence processing models in the text domain are referred to as “Large Language Models,” or LLMs. See, for example, the PaLM 2 Technical Report, GOOGLE, https: / / ai.google / static / documents / palm2techreport.pdf (nd). Other exemplary sequencing models may operate in other domains, such as the image domain, for example, referencing Dosovitskiy et al.'s *An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale*, ARXIV:2010.11929v2 (June 3, 2021); the audio domain, for example, referencing Agostinelli et al.'s *MusicLM: Generating Music From Text*, ARXIV:2301.11325v1 (January 26, 2023); or the biochemical domain, for example, referencing Jumper et al.'s *Highly accurate protein structure prediction with AlphaFold*, 596 Nature 583 (August 26, 2021). The sequencing model(s)4 may process one or more types of data simultaneously. The sequencing model(s)4 may include relatively large models (e.g., more parameters, higher computational cost), relatively small models (e.g., fewer parameters, lighter computation), or both.
[0176] In general, a sequencing model(s) 4 can use data from input(s) 1702 to obtain an input sequence(s). For example, input sequence(s) 5 may include a representation of the data from input(s) 1702 in a format understood by sequencing model(s) 4. One or more machine learning-trained components of sequencing model(s) 4 can collect data from input(s) 1702, parse the data into pieces compatible with the sequencing model(s) 4's processing architecture (e.g., via "tokenization"), and project the pieces into the input space associated with prediction layer(s) 6 (e.g., via "embedding").
[0177] Sequence processing model 4 can collect data from input 1702, parse the data into a sequence of elements, and obtain an input sequence 5. For example, a portion of the input data from input 1702 can be broken down into pieces that collectively represent the content of that portion of the input data. These pieces can provide elements for a sequence.
[0178] Elements 5-1, 5-2, ..., 5-M can, in some cases, represent building blocks for capturing or representing meaningful information in a particular data domain. For example, an element can describe an “atomic unit” across one or more domains. For instance, in the case of a text input source(s), an element might correspond to a group of one or more word or subword components, such as one or more sets of characters.
[0179] For example, elements 5-1, 5-2, ..., 5-M can represent tokens obtained using a tokenizer. For example, a tokenizer can process a given portion of an input source and output a set of tokens representing that portion of the input source (e.g., corresponding to input elements 5-1, 5-2, ..., 5-M). Various tokenization approaches can be used. For example, a text input source(s) can be tokenized using byte-pair coding (BPE) techniques. See, for example, Kudo et al., SentencePiece: A simple and language-independent subword tokenizer and detokenizer for Neural Text Processing, PROCEEDINGS OF THE 2018 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (System Demonstrations), pages 66-71 (October 31 - November 4, 2018), https: / / aclanthology.org / D18-2012.pdf. Image-based input sources can be tokenized by extracting and serializing patches from the images.
[0180] In general, any data type can be serialized and processed into input sequence 5. It should be understood that the elements 5-1, 5-2, ..., 5-M shown in Figure 18 can be tokens or their embedded representations.
[0181] The prediction layer(s) 6 can predict one or more output elements 7-1, 7-2, ..., 7-N based on the input elements. The prediction layer(s) 6 can include one or more machine learning model architectures, such as one or more layers of trained parameters that operate and transform the input(s) to extract higher-dimensional meaning from the input(s) 5-1, 5-2, ..., 5-M and the relationships between them. In this way, for example, the exemplary prediction layer(s) 6 can predict new output(s) taking into account the context provided by the input sequence 5.
[0182] Prediction layer(s) 6 can evaluate associations between parts of input sequence 5 and specific output elements. These associations can inform predictions of the likelihood that a particular output follows the input context. For example, consider the text fragment, "The carpenter's toolbox was small but heavy. It was full of ___." Prediction layer(s) 6 can identify that "it" refers back to "toolbox" by determining the relationships between each embedding. Prediction layer(s) 6 can also link "it" to attributes of the toolbox such as "small" and "heavy". Based on these associations, prediction layer(s) 6 can assign a higher probability to the word "nail" than to the word "sawdust," for example.
[0183] A transformer is an exemplary architecture that can be used in a prediction layer(s)4. See, for example, Vaswani et al., Attention Is All You Need, ARXIV:1706.03762v7 (August 2, 2023). A transformer is an example of a machine learning model architecture that uses an attention mechanism to compute associations between items in a context window. A context window can include an input sequence5 and possibly one or more output elements(s)7-1, 7-2, ..., 7-N. A transformer block can include one or more attention layers(s) and one or more post-attention layers(s) (e.g., feedforward layers(s) such as multilayer perceptrons).
[0184] The prediction layer(s) 6 may include, in addition to or instead of, a converter-based architecture, other machine learning model architectures. For example, not only convolutional neural networks (CNNs), but also recurrent neural networks (RNNs) and long-shortened memory (LSTM) models may be used. In general, the prediction layer(s) 6 can leverage various types of artificial neural networks that can understand or generate sequences of information.
[0185] The output sequence 7 may contain or represent the same or different data types as the input sequence 5. For example, input sequence 5 may represent text data, and output sequence 7 may represent text data. Input sequence 5 may represent image, audio, or audiovisual data, and output sequence 7 may represent text data (e.g., describing image, audio, or audiovisual data). It should be understood that any other intervening model component of the prediction layer 6 and the sequence processing model 4 may be configured to receive various data types in input sequence 5 and output various data types in output sequence 7.
[0186] The output sequence 7 can have various relationships with the input sequence 5. The output sequence 7 can be a continuation of the input sequence 5. The output sequence 7 can be complementary to the input sequence 5. The output sequence 7 can transform, deform, extend, or otherwise modify the input sequence 5. The output sequence 7 can respond to, evaluate, confirm, or otherwise respond to the input sequence 5. The output sequence 7 can execute (or write instructions to execute) the instructions provided through the input sequence 5.
[0187] The output sequence 7 can be generated autoregressively. For example, in some applications, the output of one or more prediction layers 6 can pass through one or more output layers (e.g., softmax layers) to obtain a probability distribution over output vocabulary (e.g., text or symbolic code) conditioned on a set of input elements in a context window. In this way, the output sequence 7 can be generated autoregressively, for example, by sampling the next most likely output element, adding that element to the context window, regenerating the probability distribution based on the updated context window, sampling the next most likely output element, and so on.
[0188] Output sequence 7 can also be generated non-autoregressively. For example, multiple output elements of output sequence 7 can be predicted together without explicit sequential conditioning of each other. See, for example, Saharia et al., Non-Autoregressive Machine Translation with Latent Alignments, ARXIV:2004.07437v3 (November 16, 2020).
[0189] The output sequence 7 can contain one or more parts or elements. In an exemplary content generation configuration, the output sequence 7 can contain multiple elements (e.g., sentences of text, discretized waveform values, computer code, etc.) corresponding to multiple parts of the generated output sequence. In an exemplary classification configuration, the output sequence 7 can contain a single element associated with the classification output. For example, the output "Vocabulary" can contain a set of classes to which the input sequence is classified. For example, a visual transducer block can pass latent state information to a multilayer perceptron that outputs class values that are likely to be associated with the input image.
[0190] Figure 19 is a block diagram of an exemplary technique for arranging an exemplary input sequence 8. The input sequence 8 may include various functional elements that form part of the model infrastructure, such as element 8-0 obtained from a task indicator 9, which signals to any model(s) that are processing the input sequence 8 on which a particular task is being performed (for example, to help fit the performance of the model(s) to that particular task). The input sequence 8 may include various data elements from different data modalities. For example, input modality 10-1 may include one modality of data. The data-versus-sequence model 11-1 processes the data from input modality 10-1 and projects the data into a format compatible with input sequence 8 (for example, one or more vectors of dimensions following the dimensions of input sequence 8), obtaining elements 8-1, 8-2, and 8-3. Another input modality 10-2 may include a different modality of data. The data-versus-sequence model 11-2 projects the data from input modality 10-2 into a format compatible with input sequence 8, obtaining elements 8-4, 8-5, and 8-6. Another input modality 10-3 may contain yet another different modality of data. The data-versus-sequence model 11-3 projects the data from input modality 10-3 into a format compatible with input sequence 8, obtaining elements 8-7, 8-8, and 8-9.
[0191] Input sequence 8 may be the same as or different from input sequence 5. Input sequence 8 can be a multimodal input sequence containing elements that represent data from different modalities using a common dimensional representation. For example, the embedding space may have P dimensions. Input sequence 8 can be configured to contain multiple elements having P dimensions. In this way, for example, exemplary embodiments can facilitate information extraction and categorization across diverse data modalities by projecting data onto elements in the same embedding space for comparison, combination, or other computations between them.
[0192] For example, elements 8-0, ..., 8-9 can represent specific locations within a multidimensional embedding space. Some elements can be mapped to a set of discrete locations within the embedding space. For example, elements corresponding to individual members of a given vocabulary of tokens can be mapped to individual locations within the embedding space associated with those tokens. Other elements can be contiguously distributed across the embedding space. For example, some data types can be decomposed into contiguously defined parts (e.g., image patches) that can be described using contiguously distributed locations within the embedding space.
[0193] In some embodiments, the expressiveness of an embedding space may not be limited to the meaning associated with any particular set of tokens or other building blocks. For example, a continuous embedding space can encode a spectrum of higher-order information. Individual pieces of information (e.g., tokens) can be mapped to specific points within that space. For example, the token for the word "dog" can be projected onto an embedding value that points to a specific location within the embedding space associated with a dog. Similarly, an image patch of a picture of a dog on grass may also be projected onto the embedding space. In some embodiments, the projection of the dog image may be similar to the projection of the word "dog," while also having similarity to the projection of the word "grass," and at the same time, both may be different. In some embodiments, the projection of an image patch may not precisely match a single projection of either of the single words. In some embodiments, the projection of an image patch can match a combination of projections of the words "dog" and "grass." In this way, for example, a higher-order embedding space can encode information that can be independent of the data modality in which the information is represented.
[0194] The task indicator 9 may include a model or model component configured to identify the task being executed and inject input values represented by elements 8-0, which signal which task is being executed, into the input sequence 8. For example, the input values may be provided as a data type associated with an input modality and projected along that input modality (e.g., the input values may be text task labels embedded in the input along with other text data, or the input values may be pixel-based representations of tasks embedded in the input along with other image data, etc.). The input values may be provided as a data type that is different from, or at least independent of, other inputs(s). For example, the input values represented by elements 8-0 may be learned within a contiguous embedding space.
[0195] Input modalities 10⁻¹, 10⁻², and 10⁻³ can be associated with various different data types (for example, as previously mentioned with respect to input(s) 1702 and output(s) 1703).
[0196] The data-sequence models 11-1, 11-2, and 11-3 may be the same or different from each other. Each of the data-sequence models 11-1, 11-2, and 11-3 can be adapted to each of the respective input modalities 10-1, 10-2, and 10-3. For example, a text data-sequence model can subdivide a portion of the input text and project that subdivision onto elements (or more) in the input sequence 8 (e.g., elements 8-1, 8-2, 8-3, etc.). An image data-sequence model can subdivide an input image and project that subdivision onto elements (or more) in the input sequence 8 (e.g., elements 8-4, 8-5, 8-6, etc.). A data-sequence model of any data type can subdivide the input of that data type and project that subdivision onto elements (or more) within the input sequence (e.g., elements 8-7, 8-8, 8-9, etc.).
[0197] The data-sequence models 11-1, 11-2, and 11-3 can form part of a machine learning-based sequencing model(s) 4. The data-sequence models 11-1, 11-2, and 11-3 can be trained jointly with or independently of the machine learning-based sequencing model(s) 4. The data-sequence models 11-1, 11-2, and 11-3 can be trained end-to-end with the machine learning-based sequencing model(s) 4.
[0198] Figure 20 is a block diagram of an exemplary model development platform 12 that can facilitate the creation, fitting, and enhancement of exemplary machine learning models (e.g., machine learning model(s) 1701, sequence processing model(s) 4, etc.). The model development platform 12 can provide several different toolkits that the developer system can employ in the development of new or fitted machine learning models.
[0199] The model development platform 12 may provide one or more model libraries 13 containing building blocks for new models. The model library 13 may include one or more pre-trained foundational models 13-1 that can provide a backbone of processing power across various tasks. The model library 13 may also include one or more pre-trained expert models 13-2 that can focus on performance in a specific domain of expertise. The model library 13 may also include various model primitives 13-3 that can provide low-level architectures or components (optionally pre-trained) that can be assembled into various configurations as needed.
[0200] The model development platform 12 can receive a selection of various model components 14. The model development platform 12 can pass the selected model components 14 to the workbench 15, which then combines the selected model components 14 into the development model 16.
[0201] The workbench 15 can facilitate further enhancement and adaptation of the development model 16 by leveraging several different toolkits integrated with the model development platform 12. For example, the workbench 15 can use the model alignment toolkit 17 to facilitate the alignment of the development model 16 with desired performance profiles for various tasks.
[0202] The model alignment toolkit 17 can provide several tools to cause the development model 16 to produce outputs aligned with desired operating characteristics. Alignment may include improving the accuracy, precision, recall, etc., of the model output. Alignment may include enforcing the output style, schema, or other preferred characteristics of the model output. Alignment may be general or domain-specific. For example, a pre-trained base model 13-1 may start with an initial level of performance across multiple domains. Alignment of the pre-trained base model 13-1 may include improving performance in a specific domain of information or task (for example, at the expense of performance in another domain of information or task).
[0203] The model alignment toolkit 17 can integrate one or more datasets 17-1 for the alignment of the development model 16. The selected datasets 17-1 may include labeled or unlabeled training data. The datasets 17-1 can be obtained from public domain datasets. The datasets 17-1 can be obtained from private datasets associated with one or more developer systems for the alignment of machine learning models customized for specific users for private use cases.
[0204] The pre-training pipeline 17-2 may include a machine learning model training workflow configured to update the development model 16 across a large and potentially noisy dataset. For example, pre-training can utilize unsupervised learning techniques (e.g., denoising) to process a large number of training instances to update model parameters from an initialized state and achieve the desired baseline performance. The pre-training pipeline 17-2 can perform pre-training using an unlabeled dataset within a dataset 17-1. The workbench 15 can execute the pre-training pipeline 17-2 to pre-train the development model 16.
[0205] The fine-tuning pipeline 17-3 may include a machine learning model training workflow configured to refine the model parameters of the development model 16 with higher quality data. The fine-tuning pipeline 17-3 can update the development model 16 by performing supervised training using the labeled dataset(s) in the dataset(s) 17-1. The fine-tuning pipeline 17-3 can update the development model 16 by performing reinforcement learning using reward signals from user feedback signals. The workbench 15 can execute the fine-tuning pipeline 17-3 to fine-tune the development model 16.
[0206] The prompt library 17-4 may include a set of inputs configured to elicit behavior consistent with a desired performance criterion. The prompt library 17-4 may include short-shot prompts (e.g., inputs that provide examples of desired model outputs to add to a desired runtime query), thought-chain prompts (e.g., inputs that provide step-by-step inference within representative examples to facilitate complete inference by the model), and so on.
[0207] Exemplary prompts can be retrieved from the available repositories of the prompt library 17-4. Exemplary prompts may be contributed by one or more developer systems using the workbench 15.
[0208] In some embodiments, a pre-trained or fine-tuned model can achieve satisfactory performance even when the input lacks representative examples. For example, a zero-shot prompt may contain input that lacks representative examples. A zero-shot prompt may be located within or outside the training domain(s) of the training dataset.
[0209] The prompt library 17-4 may include one or more prompt engineering tools. The prompt engineering tools can provide workflows for extracting or learning optimized prompt values. The prompt engineering tools can facilitate direct learning of prompt values (e.g., input element values) based on one or more training iterations. The workbench 15 can implement the prompt engineering tools on the development model 16.
[0210] The prompt library 17-4 can include a pipeline for prompt generation. For example, inputs can be generated using the development model 16 itself or other machine learning models. In this way, for example, the first model can process information about a task and output inputs that the second model processes to execute the steps of the task. The second model may be the same as or different from the first model. The workbench 15 can implement the prompt generation pipeline with the development model 16.
[0211] The prompt library 17-4 may include a pipeline for context injection. For example, the performance of the development model 16 for a particular task can be improved if additional context is provided for performing that task. The prompt library 17-4 may include software components configured to identify a desired context, retrieve the context from an external source (e.g., a database, a sensor, etc.), and add the context to an input prompt. The workbench 15 can implement the context injection pipeline in the development model 16.
[0212] The various training examples described herein with respect to the model development platform 12 refer to “pre-training” and “fine-tuning,” but it should be understood that the model fitting toolkit 17 can generally support a wide range of training techniques adapted for training a wide range of machine learning models. Exemplary training techniques can correspond to the exemplary training methods 1600 described above.
[0213] The model development platform 12 may include a model plugin toolkit 18. The model plugin toolkit 18 may include a variety of tools configured to extend the functionality of machine-learned models by integrating them with other systems, devices, and software components. For example, machine-learned models can use tools to improve the quality of their performance where appropriate. For example, deterministic tasks can be offloaded to dedicated tools instead of probabilistically performing tasks with an increased risk of error. For example, instead of autoregressively predicting the solution to a system of linear equations, a machine-learned model can recognize which tool to call to obtain the solution and pass the system of equations to the appropriate tool. The tool may be a conventional system of equations solver that can act deterministically to solve the system of equations. The output of the tool may be returned in response to the original query. In this way, by using tools, some exemplary models can focus on the strengths of the machine-learned model, such as understanding the intent of unstructured requests to a task, while at the same time increasing the model's performance by offloading specific tasks to more focused tools for the mechanical application of deterministic algorithms to well-defined problems.
[0214] The Model Plugin Toolkit 18 may include a validation tool 18-1. The validation tool 18-1 may include a tool capable of analyzing and verifying the output(s) of a machine-trained model. The validation tool 18-1 may include engineered heuristics that establish specific thresholds to apply to the model output. For example, the validation tool 18-1 may ground the output of a machine-trained model to a structured data source (e.g., to mitigate "hallucinations").
[0215] The model plugin toolkit 18 may include a tool package 18-2 for executing one or more tools, which may include scripts or other executable code that can be run alongside the development model 16. The tool package 18-2 may include one or more inputs configured to cause a machine-trained model(s) to execute the tool (e.g., a few-shot prompt that guides the model to output a tool call in the appropriate syntax). The tool package 18-2 may include, for example, fine-tuned training data for training the model to use the tool.
[0216] The model plug-in toolkit 18 may include an interface for calling an external application programming interface (API) 18-3. For example, in addition to directly executing tool calls or tool code in the development model 16, or instead, the development model 16 may be configured to output instructions that initiate API calls to send or retrieve data via an external system.
[0217] The model plugin toolkit 18 can integrate with the prompt library 17-4 to build a catalog of available tools for use in the development model 16. For example, the model can receive a catalog of available tools as input, and the model can select a tool from the available tools and generate an output that initiates a tool call to use that tool.
[0218] The model development platform 12 may include a computational optimization toolkit 19 for optimizing the computational performance of the development model 16. For example, tools for model compression 19-1 may allow for reducing the size of the development model 16 while maintaining a desired level of performance. For example, model compression 19-1 may include quantization workflows, weight pruning, and sparsification techniques. Tools for hardware acceleration 19-2 may facilitate the configuration of model storage and executables to work optimally with different hardware resources. For example, hardware acceleration 19-2 may include tools for optimally sharing the model for distributed processing across multiple processing units due to increased bandwidth, reduced integrated memory requirements, etc. Tools for distillation 19-3 may be provided for training a lighter model based on the knowledge encoded in the development model 16. For example, the development model 16 may be a high-performance, large-scale machine learning model optimized using the model development platform 12. To obtain a lightweight model for execution in resource-constrained environments, the smaller model may be a "student model," which learns to mimic the development model 16 as a "teacher model." In this way, for example, investment in learning the parameters and configuration of development model 16 can be efficiently shifted to smaller models in relation to more efficient estimation.
[0219] Workbench 15 may perform one or more of the toolkits performed on the model development platform 12, or it may perform none of them. Workbench 15 can output an output model 20 based on the development model 16. The output model 20 may be an expanded version of the development model 16. The output model 20 may be a development or training checkpoint for the development model 16. The output model 20 may be a distilled, compressed, or otherwise optimized version of the development model 16.
[0220] Figure 21 is a block diagram of an exemplary training flow for training a machine learning-prepared development model 16. One or more parts of the exemplary training flow may be performed by a computing system including one or more computing devices, such as the computing system described with reference to other figures. Each of the parts of the exemplary training flow may be performed by any (or any combination of) one or more computing devices. Furthermore, one or more parts of the exemplary training flow may be performed on the hardware components of the devices described herein, for example, to train one or more systems or models. Figure 21 shows the elements to be performed in a particular order for illustrative and explanatory purposes. Those skilled in the art will understand that, using the disclosures provided herein, any element of any of the methods discussed herein may be adapted, rearranged, extended, omitted, combined, or modified in various ways without departing from the scope of this disclosure. Figure 21 is described for the illustrative purposes and is not intended to limit to the elements / terms described in relation to other systems and figures. One or more parts of the exemplary training flow may be performed by other systems, additionally or alternatively.
[0221] First, the development model 16 can retain its initial state as the initialized model 21. The development model 16 can be initialized with weight values. The initial weight values can be random or based on an initialization schema. The initial weight values can be based on previous pre-training on the same or different models.
[0222] The initialized model 21 can undergo pre-training in the pre-training phase 22. The pre-training phase 22 can be performed using one or more pre-training pipelines 17-2 on data from a dataset 17-1. For example, if the initialized model 21 has already been pre-trained (e.g., the development model 16 includes, is, or is based on, a pre-trained foundational or expert model), then pre-training can be omitted.
[0223] The pre-trained model 23 can then become a new version of the development model 16, which can persist as the development model 16 or as a new development model. The pre-trained model 23 can be set to an initial state if the development model 16 has already been pre-trained. The pre-trained model 23 can undergo fine-tuning in the fine-tuning stage 24. The fine-tuning stage 24 can be performed using one or more fine-tuning pipelines 17-3 on data from a dataset 17-1. For example, fine-tuning can be omitted if the performance of the pre-trained model is sufficient, if the model has already been fine-tuned, or if another fine-tuning approach is preferred.
[0224] The fine-tuned model 29 can then become a new version of the development model 16, which can persist as development model 16 or as a new development model. The fine-tuned model 29 can be the initial state if the development model 16 has already been fine-tuned. The fine-tuned model 29 can undergo refinement through user feedback 26. For example, refinement through user feedback 26 may optionally include reinforcement learning based on human feedback from human users of the fine-tuned model 25. It should be understood that since reinforcement learning can be a form of fine-tuning, the fine-tuning stage 24 may include a stage for refinement through user feedback 26. Refinement through user feedback 26 may result in the creation of a refined model 27. The refined model 27 can be output to a downstream system 28 for deployment or further development.
[0225] In some embodiments, computational optimization operations may be applied before, during, or after each stage. For example, an initialized model 21 may undergo computational optimization 29-1 (e.g., using the computational optimization toolkit 19) before the pre-training stage 22. A pre-trained model 23 may undergo computational optimization 29-2 (e.g., using the computational optimization toolkit 19) before the fine-tuning stage 24. A fine-tuned model 25 may undergo computational optimization 29-3 (e.g., using the computational optimization toolkit 19) before refinement by user feedback 26. A refined model 27 may undergo computational optimization 29-4 (e.g., using the computational optimization toolkit 19) before output to a downstream system 28. The computational optimizations 29-1, ..., 29-4 may all be the same, all be different, or include at least some different optimization techniques.
[0226] Figure 22 is a block diagram of an estimation system for running one or more machine learning models 1701 to perform estimations (for example, for training, deployment, etc.). A model host 31 can receive machine learning models 1701. A model host 31 can host one or more model instances 31-1, which can be one or more instances of one or more models. A model host 31 can host model instances 31-1 using available computing resources 31-2 associated with the model host 31.
[0227] The model host 31 can perform estimations on behalf of one or more clients 32. Clients 32 can send input requests 33 to the model host 31. Using the input requests 33, the model host 31 can obtain inputs 1702 for input to machine-trained models 1701. Machine-trained models 1 can process the inputs 1702 to produce outputs 1703. Using the outputs 1703, the model host 31 can return an output payload 34 to respond to the input requests 33 from clients 32. The output payload 34 may contain or be based on the outputs 1703.
[0228] The model host 31 can extend its estimation tasks by utilizing various other resources and tools. For example, the model host 31 can communicate with a tool interface 35 to facilitate the use of tools by model instances 31-1. The tool interface 35 may include local or remote APIs. The tool interface 35 may include integrated scripts or other software functions. The model host 31 can engage with an online learning interface 36 to facilitate the continuous improvement of machine-learned models 1701. For example, the online learning interface 36 may be used within a reinforcement learning loop to extract user feedback on estimations made by the model host 31. The model host 31 may access runtime data sources 37 to extend inputs 1702 with additional contextual information. For example, the runtime data source 37 may include a knowledge graph 37-1 that facilitates the extraction of structured information for information related to input requests 33 (e.g., a search engine service). The runtime data source(s) 37 may include public or private external or local database(s) 37-2 that can store information related to input requests(s) 33 for extending the input(s) 1702. The runtime data source(s) 37 may also include account data(s) 37-3, which can be retrieved in relation to user accounts corresponding to clients 32 in order to customize the behavior of the model host(s) 31 accordingly.
[0229] Model host 31 may be implemented by one or more computing devices or systems. Client(s) 2 may be implemented by one or more computing devices or systems, which may include computing devices or systems shared with model host 31.
[0230] For example, the model host 31 can operate on a server system that provides machine learning services to client devices (multiple) running client(s) 32 (for example, across a local or wide area network). Client devices (multiple) may be end-user devices used by individuals. Client devices (multiple) may be a server system that operates client(s) 32 and provides various functions as services to downstream end-user devices.
[0231] In some embodiments, the model host 31 can operate on the same device or system as the client(s) 32. The model host 31 can be a machine learning service that runs on a device and provides machine learning capabilities to one or more applications running on client devices, which may include application-executing client(s) 32. The model host 31 may be part of the same application as the client(s) 32. For example, the model host 31 may be a subroutine or method executed by one part of the application, and the client(s) 32 may be another subroutine or method that engages with the model host 31 to perform estimation functions within the application. It should be understood that the model host 31 and client(s) 32 can have a variety of different configurations.
[0232] A model instance 31-1 may contain one or more machine learning models available for performing estimations. A model instance 31-1 may contain weights or other model components that are stored in persistent storage, temporarily cached, or loaded into high-speed memory. A model instance 31-1 may contain multiple instances of the same model (for example, to run more requests in parallel on the same model). A model instance 31-1 may contain instances of different models. A model instance 31-1 may contain cached intermediate states of active or inactive models that are used to speed up estimations of these models. For example, an estimation session with a particular model may generate a significant amount of computational results that can be reused for future estimation runs (for example, using a KV cache for a converter-based model). These computational results can be saved in relation to that estimation session so that the session can run more efficiently when resumed.
[0233] A computational resource(s) 31-2 may include one or more processors (such as a central processing unit, a graphical processing unit, a tensor processing unit, or a machine learning accelerator) connected to one or more memory devices. A computational resource(s) 31-2 may include a dynamic pool of available resources shared with other processes. A computational resource(s) 31-2 may include a memory device large enough to fit an entire model instance into a single memory instance. A computational resource(s) 31-2 may also model an instance(s) across multiple memory devices (for example, using data parallelism or tensor parallelism). This can be done to increase parallelism or run larger models by using multiple memory devices that may not individually be able to fit the entire model into memory.
[0234] Input request 33 may contain data for input(s) 1702. Model host 31 can process input request 33 and retrieve input(s) 1702. Input(s) 2 can be obtained directly from input request 33 or retrieved using input request 33. Input request 33 may be submitted to model host 31 via API.
[0235] The model host 31 can perform estimations in parallel across batches of input requests 33. For example, a model instance 31-1 can consist of an input structure having a batch dimension. Separate inputs 1702 can be distributed across the batch dimension (e.g., rows of an array). Separate inputs 1702 can contain completely different contexts. Separate inputs 1702 can be multiple estimation steps of the same task. Separate inputs 1702 can be arranged alternately within the input structure, so that any given estimation cycle can act on different parts of each input 1702. In this way, for example, the model host 31 can perform estimations in parallel for batches, and as a result, outputs 1703 can also include a batch dimension and return estimation results for the batched inputs 1702 in parallel. In this way, for example, batches of input requests 33 can be processed in parallel for higher throughput of output payloads 34.
[0236] The output payload 34 may contain or be based on the outputs 1703 from machine learning model(s) 1701. The model host 31 can process the outputs 1703 to obtain the output payload 34. This may involve chaining multiple estimations (e.g., iteratively, recursively, across the same model(s) or different models(s)) to arrive at the final output of the task returned in the output payload 34. The output payload 34 can be sent to client(s) 32 via the API.
[0237] The online learning interface(s) 36 can facilitate reinforcement learning of the machine learning models(s) 1701. The online learning interface(s) 36 can facilitate reinforcement learning with human feedback (RLHF). The online learning interface(s) 36 can facilitate federated learning of the machine learning models(s) 1701.
[0238] The model host 31 can run a machine learning model(s) 1701 to perform estimations for various tasks using various types of data. For example, various different inputs(s) 1702 and outputs(s) 1703 can be used for various different tasks. In some embodiments, the input(s) 1702 can be image data or an alternative representation of image data. A machine learning model(s) 1 can process image data to generate an output. For example, a machine learning model(s) 1701 can process image data to generate an image recognition output (e.g., recognition of image data, latent embedding of image data, encoded representation of image data, hash of image data, etc.). As another example, a machine learning model(s) 1701 can process image data to generate an image segmentation output. As yet another example, a machine learning model(s) 1701 can process image data to generate an image classification output. As yet another example, a machine learning model(s) 1701 can process image data to generate an image data modification output (e.g., modification of image data, etc.). As another example, a machine learning model(s)1701 can process image data to generate encoded image data output (e.g., an encoded / compressed representation of the image data). As yet another example, a machine learning model(s)1701 can process image data to generate upscaled image data output. As yet another example, a machine learning model(s)1701 can process image data to generate predictive output.
[0239] In some embodiments, the task is a computer vision task. In some cases, the input(s) 1702 includes pixel data from one or more images, and the task is an image processing task. For example, the image processing task may be image classification, and the output is a set of scores, each corresponding to a different object class, representing the probability that one or more images depict an object belonging to that object class. The image processing task may be object detection, and the image processing output identifies one or more regions within one or more images, and for each region, the probability that the region depicts an object of interest. As another example, the image processing task may be image segmentation, and the image processing output defines, for each pixel in one or more images, the probability of each category in a given set of categories. For example, the set of categories may be foreground and background. As yet another example, the set of categories may be object classes. As yet another example, the image processing task may be depth estimation, and the image processing output defines, for each pixel in one or more images, the respective depth value. As another example, the image processing task could be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene depicted between the pixels in the network input, for each pixel in the input images.
[0240] In some embodiments, the input(s) 1702 may be natural language data or represent natural language data in a different way. A machine learning model(s) 1 can process the natural language data to produce an output. For example, a machine learning model(s) 1701 can process natural language data to produce a language encoding output. As another example, a machine learning model(s) 1701 can process natural language data to produce a latent text embedding output. As yet another example, a machine learning model(s) 1701 can process natural language data to produce a transformation output. As yet another example, a machine learning model(s) 1701 can process natural language data to produce a classification output. As yet another example, a machine learning model(s) 1701 can process natural language data to produce a text segmentation output. As yet another example, a machine learning model(s) 1701 can process natural language data to produce a semantic intent output. As another example, a machine learning model(s)1701 can process natural language data to produce upscaled text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language). As yet another example, a machine learning model(s)1701 can process natural language data to produce predictive output (e.g., one or more predicted next parts of natural language content).
[0241] In some embodiments, the input(s) 1702 may be audio data (e.g., audio data, text data, etc., data describing spoken natural language) or may represent the audio data in a different way. A machine learning model(s) 1 may process the audio data to produce an output. For example, a machine learning model(s) 1701 may process the audio data to produce a speech recognition output. As another example, a machine learning model(s) 1701 may process the audio data to produce a speech translation output. As yet another example, a machine learning model(s) 1701 may process the audio data to produce a potential embedding output. As yet another example, a machine learning model(s) 1701 may process the audio data to produce an encoded audio output (e.g., an encoded and / or compressed representation of the audio data). As yet another example, a machine learning model(s) 1701 may process the audio data to produce an upscaled audio output (e.g., audio data of higher quality than the input audio data). As another example, a machine learning model(s)1701 can process audio data to generate a textual representation output (e.g., a textual representation of the input audio data). As yet another example, a machine learning model(s)1701 can process audio data to generate a predictive output.
[0242] In some embodiments, the input(s) 1702 may be latent encoded data (e.g., a latent spatial representation of the input) or may represent the latent encoded data in a different way. A machine learning model(s) 1 can process the latent encoded data to produce an output. For example, a machine learning model(s) 1701 can process the latent encoded data to produce a recognition output. As another example, a machine learning model(s) 1701 can process the latent encoded data to produce a reconstruction output. As yet another example, a machine learning model(s) 1701 can process the latent encoded data to produce a speech output. As yet another example, a machine learning model(s) 1701 can process the latent encoded data to produce a reclustered output. As yet another example, a machine learning model(s) 1701 can process the latent encoded data to produce a prediction output.
[0243] In some embodiments, the input(s) 1702 may be statistical data or represent statistical data in a different way. The statistical data may be computer-processed and / or computed data from several other data sources, or represent it, or include it in a different way. A machine learning model(s) 1 may process the statistical data to produce an output. For example, a machine learning model(s) 1701 may process statistical data to produce a recognition output. As another example, a machine learning model(s) 1701 may process statistical data to produce a prediction output. As yet another example, a machine learning model(s) 1701 may process statistical data to produce a classification output. As yet another example, a machine learning model(s) 1701 may process statistical data to produce a segmentation output. As yet another example, a machine learning model(s) 1701 may process statistical data to produce a visualization output. As yet another example, a machine learning model(s) 1701 may process statistical data to produce a diagnostic output.
[0244] In some embodiments, the input(s) 1702 may be sensor data or represent the sensor data in a different way. A machine learning model(s) 1 can process the sensor data to generate an output. For example, a machine learning model(s) 1701 can process sensor data to generate a recognition output. As another example, a machine learning model(s) 1701 can process sensor data to generate a prediction output. As yet another example, a machine learning model(s) 1701 can process sensor data to generate a classification output. As yet another example, a machine learning model(s) 1701 can process sensor data to generate a segmentation output. As yet another example, a machine learning model(s) 1701 can process sensor data to generate a visualization output. As yet another example, a machine learning model(s) 1701 can process sensor data to generate a diagnostic output. As yet another example, a machine learning model(s) 1701 can process sensor data to generate a detection output.
[0245] In some embodiments, a machine learning model(s)1701 may be configured to perform tasks that include encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task may be an audio compression task. The input may include audio data, and the output may include compressed audio data. In another example, the input may include visual data (e.g., one or more images or videos), and the output may include compressed visual data, and the task is a visual data compression task. In yet another example, the task may include generating embeddings for input data (e.g., input audio or visual data). In some cases, the input may include audio data representing spoken utterances, and the task is a speech recognition task. The output may include text output mapped to the spoken utterances. In some cases, the task may include encrypting or decrypting input data. In some cases, the task may include a microprocessor performance task, such as branch prediction or memory address translation.
[0246] In some embodiments, the task is a generative task, and a machine-trained model(s) 1701 may be configured to output content generated considering the input(s) 1702. For example, the input(s) 1702 may be data in one or more formats that encode a context for generating additional content, or may be represented otherwise.
[0247] In some embodiments, the task can be a text completion task. A machine learning model 1 can be configured to process an input 1702 representing text data and to produce an output 1703 representing additional text data that completes the text sequence containing the input 1702. For example, a machine learning model 1701 can be configured to produce an output 1703 to complete a sentence, paragraph, or portion of text that follows from the portion of text represented by the input 1702.
[0248] In some embodiments, a task can be a command that follows a task. A machine learning model 1 may be configured to process an input 1702 representing a command that performs a function and to produce an output 1703 that advances the goal of satisfying the command function (e.g., at least one step of a multi-step procedure for performing the function). The output 3 may represent data of the same or different modality as the input 1702. For example, the input 1702 may represent text data (e.g., a natural language command for the task to be performed), and the machine learning model 1701 may process the input 1702 to produce an output 1703 representing text data in response to the command (e.g., a natural language response, a programming language response, a machine language response, etc.). Input(s) 2 can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by text instructions), and a machine learning model(s) 1701 can process input(s) 1702 to generate outputs(s) 1703 representing text data (e.g., natural language responses, programming language responses, machine language responses, etc.) in response to the instructions. One or more outputs(s) 1703 may be generated iteratively or recursively to sequentially process and achieve steps toward achieving the requested function. For example, an initial output may be executed by an external system or processed by a machine learning model(s) 1701 to complete the initial steps toward performing the function. Multiple steps may be executed until a final output in response to the first instruction is obtained.
[0249] In some embodiments, the task may be a question-answering task. A machine-trained model 1 may be configured to process an input 1702 representing a question to be answered and to produce an output 1703 that advances the goal of returning an answer to the question (e.g., at least one step of a multi-step procedure for performing a function). The output 3 may represent data of the same or different modality as the input 1702. For example, the input 1702 may represent text data (e.g., natural language instructions for the task to be performed), and the machine-trained model 1701 may process the input 1702 to produce an output 1703 representing text data in response to the question (e.g., natural language response, programming language response, machine language response, etc.). Input(s) 2 can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by text instructions), and a machine learning model(s) 1701 can process input(s) 1702 to generate outputs(s) 1703 representing text data (e.g., natural language responses, programming language responses, machine language responses, etc.) in response to a question. One or more outputs(s) 1703 may be generated iteratively or recursively to sequentially process and achieve steps toward answering the question. For example, an initial output may be performed by an external system or processed by a machine learning model(s) 1701 to complete an initial step toward obtaining an answer to the question (e.g., querying a database, performing computer processing, executing a script, etc.). Multiple steps may be performed until a final output responding to the question is obtained.
[0250] In some embodiments, the task can be an image generation task. A machine learning model 1 can be configured to process an input 1702 that represents a context about a desired portion of an image content. The context can include text data, image data, audio data, etc. The machine learning model 1 can be configured to produce an output 1703 that represents image data depicting an image related to the context. For example, a machine learning model 1701 may be configured to generate pixel data for an image. The values of the channels associated with the pixels in the pixel data can be selected based on the context (for example, based on probabilities determined based on the context).
[0251] In some embodiments, the task can be an audio generation task. A machine learning model 1 can be configured to process an input 1702 that represents context about a desired portion of audio content. The context can include text data, image data, audio data, etc. The machine learning model 1 can be configured to produce an output 1703 that represents audio data relevant to the context. For example, a machine learning model 1701 may be configured to produce waveform data in the form of an image (e.g., a spectrogram). The channel values related to the pixels of the image can be selected based on the context. The machine learning model 1 can be configured to produce waveform data in the form of a sequence of discrete samples of a continuous waveform. The sequence values can be selected based on the context (e.g., based on probabilities determined based on the context).
[0252] In some embodiments, the task can be a data generation task. A machine learning model(s) 1 can be configured to process input(s) 1702 representing a context about a desired portion of data (e.g., data from various data domains such as sensor data, image data, multimodal data, statistical data). The desired data can be, for example, synthetic data for training other machine learning models. The context can include any data type(s). A machine learning model(s) 1 can be configured to produce outputs(s) 1703 representing data aligned with the desired data. For example, a machine learning model(s) 1701 may be configured to produce data values for arranging a dataset. The values about the data objects(s) can be selected based on the context (e.g., based on probabilities determined based on the context).
[0253] Figure 23 is a block diagram of an exemplary networked computing system that can implement an embodiment of the exemplary embodiments of the present disclosure. The system may include several computing devices and systems that are communicably coupled over a network 49. An exemplary computing device 50 is described to provide an example of a computing device that can implement any embodiment of the present disclosure (e.g., a model host 31, a client(s) 32, or both). An exemplary server computing system 60 is described as an example of a server computing system that can implement any embodiment of the present disclosure (e.g., a model host 31, a client(s) 32, or both). The computing device 50 and the server computing system(s) 60 can cooperate and interact (e.g., over the network 49) to implement any embodiment of the present disclosure (e.g., a model host 31, a client(s) 32, or both). A model development platform system 70 is an exemplary system that can host or supply a model development platform(s) 12 for developing machine learning models. The third-party system(s) 80 is an exemplary system(s) that any of the computing device(s) 50, server computing system(s) 60, or model development platform system(s) 70 can interact with in various aspects of the implementation of this disclosure (e.g., engagement of third-party tools, access to third-party databases or other resources).
[0254] Network 49 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or any combination thereof, and may include any number of wired or wireless links. Generally, communication over Network 49 can be carried out over any type of wired or wireless communication using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encoding or formatting (e.g., HTML, XML), or protection schemes (e.g., VPN, Secure HTTP, SSL). Network 49 can also be carried out over a system bus. For example, one or more devices or systems in Figure 23 may be located in the same place as, included in, or otherwise integrated with one or more other devices or systems.
[0255] The user computing device 50 can be any type of computing device, such as a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, a server computing device, a virtual machine running on a host device, or any other type of computing device. The computing device 50 can be a client computing device. The computing device 50 can be an end-user computing device. The computing device 50 can be a computing device for services provided, providing services to an end user (who may use another computing device to interact with the computing device 50).
[0256] The server computing device 50 may include one or more processors 51 and memory 52. The processor(s) 51 may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple processors connected in an operable manner. The memory 52 may include one or more non-temporary computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 52 may store data 53 and instructions 54 that can be executed by the processor(s) 51 to cause the user computing device 50 to perform operations. The operation may implement any one or more features described herein. The operation may implement exemplary methods and techniques described herein.
[0257] The computing device 50 may also include one or more input components that receive user input. For example, a user input component may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component may perform the role of implementing a virtual keyboard. Other exemplary user input components include a microphone, a camera, LiDAR, a physical keyboard or other buttons, or other means by which the user can provide user input.
[0258] The computing device 50 may store or contain one or more machine learning models 55. The machine learning models 55 may include one or more machine learning models 1701, such as a sequencing model 4. The machine learning models 55 may include one or more model instances 31-1. The machine learning models 55 may be received from a server computing system 60, a model development platform system 70, a third-party system 80 (e.g., an application distribution platform), or developed locally on the computing device 50. The machine learning models 55 may be loaded into memory 52 and used or otherwise implemented by a processor 51. The computing device 50 may implement multiple parallel instances of the machine learning models 55.
[0259] A server computing system (or multiple systems) 60 may include one or more processors 61 and a memory 62. The processor(s) 61 may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple operably connected processors. The memory 62 may include one or more non-temporary computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 62 may store data 63 and instructions 64 that can be executed by the processor(s) 61 to cause the server computing system (or multiple systems) 60 to perform operations. The operation may implement any one or more features described herein. The operation may implement exemplary methods and techniques described herein.
[0260] In some embodiments, the server computing system 60 includes one or more server computing devices, or is otherwise implemented. In examples where the server computing system 60 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or any combination thereof.
[0261] The server computing system 60 may store or otherwise include one or more machine learning models 65. The machine learning model(s) 65 may be the same as or different from the machine learning model(s) 55. The machine learning model 65 may include one or more machine learning model(s) 1701, such as the sequencing model 4. The machine learning model 65 may include one or more model instances(s) 31-1. The machine learning model(s) 65 may be received from a computing device 50, a model development platform system 70, a third-party system(s) 80, or developed locally on the server computing system(s) 60. The machine learning model(s) 65 may be loaded into memory 62 and used or otherwise implemented by a processor(s) 61. The server computing system(s) 60 may implement multiple parallel instances of the machine learning model(s) 65.
[0262] In an exemplary configuration, the machine learning model 65 may be contained within or otherwise stored and implemented in a server computing system 60 to establish a client-server relationship with a computing device 50 for providing model estimations. For example, the server computing system 60 may implement a model host 31 on the computing device 50 on behalf of a client 32. For example, the machine learning model 65 may be implemented by the server computing system 60 as part of a web service (e.g., a remote machine learning model host service such as an online interface for running machine learning model operations over a network on the server computing system 60). For example, the server computing system 60 may communicate with the computing device 50 via a local intranet or internet connection. For example, the computing device 50 may be a workstation or endpoint communicating with the server computing system 60, and the implementation of the machine learning model 65 may be managed by the server computing system 60 to remotely perform estimations (e.g., with respect to runtime or training operations) and return the output 50 (e.g., cast, streamed, etc.). The machine learning model 65 can work collaboratively or interactively with the machine learning model 55 on the computing device 50 to perform a variety of tasks.
[0263] The model development platform system(s) 70 may include one or more processors 71 and memory 72. The processor(s) 71 may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple processors connected in an operable manner. The memory 72 may include one or more non-temporary computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 72 may store data 73 and instructions 74 that can be executed by the processor(s) 71 to cause the model development platform system(s) 70 to perform operations. The operation may perform any one or more features described herein. The operation may perform exemplary methods and techniques described herein. The exemplary operation includes the functions described herein with respect to the model development platform 12. These functions and other functions may be performed by developer tools(s) 75.
[0264] The third-party system(s) 80 may include one or more processors 81 and memory 82. The processor(s) 81 may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple processors connected in an operable manner. The memory 82 may include one or more non-temporary computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 82 may store data 83 and instructions 84 that can be executed by the processor(s) 81 to cause the third-party system(s) 80 to perform operations. The operation may implement any one or more features described herein. The operation may implement exemplary methods and techniques described herein. Exemplary operations include the functionality described herein with respect to tools and other external resources that are invoked when training or performing estimations using machine learning models (or multiple models) such as 1701, 4, 16, 20, 55, 65, etc. (e.g., third-party resources (or multiple resources) such as 85, etc.).
[0265] Figure 23 shows one exemplary configuration of a computing system that can be used to implement this disclosure. Other computing system configurations can be used similarly. For example, in some embodiments, one or both of the computing system 50 or server computing system(s) 60 can implement all or part of the operation of the model development platform system 70. For example, the computing system 50 or server computing system(s) 60 can implement developer tools(s) 75 (or extensions thereof) to develop, update / train, or refine machine-learned models 1, 4, 16, 20, 55, 65, etc., using one or more techniques described herein with respect to the model alignment toolkit 17. In this way, for example, the computing system 50 or server computing system(s) 60 can develop, update / train, or refine machine-learned models based on local datasets (for example, for model personalization / customization permitted by user data preference selection).
[0266] Figure 24 is a block diagram of an exemplary computing device 98 implemented according to an exemplary embodiment of the present disclosure. The computing device 98 can be a user computing device or a server computing device (e.g., computing device 50, server computing system 60, etc.). The computing device 98 can implement a model host 31. For example, the computing device 98 can include multiple applications (e.g., applications 1 to N). Each application can include its own machine learning library and machine-trained models 1 to N. For example, each application can include machine-trained models. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. As illustrated in Figure 24, each application can communicate with multiple other components of the computing device, such as one or more sensors, a context manager, a device state component, or additional components. In some embodiments, each application can communicate with each device component using an API (e.g., a public API). In some embodiments, the API used by each application is specific to that application.
[0267] Figure 25 is a block diagram of an exemplary computing device 99 implemented according to an exemplary embodiment of the present disclosure. Computing device 99 may be the same as or different from computing device 98. Computing device 99 may be a user computing device or a server computing device (e.g., computing device 50, server computing system 60, etc.). Computing device 98 may implement a model host 31. For example, computing device 99 may include multiple applications (e.g., applications 1 to N). Each application may communicate with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some embodiments, each application may communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).
[0268] The central intelligence layer can include multiple machine learning models. For example, as shown in Figure 25, each machine learning model can be provided to each application and managed by the central intelligence layer. In other embodiments, two or more applications can share a single machine learning model. For example, in some embodiments, the central intelligence layer can provide a single model to all applications. In some embodiments, the central intelligence layer is contained within the operating system of the computing device 99 or otherwise implemented by that operating system.
[0269] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized repository of data from the computing device 99. As illustrated in Figure 25, the central device data layer can communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, or additional components. In some embodiments, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0270] The technologies described herein refer to servers, databases, software applications, and other computer-based systems, as well as actions performed and information transmitted to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of feasible configurations, combinations, and divisions of tasks and functions between and within their components. For example, the processes described herein can be implemented using a single device or component, or multiple devices or components working together. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0271] While the subject matter has been described in detail with respect to various specific exemplary embodiments, each example is provided for illustrative purposes only and does not limit the disclosure. Those skilled in the art, having attained the foregoing understanding, will readily be able to create modifications, variations, and equivalents to such embodiments. Therefore, the disclosure of the subject matter does not exclude the inclusion of such modifications, variations, or additions to the subject matter, as would be readily apparent to those skilled in the art. For example, features illustrated or described as part of one embodiment can be used in conjunction with another embodiment to create yet another embodiment. Thus, the disclosure is intended to cover such modifications, variations, and equivalents.
[0272] Aspects of this disclosure have been described in relation to their exemplary embodiments. All and all features of the following claims can be combined or rearranged in any possible way, including combinations of claims not expressly enumerated together, so that the dependency of the exemplary claims enumerated herein should not be read as limiting the scope of possible combinations of features disclosed herein. Accordingly, the scope of this disclosure is illustrative and not restrictive, and the disclosure of the subject matter does not exclude such modifications, variations, or additions to the subject matter, as would be readily apparent to those skilled in the art. Furthermore, terms are used herein with enumeration of exemplary elements joined by conjunctions such as “and,” “or,” and “but.” It should be understood that such conjunctions are provided for illustrative purposes only. For example, clauses and other sets of items joined by certain conjunctions such as “or” may refer to “and / or,” “at least one of the exemplary elements enumerated therein,” “any combination of,” etc. Terms such as “based on” should be understood as “at least partially based on.”
[0273] The term "can" should be understood not as referring to a capability that necessarily exists in all embodiments, but rather as referring to the potential of features in various embodiments. For example, the phrase "X can do Y" should be understood as indicating that in various embodiments X may be configured to do Y, and not as indicating that X is always able to perform Y in all examples. It should be understood that in various embodiments X may not be able to perform Y and may remain within the scope of this disclosure.
[0274] The term "may" should be understood as referring to the possibility of features in various embodiments, rather than prescribing a capability that necessarily exists in all embodiments. For example, the phrase "X may do Y" should be understood as indicating that in various embodiments X may be configured to do Y, and not as indicating that X is always able to perform Y in all examples. It should be understood that in various embodiments X may not be able to perform Y and may remain within the scope of this disclosure.< / platform> < / surroundings> < / background>
Claims
1. A computer implementation method, Receiving data indicating requests for multiple media assets with multiple media modalities generated based on existing media assets, This involves generating multiple media asset components, (i) Extracting one or more signals from the existing media assets, each comprising at least one of background image data, color palette data, text data, or logo image data, (ii) Generating multiple background image assets using a first generative machine learning model, A mask image including one or more bounding boxes is generated based on at least one of the background image data, the text data, or the logo image data. The mask image is repaired by generating pixels to fill one or more bounding boxes of the mask image using a machine learning model. The generation of multiple images, wherein each image includes a different aspect ratio, By doing so, (iii) Using a second generative machine learning model to generate multiple text assets, To obtain the text data associated with one or more bounding boxes, Inputting the aforementioned text data into a machine learning model, The aforementioned machine learning model generates multiple text assets as output, By doing so, (iv) Using a third generative machine learning model, generate unique profile data based on at least one of the color palette data or the logo image data, This involves generating the aforementioned multiple media asset components, A computer-aided method comprising: automatically sending a plurality of media assets, including one or more image assets, text assets, and unique profile data, to a content item generation pipeline in order to generate a plurality of candidate content items.
2. The mask image containing one or more bounding boxes is generated. Identifying one or more text objects, For each of the one or more text objects, Determining that the first text object among the one or more text objects is an overlay text object, Based on determining that the first text object is an overlay text object, a first bounding box is generated for the first text object, The computer implementation method according to claim 1, including the method described in claim 1.
3. The generation of the aforementioned multiple media asset components is Extracting the logo image, Selecting a dominant color from the image color palette, This involves upscaling the aforementioned logo image, Enlarging the aforementioned logo image, Sharpening the aforementioned logo image, The logo image is upscaled by the above, Generatively expanding the aforementioned logo image to generate one or more images, wherein each image includes a set of different aspect ratios, A computer implementation method according to claim 1 or 2, including the method described in claim 1 or 2.
4. The computer implementation method according to claim 3, wherein generatively expanding the logo image includes generating pixels to be mixed with existing pixels of the logo image to generate an image larger than the original logo image.
5. The computer implementation method according to any one of claims 1 to 4, wherein generating the plurality of background image assets includes adjusting at least one of the brightness, saturation, or contrast of each of the background image assets.
6. The computer implementation method according to any one of claims 1 to 5, wherein the unique profile data includes at least one of logo data, color palette data, font data, or image styling data.
7. Inputting multiple output assets into the content creation pipeline, A computer-aided method according to any one of claims 1 to 6, comprising generating a plurality of content items through the content creation pipeline, wherein each of the plurality of content items includes a unique combination of a content asset and an aspect ratio.
8. The generation of the aforementioned multiple text assets is Extracting signals that contain text data, Compiling the aforementioned text data, The compiled text data is input into the second generative machine learning model, The output from the second generative machine learning model described above is to obtain at least one of either a short headline or a long headline, Inputting at least one of the short or long headings into the second generative machine learning model, A computer implementation method according to any one of claims 1 to 7, comprising obtaining a description as output from the machine-trained model.
9. The computer implementation method according to any one of claims 1 to 8, wherein the one or more bounding boxes include the position and size of the bounding boxes.
10. The computer implementation method according to any one of claims 1 to 9, wherein the content item generation pipeline includes a prompt generation component and a content item generation component.
11. The prompt generation component generates input prompt data based on the background image, the text asset, and the unique profile data. The input prompt data is provided to the content item generation component, The computer implementation method according to claim 10, comprising generating one or more candidate content items based on the input prompt data using the content item generation component.
12. The computer implementation method according to claim 1, wherein the content item generation component includes a generative machine learning model.
13. Training the aforementioned generative machine learning model, Based on comparing the generated content item components with existing media assets, a training dataset is generated, Based on comparing the generated content item components with the existing media assets, one or more parameters of the generative machine learning model are automatically adjusted to reduce the differences between the generated content items and the existing media assets. This involves training the aforementioned generative machine learning model, The computer implementation method according to claim 12, including the method described in claim 12.
14. A computing system, One or more processors, One or more temporary or non-temporary computer-readable media storing instructions that can be executed to cause one or more processors to perform an operation, wherein the operation is Receiving data indicating requests for multiple media assets with multiple media modalities generated based on existing media assets, This involves generating multiple media asset components, (i) Extracting one or more signals comprising at least one of background image data, color palette data, text data, or logo image data, (ii) Generating multiple background image assets using a first generative machine learning model based on at least one of the background image data, the text data, the logo image data, or the color palette data, (iii) Generating multiple text assets using a second generative machine learning model based on at least one of the text data or background image data, (iv) Using a third generative machine learning model, generate unique profile data based on at least one of the following: color palette data, text data, or logo image data. By doing so, the plurality of media asset components are generated, To generate multiple candidate content items, the system automatically sends the multiple media assets, including one or more image assets, text assets, and unique profile data, to the content item generation pipeline. A computing system comprising the one or more temporary or non-temporary computer-readable media, including the above.
15. The computing system according to claim 14, wherein the operation includes receiving data indicating that the existing media asset has been uploaded.
16. The computing system according to claim 14 or 15, wherein the operation includes analyzing web resources associated with unique profile data to obtain the existing media assets.
17. The computing system according to any one of claims 14 to 16, wherein the existing media asset includes flattened image data.
18. The computing system according to claim 17, wherein the request is associated with a client account, and the client account is associated with an account profile that stores inputs to a machine learning-based media asset generation pipeline.
19. The computing system according to claim 18, wherein the account profile is retrieved from a database and the account profile is pre-generated prior to the request.
20. One or more temporary or non-temporary computer-readable media storing instructions that can be executed by one or more processors in order to perform an operation including the computer implementation method described in any one of claims 1 to 19.
Citation Information
Patent Citations
Image processing apparatus, image processing system, image processing method, and program
JP2020052530A
Image Replacement Repair
JP2023500203A
Method and system for reconciling unreconciled content based on multimodal metadata through data filtering and synchronization to generate a composite media asset - Patents.com
JP2024508363A
Content-adaptive guided tutorial generation
US10628185B1
System and methods for generating media assets
US20190096438A1