Generating digital components using customized machine learning models
By training basic machine learning models to generate user-specific or domain-specific models, the problem of generation does not meet the needs and data privacy protection is solved, and efficient digital component generation and resource utilization is achieved.
Patent Information
- Application Number
- CN202380070924.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to generate digital components that meet the specific needs of users or fields, and there are problems of data privacy protection and waste of computing resources.
By training basic machine learning models, generate user-specific or domain-specific machine learning models, use user-specific training data for refinement, and ensure that relevant data is not shared with other models, and use the generative model to generate digital components.
It realizes the generation of digital components that are more in line with user needs, protects data privacy, improves the efficiency of computing and network resources, and reduces the generation of unsatisfactory components.
Smart Images

Figure CN120359510A_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] This specification relates to using customized machine learning models for data processing and generating digital components.
[0002] Advances in machine learning have enabled artificial intelligence to be implemented in more applications. For example, generative models are a type of machine learning model designed to learn and mimic the underlying distribution of a given data set. Different from discriminative models that focus on classifying data into predefined categories, generative models are designed to generate new data similar to the original training data. Generative models are used in various applications such as image generation, text synthesis, and data augmentation. SUMMARY OF THE INVENTION
[0003] Generally, an innovative aspect of the subject matter described in this specification can be embodied in a method that includes the following actions: obtaining, by an artificial intelligence (AI) system, first model parameters associated with a first machine learning model at a first hierarchical layer; generating, by the AI system, a second machine learning model that includes the first model parameters, the second machine learning model being at a second hierarchical layer below the first hierarchical layer; and refining, by the AI system, the second machine learning model to obtain an updated second machine learning model that includes second model parameters, where the first model parameters and the second model parameters have different values for at least one model parameter, and the second model parameters cannot be shared with the first machine learning model.
[0004] These and other embodiments may each optionally include one or more of the following features.
[0005] The first model parameters may include at least one of a model type of a machine learning model, model weights of a machine learning model, a loss function of a machine learning model, or an optimization algorithm of a machine learning model.
[0006] The method may include providing, by the AI system, a first data sharing option and a second data sharing option, where the first data sharing option indicates that data associated with the second machine learning model cannot be shared with the first machine learning model, and where the second data sharing option indicates that at least a portion of the data associated with the second machine learning model can be shared with the first machine learning model.
[0007] The method may include receiving, by the AI system, an indication that the first data sharing option has been selected.
[0008] Data associated with the second machine learning model may include at least one of model parameters associated with the second machine learning model, queries input to the second machine learning model, outputs of the second machine learning model, or training data used to train the second machine learning model.
[0009] The method may include: receiving a query by an AI system; generating digital components by the AI system using a second machine learning model; and generating training data by the AI system based on the digital components, where the query, the digital components, and the training data are not shareable with the first machine learning model.
[0010] The method may include: obtaining, by the AI system, additional query data including data different from the query, where the additional query data restricts the digital components generated by the second machine learning model and is not shareable with the first machine learning model; generating, by the AI system, a prompt including the query and the additional query data; inputting, by the AI system, the prompt into the second machine learning model; and generating, by the second machine learning model, digital components.
[0011] The second machine learning model may be a supervised machine learning model. Generating, by the AI system, digital components using the second machine learning model may include: including the query as a feature of the training data; and including at least one of the digital components or an algorithm for generating the digital components as a label of the training data.
[0012] The second machine learning model may be trained using a reinforcement learning (RL) algorithm. Generating, by the AI system, digital components using the second machine learning model may include including at least one of the digital components, an algorithm for generating the digital components, or a reward for the digital components in the training data.
[0013] A third machine learning model may be at a third hierarchical layer below the second hierarchical layer, a fourth machine learning model may be at a fourth hierarchical layer below the third hierarchical layer, and data associated with the third machine learning model may not be shareable with the second machine learning model but may be shareable with the fourth machine learning model.
[0014] Training data for training the first machine learning model may be shareable with the second machine learning model.
[0015] Each of the first machine learning model and the second machine learning model may be a generative model.
[0016] The techniques described herein can be implemented to achieve the following advantages. In some cases, user data privacy can be protected by training a user-specific machine learning model that can be accessed exclusively by the user. In particular, the model parameters of a base machine learning model can be used to generate a user-specific machine learning model. The user-specific machine learning model can be refined using user-specific training data, and data associated with the user-specific machine learning model (such as model parameters) cannot be shared with other machine learning models (such as another machine learning model). By doing so, the techniques described herein can leverage an established base machine learning model to generate a user-specific machine learning model and enable the user-specific machine learning model to provide a secure, customized digital component generation service to the user.
[0017] In some cases, the techniques described herein enable the generation of sector-specific machine learning models or even more fine-grained sub-sector-specific machine learning models to generate sector-specific digital components. In some cases, the design styles from one sector (such as the fashion industry) may be very different from those of another sector (such as the automotive industry). Therefore, two different sectors may expect different machine learning models. On the other hand, the machine learning models for two different sectors may share a lot in common (for example, advertisements in both the fashion industry and the automotive industry need to meet certain regulations). The techniques described herein enable the training of a base machine learning model and the use of the model parameters of the base machine learning model to generate at least two separate sector-specific machine learning models (for example, one for the fashion industry and another for the automotive industry) for generating sector-specific digital components. By doing so, the at least two separate sector-specific machine learning models can leverage the model parameters of a common base machine learning model and generate sector-specific digital components for the user from the user's respective sector.
[0018] In this regard, the techniques described herein achieve significant advantages of machine learning-based digital component generation techniques. Additionally, the improved machine learning-based digital component generation techniques further bring computational and network resource efficiency. For example, the techniques described herein reduce the amount of unsatisfactory digital components, which can utilize a large amount of resources from generating and recommending associated results and evaluating and investigating user feedback on unsatisfactory digital components.
[0019] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1A block diagram of an example environment for implementing digital component generation using a customized machine learning model, according to an implementation of the present disclosure.
[0021] Figure 2 A block diagram showing the interaction between an AI system, a generative model, and a client device, according to an implementation of the present disclosure.
[0022] Figure 3 A block diagram of an example system for implementing a user-specific machine learning model, according to an implementation of the present disclosure.
[0023] Figure 4 A block diagram of an example system for implementing a domain-specific machine learning model, according to an implementation of the present disclosure.
[0024] Figure 5 A flowchart of an example process for implementing a user-specific machine learning model, according to an implementation of the present disclosure.
[0025] Figure 6 A flowchart of an example process for implementing a domain-specific machine learning model, according to an implementation of the present disclosure.
[0026] Figure 7 A block diagram of an example computer system that can be used to perform the described operations, according to an implementation of the present disclosure.
[0027] Like reference numerals and names in the various figures indicate like elements. Detailed Description
[0028] This specification describes techniques for using customized (e.g., user-specific or domain-specific) machine learning models to generate digital components, and is presented to enable any person skilled in the art to make and use the disclosed subject matter in the context of one or more specific implementations. Various modifications, changes, and permutations can be made to the disclosed implementations, and such modifications, changes, and permutations will be apparent to those of ordinary skill in the art, and the defined general principles can be applied to other implementations and applications without departing from the scope of the present disclosure. In some instances, one or more technical details that are not necessary for obtaining an understanding of the described subject matter and are within the skills of those of ordinary skill in the art may be omitted so as not to obscure one or more of the described implementations. The present disclosure is not intended to be limited to the described or illustrated implementations, but is accorded the widest scope consistent with the described principles and features.
[0029] Artificial intelligence (AI) is a branch of computer science focused on creating agents that can learn and act autonomously (e.g., without human intervention). AI can utilize machine learning, which focuses on developing algorithms that can learn from data; natural language processing, which focuses on understanding and generating human language; and / or computer vision, which is a field focused on understanding and interpreting images and videos.
[0030] In the prior art, machine learning models can be trained to generate digital components for various users, such as advertisements. However, some users, such as VIP advertisers, may be reticent about sharing their data (such as design concepts or advertising performance data) with others. This may lead to avoiding the use of machine learning models in order to protect their data from unauthorized access.
[0031] Furthermore, current technologies typically train machine learning models to create digital components for different domains (such as vertical markets or geographic regions). However, design styles within one domain (such as the fashion industry) may be significantly different from those in another domain (such as the automotive industry). Although machine learning models can be domain-neutral and trained without specific weights for training data assigned to any domain, this may result in digital components that do not fully meet the unique needs or requirements of users from a specific domain.
[0032] In some implementations, the techniques described throughout this specification enable the generation of machine learning models specifically trained for a user or a domain, and using the machine learning model to generate user-specific or domain-specific digital components. Initially, a base machine learning can be trained. The model parameters of the base machine learning model can be used to generate a user-specific or domain-specific machine learning model, which can be continuously refined using user-specific or domain-specific training data. The data associated with the user-specific or domain-specific machine learning model may not be shareable with others to protect the data from unauthorized access. Moreover, due to the specific training, the user-specific or domain-specific machine learning model can generate digital components that are more satisfactory to users with specific needs or users from a specific domain than the digital components generated by the base machine learning model.
[0033] As used throughout this document, the phrase "digital component" refers to a discrete unit of digital content or digital information (e.g., a video clip, an audio clip, a multimedia clip, game content, an image, text, a bullet point, AI output, language model output, or another content unit). Digital components can be electronically stored in a physical memory device as a single file or as a collection of files, and digital components can take the form of a video file, an audio file, a multimedia file, an image file, or a text file, and include advertising information such that an advertisement is a type of digital component.
[0034] Figure 1 is a block diagram of an example environment 100 for implementing digital component generation using a customized machine learning model according to an implementation of the present disclosure. The example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. The network 102 connects an electronic document server 104, a client device 106, a digital component server 108, and a service device 110. The example environment 100 can include many different electronic document servers 104, client devices 106, and digital component servers 108.
[0035] The client device 106 is an electronic device capable of requesting and receiving online resources via the network 102. Example client devices 106 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality devices, virtual reality devices, and other devices that can send and receive data via the network 102. The client device 106 typically includes a user application (such as a web browser) to facilitate sending and receiving data via the network 102, but native applications executed by the client device 106 can also facilitate sending and receiving data via the network 102.
[0036] A gaming device is a device that enables a user to participate in a gaming application. For example, in the device, the user can control one or more characters, avatars, or other rendered content presented in the gaming application. A gaming device typically includes a computer processor, a memory device, and a controller interface (physical or visually rendered) that enables the user to control the content rendered by the gaming application. A gaming device can store and execute a gaming application locally, or execute a gaming application that is at least partially stored and / or serviced by a cloud server (e.g., an online gaming application). Similarly, a gaming device can interface with a gaming server that executes the gaming application and "streams" the gaming application to the gaming device. A gaming device can be a tablet device, a mobile telecommunications device, a computer, or another device that performs other functions in addition to executing a gaming application.
[0037] The digital assistant device includes a device having a microphone and a speaker. The digital assistant device is generally capable of receiving input by voice, and responding to content using audible feedback, and may present other audible information. In some cases, the digital assistant device also includes a visual display or communicates with a visual display (e.g., via a wireless or wired connection). When there is a visual display, feedback or other information may also be provided visually. In some cases, the digital assistant device may also control other devices, such as lights, locks, cameras, climate control devices, alarm systems, and other devices registered with the digital assistant device.
[0038] As shown, client device 106 is presenting electronic document 150. An electronic document is a collection of data for presenting content at client device 106. Examples of electronic documents include web pages, word processing documents, Portable Document Format (PDF) documents, images, videos, search result pages, and feeds. Native applications (e.g., “apps” and / or game applications) such as those installed on mobile, tablet, or desktop computing devices are also examples of electronic documents. The electronic document may be provided to client device 106 by an electronic document server 104 (“Electronic Doc Servers”).
[0039] For example, electronic document server 104 may include a server hosting a publisher website. In this example, client device 106 may initiate a request for a given publisher web page, and the electronic server 104 hosting the given publisher web page may respond to the request by sending machine-executable instructions that initiate the presentation of the given web page at client device 106.
[0040] In another example, electronic document server 104 may include an app server from which client device 106 can download an app. In this example, client device 106 may download the files required to install the app at client device 106 and then execute the downloaded app locally (i.e., on the client device). Alternatively or additionally, client device 106 may initiate a request to execute the app, which is sent to a cloud server. In response to receiving the request, the cloud server may execute the application and stream the user interface of the application to client device 106, such that client device 106 does not have to execute the app itself. Instead, client device 106 may present the user interface generated by the cloud server executing the app and communicate any user interactions with the user interface back to the cloud server for processing.
[0041] An electronic document can include a variety of content. For example, electronic document 150 can include native content 152 that is within the electronic document 150 itself and / or does not change over time. The electronic document can also include dynamic content that can change over time or based on each request. For example, the publisher of a given electronic document (e.g., electronic document 150) can maintain a data source for populating portions of the electronic document. In this example, a given electronic document can include a script, such as script 154, that causes a client device 106 (or cloud server) to request content (e.g., digital components) from the data source when the given electronic document is processed (e.g., rendered or executed) by the client device 106. The client device 106 (or cloud server) integrates the content (e.g., digital components) obtained from the data source into the given electronic document to create a composite electronic document that includes the content obtained from the data source.
[0042] In some cases, a given electronic document (e.g., electronic document 150) can include a digital component script (e.g., script 154) that references a service device 110 or a specific service provided by the service device 110. In these cases, the digital component script is executed by the client device 106 when the given electronic document is processed by the client device 106. Execution of the digital component script configures the client device 106 to generate a request for a digital component 112 (referred to as a "component request") that is sent over the network 102 to the service device 110. For example, the digital component script can enable the client device 106 to generate a packetized data request that includes a header and payload data. The component request 112 can include event data specifying characteristics such as the name (or network location) of the server from which the digital component is being requested, the name (or network location) of the requesting device (e.g., client device 106), and / or information that the service device 110 can use to select one or more digital components or other content to provide in response to the request. The component request 112 is sent by the client device 106 over the network 102 (e.g., a telecommunications network) to the server of the service device 110.
[0043] Component request 112 may include event data specifying other event characteristics, such as the electronic document being requested and characteristics of the location where the digital component may be presented. For example, event data specifying a reference (e.g., a Uniform Resource Locator (URL)) to the electronic document (e.g., a web page) in which the digital component will be presented, the available location within the electronic document for presenting the digital component, the size of the available location, and / or the media type eligible to be presented at the location may be provided to service device 110. Similarly, event data specifying keywords associated with the electronic document ("document keywords") or entities (e.g., people, places, or things) referenced by the electronic document may also be included in component request 112 (e.g., as payload data) and provided to service device 110 to facilitate identification of digital components eligible to be presented with the electronic document. The event data may also include a search query submitted from client device 106 to obtain a search results page.
[0044] Component request 112 may also include event data related to other information, such as information provided by the user of the client device, geographic information indicating the state or region from which the component request is submitted, or other information providing the context of the environment in which the digital component will be displayed (e.g., the time of day of the component request, the date of the week of the component request, the type of device on which the digital component will be displayed, such as a mobile device or a tablet device). Component request 112 may be sent, for example, over a packetized network, and component request 112 itself may be formatted as packetized data having a header and payload data. The header may specify the destination of the packet, and the payload data may include any of the information discussed above.
[0045] Service device 110 selects digital components (e.g., third-party content such as video files, audio files, images, text, game content, augmented reality content, and combinations thereof, all of which may take the form of advertising content or non-advertising content) to be presented with a given electronic document (e.g., at the location specified by script 154) in response to receiving component request 112 and / or using the information included in component request 112.
[0046] In some implementations, the digital components are selected in less than one second to avoid errors that may result from a delayed selection of the digital components. For example, a delay in providing the digital components in response to component request 112 may cause a page load error at client device 106 or result in parts of the electronic document remaining unfilled even after other parts of the electronic document are presented at client device 106.
[0047] Moreover, as the latency increases when providing digital components to the client device 106, it is more likely that the electronic document will no longer be presented at the client device 106 when the digital components are delivered to the client device 106, thereby negatively affecting the user's experience of the electronic document. Additionally, the latency in providing digital components may cause the delivery of the digital components to fail, for example, if the electronic document is no longer presented at the client device 106 when the digital components are provided.
[0048] In some implementations, the service device 110 is implemented in a distributed computing system that includes, for example, a server and a set 114 of multiple computing devices, where the server and the set of multiple computing devices are interconnected and identify and distribute digital components in response to a request 112. The set 114 of multiple computing devices operates together to identify a set of digital components eligible for presentation in an electronic document from a corpus of millions of available digital components (DC 1-x ). The millions of available digital components may be indexed, for example, in a digital component database 116. Each digital component index entry may reference the corresponding digital component and / or include distribution parameters (DP1 to DP x ) that contribute to (e.g., trigger, condition, or limit) the distribution / transmission of the corresponding digital component. For example, the distribution parameters may contribute to (e.g., trigger) the transmission of a digital component by requiring that the component request include at least one criterion that matches (e.g., exactly or at some pre-specified level of similarity) one of the distribution parameters of the digital component.
[0049] In some implementations, the distribution parameters for a particular digital component may include distribution keywords that must match (e.g., match the electronic document, document keywords, or terms specified in the component request 112) for the digital component to be eligible for presentation. Additionally or alternatively, the distribution parameters may include embeddings that can use various different data dimensions, such as website details and / or consumption details (e.g., page viewport, user scroll speed, or other information regarding data consumption). The distribution parameters may also require that the component request 112 include information specifying a particular geographical region (e.g., country or state) and / or information specifying that the component request 112 originates from a particular type of client device (e.g., a mobile device or a tablet device) for the digital component to be eligible for presentation. The distribution parameters may also specify an eligibility value (e.g., a ranking score, or some other specified value) that is used to evaluate the eligibility of the digital component for distribution / transmission (e.g., as compared to other available digital components).
[0050] The identification of eligible digital components can be split into multiple tasks 117a through 117c, which are then assigned among the computing devices within a set 114 of multiple computing devices. For example, different computing devices in set 114 can each analyze different portions of a digital component database 116 to identify various digital components having distribution parameters that match the information included in a component request 112. In some implementations, each given computing device in set 114 can analyze different data dimensions (or sets of dimensions) and pass (e.g., send) the results (Res 1 through Res 3) 118a through 118c of the analysis back to a service device 110. For example, the results 118a through 118c provided by each of the computing devices in set 114 can identify a subset of digital components eligible for distribution in response to a component request and / or a subset of digital components having certain distribution parameters. The identification of a subset of digital components can include, for example, comparing event data with distribution parameters and identifying a subset of digital components having distribution parameters that match at least some characteristics of the event data.
[0051] The service device 110 aggregates the results 118a through 118c received from the set of multiple computing devices 114 and uses the information associated with the aggregated results to select one or more digital components to be provided in response to the request 112. For example, the service device 110 can select a winning set of digital components (one or more digital components) based on the results of one or more content evaluation processes, as described below. Further, the service device 110 can generate and send via a network 102 reply data 120 (e.g., digital data representing a reply) that enables the client device 106 to integrate the winning set of digital components into a given electronic document such that the winning set of digital components (e.g., winning third-party content) and the content of the electronic document are presented together at a display of the client device 106.
[0052] In some implementations, the client device 106 executes instructions included in the reply data 120 that configure the client device 106 and enable it to obtain the winning set of digital components from one or more digital component servers 108. For example, the instructions in the reply data 120 can include a network location (e.g., a URL) and a script that causes the client device 106 to send a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component from the digital component server 108. In response to the request, the digital component server 108 will identify the given winning digital component specified in the server request 121 (e.g., within a database storing multiple digital components) and send digital component data 122 to the client device 106 that presents the given winning digital component in the electronic document at the client device 106.
[0053] When client device 106 receives digital component data 122, the client device will render the digital component (e.g., third-party content) and present the digital component at the location specified by or assigned to script 154. For example, script 154 can create a walled garden environment (such as a frame) that is presented within the native content 152 of electronic document 150, e.g., beside the native content. In some implementations, the digital component overlays (or is adjacent to) a portion of the native content 152 of electronic document 150, and service device 110 can specify the presentation location within electronic document 150 in response 120. For example, when native content 152 includes video content, service device 110 can specify a location or object within the scene depicted in the video content above which the digital component will be presented.
[0054] Service device 110 may also include AI system 160, which is configured to autonomously generate digital components before request 112 (e.g., offline) and / or in response to request 112 (e.g., online or in real time). As described in more detail throughout this specification, AI system 160 can collect online content about a particular entity (e.g., a digital component provider or another entity) and use one or more generative models 170 to generate digital components based on the collected online content.
[0055] Generative models are designed to generate new data similar to a given training dataset and operate by learning the underlying patterns, structures, and relationships present in the training dataset, enabling them to create new samples with similar characteristics. The primary goal of generative models is to capture the inherent complexity of the data distribution, allowing them to produce outputs that exhibit the same diversity and variability found in the original dataset.
[0056] One of the fundamental concepts of generative models is to generate data from random noise or latent variables. Generative models create a mapping between the latent space and the data space, thereby allowing the generation of entirely new instances with meaningful features. Generative models can be broadly classified into two main types: likelihood-based and adversarial.
[0057] Likelihood-based generative models (such as variational autoencoders (VAEs) and autoregressive models) focus on learning the probability distribution of the data. For example, VAEs employ an encoder-decoder architecture to map data points into the latent space and then decode them back into the data space. This process encourages the model to learn a more structured and continuous representation of the data distribution.
[0058] Adversarial-based generative models, most notably generative adversarial networks (GANs), utilize different methods. A GAN consists of two neural networks: a generator and a discriminator. The generator aims to produce data that is indistinguishable from real data, while the discriminator attempts to distinguish between real data and the generated data. This adversarial process causes the generator to improve over time and produce increasingly convincing outputs.
[0059] Figure 2 FIG. 200 is a block diagram showing the interaction between an AI system, a generative model, and a client device according to an implementation of the present disclosure. In some cases, the generative model 202 and the client device 204 may be the same as or similar to Figure 1 the generative model 170 and the client device 106, respectively. The generative model 202 may be, for example, a text-to-text generative model, a text-to-image generative model, a text-to-video generative model, an image-to-image generative model, or any other type of generative model. Although Figure 2 a single generative model 202 is depicted in FIG., the generative model 202 may be a set of different generative models, and the set of different generative models may be invoked for different tasks, and the different generative models are specifically trained for the different tasks. For example, one generative model within the set of generative models may be specifically trained to perform a content summarization task, while another model may be specifically trained to generate digital components, for example, using the output of the specifically trained generative model. In addition, the set of models may include a generalized generative model that is larger in size and capable of generating a large and diverse dataset, but this generalized model may have a higher latency than the specialized models, which may make the generalized model less desirable for use in real-time operations depending on the time latency constraints required for generating the content.
[0060] The AI system 160 includes a data collection device 206, a prompt device 208, a digital component service device 210, a training data generation device 212, and a model refinement device 214. The following description refers to these different devices as independently implemented and each configured to perform a set of operations, but any of these devices may be combined to perform the operations discussed below.
[0061] The AI system 160 communicates with a memory structure 232. The memory structure 232 may include one or more databases. As shown, the memory structure 232 includes a collected data database 216, a digital component database 218, and a training data database 220. Each of these databases 216, 218, and 220 may be implemented in the same hardware memory device, separate hardware memory devices, and / or in a distributed cloud computing environment.
[0062] At a high level, client device 204 sends query 246 to AI system 160. In some examples, a user may use a front-end interface of AI system 160 (e.g., a website or an application of a computing device) to submit a query. In some cases, query 246 may be, for example, a request for AI system 160 to generate a digital component (e.g., an advertisement). For example, a user may enter a prompt to request AI system 160 to generate an advertisement.
[0063] In some cases, a user may upload to AI system 160 one or more raw digital components (e.g., images, text, and videos) associated with the query (whether or not as part of the query), and the raw digital components may be used to create a digital component. For example, the raw digital component may be an image of a product, and the image of the product may be included in one or more advertisements generated by AI system 160.
[0064] In some embodiments, a user may submit additional query data to AI system 160, where the additional query data may include data not in the query and may limit the digital components generated by AI system 160. For example, the additional query data may include, but is not limited to, the geographic location targeted by the advertisement, the language of the advertisement, and / or the vertical industry targeted by the advertisement. For example, an advertiser may indicate that the advertisement is targeted at the North American market, should be in English, and / or is targeted at the fashion apparel vertical industry. In some examples, the user provides the additional query data in the same prompt as the request to generate the digital component. In other examples, the additional query data is entered separately from the prompt. For example, AI system 160 may generate one or more follow-up questions in response to a user's prompt, where the one or more follow-up questions are for soliciting input from the user on the additional query data. For example, the follow-up questions may be "which geographic location(s) is targeted by the advertisement", "which language should the advertisement be in", and / or "which vertical market(s) is targeted by the advertisement?"
[0065] In some examples, the AI system 160 can use the data collection device 206 to collect additional query data that is not directly input by the user. The data collection device 206 is implemented using at least one computing device (e.g., one or more processors) and can include one or more machine learning models. In some cases, the data collection device 206 can obtain an identification of an entity associated with the query. The identification can include at least one identifier, such as a company or business name, URL, phone number, employer ID number, or other means of identifying the entity. The data collection device 206 can use, for example, the account of the user who submitted the query or obtain at least one identifier from a partner system. The data collection device 206 can automatically identify data sources that include information about the entity based on the identification of the entity. These data sources can be, but are not limited to, web pages (e.g., the landing page of the entity), review compilation pages (e.g., google.com, yelp.com, and crunchbase.com), federal and / or state registries (e.g., Delaware entity search tool), private databases, news articles, or other suitable sources. In some implementations, a data crawler application automatically queries multiple databases, performs searches, and extracts information from the results in response to a process being triggered. The information obtained from these data sources can be large amounts of text data, a combination of text and images, metadata, or other suitable data and / or media.
[0066] In some examples, the data collection device 206 can perform semantic analysis on the collected information of at least one data source. In some implementations, semantic analysis is used to analyze a single data source. In some implementations, all of the collected information is analyzed. The semantic analysis can be performed by one or more machine learning algorithms, the overall goal of which is to generate one or more entity attributes associated with the entity. In some cases, the data collection device 206 can perform semantic analysis using a concatenation operation or an array of neural networks of machine learning algorithms that can include parallel operations or operate independently of each other. In some implementations, traditional data analysis can be performed in addition to or separately from the machine learning process. Similar to the additional query data, one or more entity attributes can include, for example, the geographical location targeted by the advertisement, the preferred language of the advertisement, and / or the vertical market targeted by the advertisement. In some examples, the data collection device 206 can include one or more entity attributes in the additional query data.
[0067] The data collection device 206 may store the collected data in the collected data database 216. For example, the data collection device 206 may index the collected data to the query for which the data was collected and / or the entity characterized by the collected data, such that the collected data may be retrieved from the collected data database 216 for additional operations performed by the data collection device 206 and / or any operations performed by the AI system 160.
[0068] The AI system 160 may use the prompting device 208 to generate an input prompt 242 using the query 246 and / or additional query data. The prompting device 208 may be implemented using at least one computing device (e.g., a device including one or more processors) and may include one or more language models. In some cases, the input prompt 242 may include the query 246 and a set of constraints generated based on, for example, additional query data. For example, the prompting device 208 may insert into the input prompt 242 one or more of the entity attributes corresponding to the entity identified by the data collection device 206. In some implementations, one or more of the entity attributes inserted into the prompt operate as context constraints that limit the content created by the generation model 202 in response to the input prompt 242. For example, the entity attributes may limit the content created by the generation model to the topic specified by the entity attributes included as context constraints in the prompt.
[0069] The AI system 160 may send the input prompt 242 to the generation model 202. The generation model 202 may then generate a plurality of candidate digital components based on the input prompt 242 and send the candidate digital components as the model output 244 to the AI system 160. In some cases, the AI system 160 may receive a plurality of original digital components (e.g., original images) associated with the query. The generation model 202 may use the original digital components to generate a plurality of candidate digital components (e.g., candidate advertisements), where each of the plurality of candidate digital components includes at least one of the plurality of original digital components.
[0070] The AI system 160 may store the generated candidate digital components in the digital component database 218. For example, the AI system 160 may index the generated candidate digital components to the query for which the candidate digital components were generated and / or the entity associated with the candidate digital components, such that the candidate digital components may be retrieved from the digital component database 218 for additional operations performed by the AI system 160.
[0071] For example, assume that query 246 is "Generate an advertisement for sunglasses" and the user uploads an image of sunglasses. Also assume that the additional query data indicates that the entity is for the fashion clothing vertical market in Japan. Input prompt 242 can take the following form:
[0072] Generate good_output: an advertisement where the query is "Generate an advertisement for sunglasses". good_output should be for the fashion clothing vertical market in Japan.
[0073] The generation model 202 can generate multiple candidate advertisements that include images of sunglasses and have different backgrounds. For example, the advertisement can include a scene of Mount Fuji in the background, the advertisement can include a snow scene in the background, the advertisement can include a backyard scene in the background, and the advertisement can include a scene of the Eiffel Tower in the background.
[0074] The AI system 160 can use the digital component service device 210 to serve one or more of the candidate digital components. The digital component service device 210 can be implemented using at least one computing device (e.g., a device including one or more processors) and can include one or more machine learning models. Assuming the digital component is an advertisement, in some cases, the digital component service device 210 can perform advertisement rendering, which includes rendering and formatting the advertisement to match the layout of the publisher's website or app. The digital component service device 210 can generate the necessary hypertext markup language (HTML), image, or video components to display the advertisement. In some examples, the digital component service device 210 can perform advertisement delivery, i.e., sending the rendered advertisement to the publisher's website or app, where the advertisement is displayed to the user in the specified advertisement space. In some examples, the digital component service device 210 can serve multiple candidate digital components generated in response to query 246 and collect performance data for the candidate digital components. In some embodiments, the performance data can indicate the acceptance level of the candidate digital components and can be used to evaluate and rank the candidate digital components. The performance data can be based on, for example, user interactions with the candidate digital components. For example, the user can interact with the advertisement by clicking on the advertisement, watching the video, purchasing the product promoted by the advertisement, or taking other actions. Examples of performance data include, but are not limited to, click-through rate (CTR), conversion rate (CVR), cost per day (CPD), and other user actions.
[0075] In some implementations, the digital component service device 210 may operate in an exploration mode or an exploitation mode. For example, when performance data needs to be collected for the evaluation of candidate digital components, it may operate in the exploration mode. When the digital component service device 210 operates in the exploration mode, the digital component service device 210 may randomly select one of the candidate digital components in each service of the candidate digital components (e.g., in each delivery of an advertisement) to deliver to the user. After multiple deliveries, each candidate digital component has the opportunity to be delivered to the user (e.g., the audience of the advertisement), and the performance data of each candidate digital component has the opportunity to be monitored and recorded. On the other hand, for example, when the generation model has been trained to a certain extent (e.g., when a predetermined amount of performance data has been received, when a predetermined number of training iterations have been performed, or other suitable conditions), it may operate in the exploitation mode. When the digital component service device 210 operates in the exploitation mode, the digital component service device 210 does not test multiple candidate digital components in response to a query. Instead, the digital components generated by the digital component service device 210 may be directly delivered to the user and / or the querier who requests the generation of the digital components.
[0076] In some cases, by serving candidate digital components, the AI system 160 may determine the acceptance level of each of the candidate digital components and identify the candidate digital component with the highest acceptance level (e.g., the highest CTR, the highest CVR, and / or the highest CPD) among the candidate digital components. In some embodiments, the AI system 160 may perform the identification of the candidate digital components when one or more predetermined conditions occur, such as a predetermined time period has elapsed after delivering the digital component, the acceptance level of the candidate digital component has met (e.g., met or exceeded) a predetermined threshold, or other suitable conditions.
[0077] The AI system 160 can determine the acceptance level in various ways. In some cases, the acceptance level is determined based on a metric of performance data. For example, the candidate digital component with the highest CTR, highest CVR, or highest CPD has the highest acceptance level. In some cases, the performance data includes at least two different metrics, and the acceptance level is determined based on a combination of at least two different metrics. In one example, the acceptance level can be determined based on the weighted sum of at least two different metrics. In another example, at least two different metrics can be ranked from the most important metric to the least important metric. First, the candidate digital components can be ranked from highest to lowest based on the most important metric. The candidate digital component ranked highest using this metric has the highest acceptance level. For candidate digital components with at least one metric equal, the candidate digital components can be ranked (i.e., broken ties) based on the lower-ranked metric to determine their ranking.
[0078] In some examples, regardless of whether performance data is considered, the AI system 160 can identify candidate digital components based on one or more other metrics. In some cases, the candidate digital component with the best performance data may not necessarily be the desired output. An example is a clickbait ad, which is an online ad designed to attract viewers to click. Clickbait ads typically use engaging or sensational headlines, images, or phrases to attract users' attention and encourage them to click on the ad to learn more. Thus, clickbait ads may perform well in a metric of performance data such as CTR. If CTR is the only metric used to determine the acceptance level, the clickbait ad may have the highest acceptance level. However, clickbait ads may not be the desired output because, for example, they may perform poorly in another metric such as CVR.
[0079] To prevent this outcome, other metrics can be implemented to identify candidate digital components. For example, the candidate digital components can be ranked from best to worst based on performance data. Starting from the beginning of the ranked candidate digital components, each candidate digital component can be evaluated to determine whether the attributes of the candidate digital component meet a predetermined condition (e.g., the candidate digital component does not include any clickbait ad information). The first candidate digital component whose attributes meet the predetermined condition can be identified as the candidate digital component. In some cases, the candidate digital components can be identified based on the performance data associated with multiple candidate digital components (e.g., using similar operations described above regarding identifying candidate digital components based on the acceptance level).
[0080] In some embodiments, the AI system uses a training data generation device 212 to generate training data. The training data generation device 212 can be implemented using at least one computing device (e.g., a device including one or more processors), and can include one or more machine learning models.
[0081] In some embodiments, the generative model 202 is an unsupervised machine learning model trained using a reinforcement learning algorithm. Reinforcement learning (also known as RL) is a machine learning method for solving problems by maximizing rewards or achieving specific goals through the interaction between an agent and an environment, which is modeled as a Markov decision process (MDP). RL is an unsupervised learning method that relies on sequential feedback (e.g., rewards) from the environment. During the learning process, the agent observes the state of the environment, selects actions based on a policy, and receives feedback in the form of rewards or scores.
[0082] Through trial and error, the agent iteratively interacts with the environment, aiming to obtain maximum rewards or reach a specific goal. The reward signal from the environment is used to evaluate the quality of the agent's actions, rather than guiding the agent on how to perform correct actions. Since the environment provides limited feedback, the agent learns through experience, acquires knowledge during the interaction, and enhances the action selection policy to adapt to the environment.
[0083] More specifically, the learning process can involve the agent repeatedly observing the state of the environment, making decisions based on behavior, and receiving feedback. The goal of this learning can be to achieve an ideal state value function or policy. In some cases, the state value function can represent the expected cumulative reward that can be obtained by following the policy.
[0084] In one example, the state value function can be defined as:
[0085] V π (s) = E π [R t |s t = s]
[0086] In this equation, R t represents the long-term cumulative reward obtained by performing actions based on the policy π. The state value function represents the expectation of the cumulative reward brought by using the policy π starting from the state s.
[0087] As an example, assume that the digital component is an image, and the generative model 202 is trained to generate images based on queries and / or additional query data. The state of the environment can include the following elements:
[0088] - Queries and / or additional query data.
[0089] -Feedback on previously generated images: The feedback can be based on, for example, performance data associated with multiple candidate images.
[0090] The state evolves as the generative model 202 iteratively generates images, receives feedback, and updates its policy.
[0091] The actions of the agent can be the actions taken by the generative model 202 in response to its current state. In this example, the actions of the agent can be to generate an image based on its current policy, queries, and / or additional query data, as well as feedback on previously generated images. The agent aims to learn a policy that causes the generation of images that receive positive feedback, and thus maximizes the received rewards while minimizing the penalties. The rewards and / or penalties can be determined based on a reward function.
[0092] The goal of the generative model 202 is to learn from these rewards and penalties to iteratively improve its digital component generation ability. Over time, the generative model 202 should generate digital components that are more likely to receive positive feedback, resulting in better digital component generation performance. In this case, the reward function acts as a "reinforcement signal" that guides the learning process of the generative model 202.
[0093] In some implementations, when the machine learning model is an unsupervised machine learning model trained using an RL algorithm, the AI system 160 can include at least one of the following in the training data: identified candidate digital components (e.g., pixels of the generated image), the algorithm used to generate the identified candidate digital components, or the rewards of the identified candidate digital components. In some cases, the training data can include other candidate digital components and / or their corresponding data (e.g., other candidate digital components, the algorithms used to generate other candidate digital components, and / or the rewards of other candidate digital components).
[0094] The algorithm used to generate the identified candidate digital components can include, for example, one or more steps associated with generating a background image for the original image. In some cases, the algorithm used to generate the candidate digital components can occupy less memory space than the candidate digital components themselves. Therefore, in some cases, including the algorithm used to generate the candidate digital components in the training data can save storage space compared to including the entire candidate digital component in the training data.
[0095] In some implementations, a reward function can be used to generate the rewards of the candidate digital components. To output the rewards of the candidate digital components, the input to the reward function can include, for example, at least one of the acceptance level or performance data of the candidate digital components.
[0096] In some examples, the generation model 202 is a supervised machine learning model. The input to a supervised machine learning model can include one or more features, such as an input prompt (e.g., input prompt 242), a query (e.g., query 246), and / or additional query data. The output of a supervised machine learning model can be, for example, a digital component (e.g., an image). A supervised machine learning model can be trained using a set of training data and a corresponding set of labels, where the training data can include multiple sets of data related to multiple queries and the generated digital components for the multiple queries. For example, a piece of training data can include an input prompt, a query, and / or additional query data as features of the sample. Given the features of the sample, the label of the piece of training data can be, for example, a digital component with a high acceptance level. A machine learning model can be trained by optimizing a loss function based on the difference between the output of the model and the corresponding label during training.
[0097] The training data generation device 212 can store the generated training data in the training data database 220. For example, the training data database 220 can index the generated training data to the query for which the training data is generated and / or the entity associated with the generated training data, such that the generated training data can be retrieved from the training data database 220 for additional operations performed by the training data generation device 212 and / or the AI system 160.
[0098] In some cases, the AI system 160 can use the model refinement device 214 to refine the generation model 202 using the training data. The model refinement device 214 can be implemented using at least one computing device (e.g., a device including one or more processors) and can include one or more machine learning models. In some cases, the model refinement device 214 can refine the generation model 202 immediately after a specific event occurs. For example, when the accuracy of the generation model 202 meets (meets or is below) a predetermined threshold, the generation model 202 can be retrained. In some cases, the generation model 202 can be retrained periodically (e.g., every seven days or thirty days) and / or when a certain amount of training data has been generated.
[0099] In some implementations, after a training and / or refinement period, the generation model 202 can meet one or more predetermined conditions. One or more predetermined conditions can include, for example, that the accuracy of the generation model 202 meets (meets or exceeds) a predetermined threshold (e.g., the CTR of the images generated by the generation model 202 meets a predetermined threshold). When the generation model 202 meets one or more predetermined conditions, the AI system 160 can enter an exploitation mode, in which the AI system 160 can return the output digital component 248 to the client device 204 in response to a query from the client device 204.
[0100] In some examples, the generation model 202 can be a machine learning model that is specifically trained for a user and / or a domain (e.g., the user-specific machine learning model 308 described with respect to Figure 3 the domain-specific machine learning model 410 described with respect to Figure 4 and / or the subdomain-specific machine learning model 412 described with respect to Figure 4 ). In some cases, the AI system 160 can interface with more than one generation model, and the AI system 160 can determine which generation model to invoke based on the query 246. For example, assume there is a base machine learning model and a user-specific machine learning model. When the AI system 160 receives the query 246, the AI system 160 can determine whether the user associated with the query 246 has subscribed to the user-specific machine learning model. If the user associated with the query 246 has subscribed to the user-specific machine learning model, the AI system 160 can invoke the user-specific machine learning model. Otherwise, the AI system 160 can invoke the base machine learning model. More details are described with respect to Figures 3 to 6 ).
[0101] Figure 3 is a block diagram of an example system 300 for implementing a user-specific machine learning model according to an implementation of the present disclosure. At a high level, the example system 300 relates to a solution for generating a machine learning model specifically trained for a user and using the machine learning model to generate user-specific digital components. Initially, a base machine learning can be trained. The model parameters of the base machine learning model can be used to generate a user-specific machine learning model, and the machine learning model can be continuously refined using user-specific training data. Due to the user-specific training, the user-specific machine learning model can generate digital components that are more satisfactory to the user than the digital components generated by the base machine learning model. The example system 300 can be extended deeper using similar operations. More details are described below.
[0102] The example system 300 includes two hierarchical layers of machine learning models, namely a first hierarchical layer 302 and a second hierarchical layer 304 below the first hierarchical layer 302. The first hierarchical layer 302 includes a base machine learning model 306. The second hierarchical layer 304 can include one or more user-specific machine learning models 308 (e.g., a machine learning model for user A to a machine learning model for user N, as shown). In some cases, each of the base machine learning model 306 and the user-specific machine learning models 308 is a generation model. A user 310 (e.g., any one of users A to N, as shown) can invoke the user-specific machine learning model 308 at the second hierarchical layer 304 to generate digital components.
[0103] Initially, a general training data that can be shared with the base machine learning model 306 can be used to train the base machine learning model 306. In some cases, the base machine learning model 306 is an unsupervised machine learning model. Training the base machine learning model 306 can be similar to the operations associated with training an unsupervised machine learning model as described previously with respect to Figure 2 and details are omitted here for brevity. In some cases, the base machine learning model 306 is a supervised machine learning model. Training the base machine learning model 306 can be similar to the operations associated with training a supervised machine learning model as described previously with respect to Figure 2 and details are omitted here for brevity.
[0104] The model parameters of the base machine learning model 306 can be continuously updated during the training process and converge to a certain value after training. In some implementations, the model parameters can include at least one of the model type of the base machine learning mode 306, the model weights of the base machine learning mode 306, the loss function of the base machine learning mode 306, or the optimization algorithm of the base machine learning mode 306. The model type of the base machine learning mode 306 can be, for example, supervised, unsupervised, or other suitable model types. The model weights of the base machine learning mode 306 can include, for example, coefficients, intercepts, feature weights, weights in a neural network, biases in a neural network, or other suitable model weights.
[0105] In some cases, a user can request to create a user-specific machine learning model (e.g., any of the user-specific machine learning models 308) that can be called, for example, exclusively by the user. For example, the user may be a VIP advertiser who does not want to share their machine learning model data with others and can thus request a user-specific machine learning model. In some cases, creating the user-specific machine learning model 308 can include obtaining the model parameters associated with the base machine learning model 306 and generating the user-specific machine learning model 308 that includes the model parameters associated with the base machine learning model 306. In one example, the user-specific machine learning model 308 is an exact copy of the base machine learning model 306. In another example, the user-specific machine learning model 308 can include some, but not all, of the model parameters of the base machine learning model 306.
[0106] The user-specific machine learning model 308 can be refined using user-specific training data. Examples of user-specific training data can include, but are not limited to, performance data associated with digital components generated by the user-specific machine learning model 308 (e.g., performance data associated with an advertisement), input prompts to the user-specific machine learning model 308, query context of the user-specific machine learning model 308, personalized data of the user (e.g., the user's landing page information, the user's design preferences, or the user's background information), or other types of user-specific training data. In some implementations, the general training data used to train the base machine learning model 306 can be shared with the user-specific machine learning model 308 and can be used to train the user-specific machine learning model 308.
[0107] In some cases, the user-specific machine learning model 308 is an unsupervised machine learning model. Refining the user-specific machine learning model 308 can be similar to the operations associated with training an unsupervised machine learning model as previously described Figure 2 and details are omitted here for brevity. In some cases, the user-specific machine learning model 308 is a supervised machine learning model. Refining the user-specific machine learning model 308 can be similar to the operations associated with training a supervised machine learning model as previously described Figure 2 and details are omitted here for brevity.
[0108] In some implementations, the model parameters of the user-specific machine learning model 308 can be updated after refinement. Thus, the model parameters of the user-specific machine learning model 308 and the model parameters of the base machine learning model 306 (which are used to generate the user-specific machine learning model 308) can have different values for at least one model parameter. In some cases, the model parameters of the user-specific machine learning model 308 cannot be shared with the base machine learning model 306 and / or any other user-specific machine learning model 308. This can ensure that the base machine learning model 306 and / or any other user-specific machine learning model 308 cannot use the model parameters of the user-specific machine learning model 308 to generate the same or similar digital components for which the user's intent is to remain private.
[0109] In some cases, more than one data sharing option can be provided to a user, with each data sharing option corresponding to a respective data privacy level. For example, two data sharing options can be provided to the user, namely a first data sharing option and a second data sharing option. The first data sharing option indicates that data associated with the user-specific machine learning model 308 cannot be shared with others (including the base machine learning model 306 and any other user-specific machine learning models 308). The second data sharing option indicates that at least a portion of the data associated with the user-specific machine learning model 308 can be shared with the base machine learning model 306 and / or other user-specific machine learning models 308. The data associated with the user-specific machine learning model 308 can include at least one of model parameters associated with the user-specific machine learning model 308, queries input to the user-specific machine learning model 308, outputs of the user-specific machine learning model 308, or training data used to train the user-specific machine learning model 308. When the user selects the first data sharing option, none of the data associated with the user-specific machine learning model 308 can be shared with the base machine learning model 306 or any other user-specific machine learning model 308. This prevents the base machine learning model 306 and / or another user-specific machine learning model 308 from using the data associated with the user-specific machine learning model 308 to generate digital components.
[0110] User 310 can invoke one or more user-specific machine learning models 308 to generate digital components. For example, the user can submit a query via an AI system (e.g., AI system 160), and the AI system can use one or more user-specific machine learning models 308 to generate a digital component based on the query. These operations can be similar to the operations described previously with respect to Figure 2 and details are omitted here for brevity.
[0111] In some cases, the user can invoke only one user-specific machine learning model 308 (e.g., using a subscription) to generate a digital component. In some cases, the user can invoke more than one user-specific machine learning model 308 to generate a digital component. In some implementations, the user can invoke both the base machine learning model 306 and the user-specific machine learning model 308 to generate a digital component.
[0112] Although Figure 3Two hierarchical levels are depicted, but system 300 can be implemented with more than two hierarchical levels. For example, a third hierarchical level can exist below the second hierarchical level 304, and the third hierarchical level can include one or more machine learning models. Similar to the operation of the machine learning model at the second hierarchical level 304, the model parameters of the machine learning model at the second hierarchical level 304 can be used to generate the machine learning model at the third hierarchical level. If the user chooses not to share data, the data associated with the machine learning model at the third hierarchical level cannot be shared with any other machine learning model.
[0113] In some implementations, the techniques described with respect to Figure 3 can be used in the context of using a generative model to generate advertisements. In one example use case, the base machine learning model 306 can be a generative model that is trained to generate advertisements using general data that can be shared with the base machine learning model 306. Each of the users 310 can be an advertiser. In some cases, users do not want to share their data (e.g., product information, user preferences, or trade secrets) with others, so they can request the generation of user-specific machine learning models 308. The user-specific machine learning models 308 can be trained using the users' data, and the model parameters and / or other data associated with the user-specific machine learning models 308 cannot be shared with others (including the base machine learning model 306 and any other user-specific machine learning models 308). By doing so, advertisers can utilize the generative model to create advertisements while maintaining their data privacy. Moreover, since the user-specific machine learning models 308 are trained using the users' data, the user-specific machine learning models 308 can generate digital components that are more likely to meet the users' preferences.
[0114] Figure 4FIG. 0 is a block diagram of an example system 400 for implementing a domain - specific machine - learning model according to an implementation of the present disclosure. At a high level, the example system 400 relates to a solution for generating a machine - learning model specifically trained for a domain (e.g., a vertical market or a geographic region) and using the machine - learning model to generate domain - specific digital components. Initially, a base machine - learning model can be trained. The model parameters of the base machine - learning model can be used to generate a domain - specific machine - learning model, which can be continuously refined using domain - specific training data. Due to the domain - specific training, the domain - specific machine - learning model can generate digital components that are more suitable for the domain than the digital components generated by the base machine - learning model. Using similar operations, a sub - domain - specific machine - learning model can be generated under the domain - specific machine - learning model to generate digital components for a more fine - grained sub - domain (e.g., a vertical sub - market under a vertical market or a geographic sub - region under a geographic region). The example system 400 can be extended deeper using similar operations. More details are described below.
[0115] The example system 400 includes three hierarchical levels of machine - learning models, namely a first hierarchical level 402, a second hierarchical level 404 below the first hierarchical level 402, and a third hierarchical level 406 below the second hierarchical level 404. The first hierarchical level 402 includes a base machine - learning model 408. The second hierarchical level 404 can include one or more domain - specific machine - learning models 410 (e.g., machine - learning models for domains A, B, and N, as shown). The third hierarchical level 406 can include one or more sub - domain - specific machine - learning models 412 (e.g., machine - learning models for sub - domains A1, A2, B1, N1, and N2, as shown). In some cases, each of the base machine - learning model 408, the domain - specific machine - learning model 410, and the sub - domain - specific machine - learning model 412 is a generative model. A user 414 (e.g., any one of users A to N, as shown) can invoke the sub - domain - specific machine - learning model 412, the domain - specific machine - learning model 410, and / or the base machine - learning model 408 to generate digital components. The base machine - learning model 408 can be similar to the base machine - learning model 306 operationally and / or structurally. Similarly, the domain - specific machine - learning model 410 and / or the sub - domain - specific machine - learning model 412 can be similar to the user - specific machine - learning model 308 operationally and / or structurally.
[0116] As shown, all the machine - learning models in the example system 400 are included in a tree, where the base machine - learning model 408 is the root of the tree. The tree includes multiple sub - trees, where each domain - specific machine - learning model 410 is the root of a sub - tree. Each sub - domain - specific machine - learning model 412 is a leaf of the tree.
[0117] A tree can be defined by defining one or more domains and / or one or more sub-domains. One or more domains can be defined, where each domain represents, for example, a vertical market, a geographical region, or other suitable classification. In one example, a domain is defined based on a vertical market and can include fashion, travel, automotive, and / or other suitable vertical markets. In another example, a domain is defined based on a geographical region and can include North America, Europe, Asia, and / or other suitable geographical regions.
[0118] In some cases, one or more sub-domains can be defined under a domain, where each sub-domain represents, for example, a vertical sub-market, a geographical sub-region, or other suitable classification. In an example where domains are divided based on vertical markets, sub-domains such as clothing, jewelry, and / or other suitable classifications can be defined under the fashion domain. In another example where domains are divided based on geographical regions, sub-domains such as the United States and Canada can be defined under the North America domain.
[0119] Training the base machine learning model 408 can be similar to the operations associated with training the base machine learning model 306 as described previously Figure 3 and details are omitted here for brevity. For example, the base machine learning model 408 can be domain-neutral and can be trained without assigning specific weights to training data from any particular domain.
[0120] Similarly, generating and refining the domain-specific machine learning model 410 and / or the sub-domain-specific machine learning model 412 can be similar to the operations associated with generating and refining the user-specific machine learning model 308 as described previously Figure 3 and details are omitted here for brevity. For example, the model parameters of the base machine learning model 408 can be used to generate the domain-specific machine learning model 410, and the model parameters of the domain-specific machine learning model 410 can be used to generate the sub-domain-specific machine learning model 412. In some cases, the data associated with the domain-specific machine learning model 410 cannot be shared with the base machine learning model 408. In some cases, the data associated with the sub-domain-specific machine learning model 412 cannot be shared with any other machine learning model (including the base machine learning model 408, any domain-specific machine learning model 410, and any other sub-domain-specific machine learning model 412).
[0121] In some cases, the lower a machine learning model is in the hierarchy, the more specialized the machine learning model is in a particular domain. Thus, for example, the domain-specific machine learning model 410 can be more specialized in generating digital components for a particular domain than the base machine learning model 408, and the sub-domain-specific machine learning model 412 can be more specialized in generating digital components for a particular sub-domain than the domain-specific machine learning model 410. For example, the base machine learning model 408 can be trained to generate general advertisements. When initializing using the model parameters of the base machine learning model 408, the domain-specific machine learning model 410 can be refined to generate advertisements for the fashion vertical market. Similarly, when initializing using the model parameters of the domain-specific machine learning model 410, the sub-domain-specific machine learning model 412 can be refined to generate advertisements in a sub-domain (e.g., clothing or jewelry) within the fashion vertical market. Thus, a machine learning model at a lower hierarchy level can be derived from another machine learning model at a higher hierarchy level by leveraging the model parameters of the trained machine learning model at the higher hierarchy level without overtraining. At the same time, the machine learning model at the lower hierarchy level can be refined to generate digital components that are more likely to meet the unique needs or requirements of users from a particular domain.
[0122] User 414 can invoke one or more sub-domain-specific machine learning models 412 to generate digital components. For example, a user can submit a query (e.g., Figure 2 query 246) via an AI system (e.g., AI system 160), and the AI system can invoke one or more sub-domain-specific machine learning models 412 to generate digital components based on the query. These operations can be similar to the operations previously described with respect to Figure 2 and details are omitted here for brevity.
[0123] In some cases, a user (e.g., user A) can invoke (e.g., using a subscription) only one sub-domain-specific machine learning model 412 to generate digital components. In some cases, a user (e.g., user B or user N) can invoke more than one sub-domain-specific machine learning model 412 to generate digital components. In some implementations, a user can invoke any number of base machine learning models 408, domain-specific machine learning models 410, and / or sub-domain-specific machine learning models 412 to generate digital components.
[0124] In some implementations, an AI system (e.g., AI system 160) can identify a machine learning model at an appropriate hierarchical level for generating digital components based on context data. For example, the AI system can identify the hierarchical level corresponding to (e.g., most similar to) the context data and invoke the machine learning model of that hierarchical level to generate digital components. Context data can be obtained from, for example, queries, additional query data (e.g., the additional query data described regarding Figure 2 the described additional query data) or other suitable sources. In some cases, the AI system can perform an analysis of the context data from at least one data source (e.g., semantic analysis, image analysis, audio analysis, video analysis, or other suitable analysis). In some implementations, the context data from a single data source is analyzed. In some implementations, the context data from more than one data source is analyzed. The analysis can be performed by one or more machine learning algorithms, the overall goal of which is to identify one or more appropriate hierarchical levels for generating digital components. In some cases, the AI system can perform the analysis using a neural network array of machine learning algorithms that can use concatenated operations or can include parallel operations or operate independently of each other. In some implementations, traditional data analysis can be appended to or performed separately from the machine learning process.
[0125] For example, assume that the tree of machine learning models includes three hierarchical levels. The first hierarchical level includes a base machine learning trained to generate general advertisements. The second hierarchical level includes a first domain-specific machine learning model trained to generate advertisements for clothing and a second domain-specific machine learning model trained to generate advertisements for sports products. Under the first domain-specific machine learning model, there can be three sub-domain-specific machine learning models respectively trained to generate advertisements for coats, shoes, and holiday clothing. Under the second domain-specific machine learning model, there can be three sub-domain-specific machine learning models respectively trained to generate advertisements for hockey, basketball, and football. Assume that a user input provides a prompt indicating that the user wants to generate an advertisement for a product and the input includes additional query data of an image of the product. The AI system can perform image analysis on the image provided by the user and identify the objects included in the image. In one example, assume that the image includes a coat. The AI system can infer that the user may be more interested in an advertisement for the vertical sub-market related to coats. Thus, the AI system can identify the sub-domain-specific machine learning model trained for coats to generate the advertisement. In another example, assume that the image includes multiple clothing objects, such as a coat, shoes, and pants. The AI system can infer that the user may be more interested in an advertisement for the vertical market related to clothing. Thus, the AI system can identify the first domain-specific machine learning model trained for clothing to generate the advertisement.
[0126] By identifying a pre-trained machine learning model at an appropriate hierarchical level for generating digital components, the digital components generated by the identified machine learning model are more likely to be accepted by the user and will likely require less post-processing when modifying the generated digital components to meet the user's needs. Thus, resources can be saved by reducing post-processing as compared to using a general machine learning model to generate digital components and modifying the generated digital components at run-time or post-generation.
[0127] Although Figure 4 Three hierarchical levels are depicted in, the system 400 can be implemented with more than three hierarchical levels. For example, a fourth hierarchical level can exist below the third hierarchical level 406, and the fourth hierarchical level can include one or more machine learning models (e.g., machine learning models for even more fine-grained domains under a sub-domain). Similar to the operation of the machine learning models at the second or third hierarchical levels, the model parameters of the machine learning models at the third hierarchical level 406 can be used to generate the machine learning models at the fourth hierarchical level. The data associated with the machine learning models at the fourth hierarchical level cannot be shared with any other machine learning models.
[0128] In some implementations, the techniques described with respect to Figure 4 can be used in the context of using a generative model to generate advertisements. The base machine learning model 408 can be a generative model that is trained to generate advertisements using general data that can be shared with the base machine learning model 408. Each of the users 414 can be an advertiser. One example use case involves generating advertisements for a specific vertical market. In some cases, the design styles from one vertical market (e.g., the fashion industry) may be very different from another vertical market (e.g., the automotive industry). Thus, two different vertical markets may expect different machine learning models. On the other hand, the machine learning models for two different vertical markets can share a lot in common (e.g., advertisements for both the fashion industry and the automotive industry need to meet certain regulations). The techniques described herein enable training a base machine learning model for generating general advertisements and generating at least two separate vertical market-specific machine learning models (e.g., one for the fashion industry and another for the automotive industry) for generating vertical market-specific digital components. Additionally, similar operations can be used to generate machine learning models for vertical sub-markets, and the machine learning models for vertical sub-markets can generate advertisements that are more specific to their respective vertical sub-markets. Thus, for example, the model parameters of a fashion-specific machine learning model can be used to generate a clothing-specific machine learning model and a jewelry-specific machine learning model to generate advertisements in their respective vertical sub-markets.
[0129] Another example use case involves generating advertisements for a specific geographical region. Similar to the example use case of generating advertisements for a specific vertical market, the design style from one geographical region (e.g., North America) may be very different from another geographical region (e.g., Asia). The techniques described herein allow for training a base machine learning model to generate general advertisements, as well as generating at least two separate geographical region-specific machine learning models (e.g., one for North America and another for Asia) for generating geographical region-specific digital components.
[0130] Figure 5 is a flowchart of an example process 500 for implementing a user-specific machine learning model according to an implementation of the present disclosure. The operations of process 500 may be performed, for example, by Figure 1 service device 110 or another data processing device. The operations of process 500 may also be implemented as instructions stored on a computer-readable medium, which may be non-transitory. Executing the instructions by one or more data processing devices causes the one or more data processing devices to perform the operations of process 500.
[0131] At 502, an AI system (e.g., AI system 160) obtains first model parameters associated with a first machine learning model (e.g., base machine learning model 306) at a first hierarchical layer (e.g., first hierarchical layer 302). In some cases, the first model parameters include at least one of a model type of the machine learning model, model weights of the machine learning model, a loss function of the machine learning model, or an optimization algorithm of the machine learning model.
[0132] At 504, the AI system generates a second machine learning model (e.g., user-specific machine learning model 308) that includes first model parameters, and the second machine learning model is at a second hierarchy level (e.g., second hierarchy level 304) below the first hierarchy level. In some cases, the AI system provides a first data sharing option and a second data sharing option, where the first data sharing option indicates that data associated with the second machine learning model is not shareable with the first machine learning model, and where the second data sharing option indicates that at least a portion of the data associated with the second machine learning model is shareable with the first machine learning model. In some implementations, the AI system receives an indication that the first data sharing option has been selected. In some examples, the data associated with the second machine learning model includes at least one of model parameters associated with the second machine learning model, queries input to the second machine learning model, outputs of the second machine learning model, or training data used to train the second machine learning model. In some implementations, the training data used to train the first machine learning model may be shared with the second machine learning model. In some cases, each of the first machine learning model and the second machine learning model is a generative model. In some implementations, a third machine learning model is at a third hierarchy level below the second hierarchy level, a fourth machine learning model is at a fourth hierarchy level below the third hierarchy level, and data associated with the third machine learning model is not shareable with the second machine learning model but is shareable with the fourth machine learning model.
[0133] At 506, the AI system refines the second machine learning model to obtain an updated second machine learning model that includes second model parameters, where the first model parameters and the second model parameters have different values for at least one model parameter, and the second model parameters are not shareable with the first machine learning model. In some examples, the AI system receives a query (e.g., Figure 2 query 246), uses the second machine learning model to generate digital components, and generates training data based on the digital components, where the query, the digital components, and the training data are not shareable with the first machine learning model.
[0134] In some cases, the AI system obtains additional query data (e.g., regarding Figure 2 additional query data described) that includes data different from the query, where the additional query data restricts the digital components generated by the second machine learning model and is not shareable with the first machine learning model. The AI system generates a prompt that includes the query and the additional query data. The AI system inputs the prompt into the second machine learning model. The second machine learning model generates digital components.
[0135] In some examples, the second machine learning model is a supervised machine learning model. The AI system includes queries as features of the training data and includes at least one of digital components or algorithms for generating digital components as labels of the training data.
[0136] In some examples, the second machine learning model is trained using an RL algorithm. The AI system includes at least one of digital components, algorithms for generating digital components, or rewards of digital components in the training data.
[0137] Figure 6 is a flowchart of an example process 600 for implementing a domain-specific machine learning model according to an implementation of the present disclosure. The operations of process 600 can be performed, for example, by Figure 1 service device 110 or another data processing device. The operations of process 600 can also be implemented as instructions stored on a computer-readable medium, which can be non-transitory. Executing the instructions by one or more data processing devices causes the one or more data processing devices to perform the operations of process 600.
[0138] At 602, the AI system obtains first model parameters of a first machine learning model (e.g., base machine learning model 408). The first machine learning model can be at a first hierarchical layer (e.g., first hierarchical layer 402). In some cases, the first model parameters include at least one of the model type of the machine learning model, the model weights of the machine learning model, the loss function of the machine learning model, or the optimization algorithm of the machine learning model.
[0139] At 604, the AI system generates a second machine learning model (e.g., domain-specific machine learning model 410 or sub-domain-specific machine learning model 414) based on the first model parameters, where the second machine learning model is at a given hierarchical layer (e.g., second hierarchical layer 404 or third hierarchical layer 406) of the hierarchical model structure.
[0140] At 606, the AI system refines the second machine learning model to obtain an updated second machine learning model including second model parameters, where the first model parameters and the second model parameters have different values for at least one model parameter. In some cases, the second model parameters cannot be shared with the first machine learning model.
[0141] At 608, the AI system generates a third machine learning model that includes second model parameters, where (i) the third machine learning model is at a lower hierarchical layer of the hierarchical model structure below a given hierarchical layer of the second machine learning model, and (ii) the second machine learning model and the third machine learning model have a first common attribute. In some implementations, the first common attribute includes at least one of a common vertical market or a common geographical region. In some cases, the AI system refines the third machine learning model to obtain an updated third machine learning model that includes third model parameters, where the third model parameters and the second model parameters have at least one different model parameter, and the third model parameters are not shareable with the second machine learning model.
[0142] In some examples, the first machine learning model is the root of a tree of machine learning models, the first machine learning model has one or more machine learning models as one or more child nodes of the first machine learning model, and each child node of the first machine learning model is the root of a subtree, where the first subtree of machine learning models includes the second machine learning model that is the root of the first subtree and all machine learning models below the second machine learning model, and where each machine learning model of the first subtree of machine learning models has the first common attribute. In some cases, an additional second machine learning model is at a given hierarchical layer, the second subtree of machine learning models includes the additional second machine learning model that is the root of the second subtree and all machine learning models below the additional second machine learning model, and where each machine learning model of the second subtree of machine learning models has a second common attribute that is different from the first common attribute.
[0143] In some examples, data associated with the first subtree of machine learning models is not shareable with the second subtree of machine learning models, and data associated with the second subtree of machine learning models is not shareable with the first subtree of machine learning models. In some examples, data associated with the first subtree of machine learning models includes at least one of model parameters of the first subtree of machine learning models, queries to the machine learning models of the first subtree of machine learning models, outputs of the machine learning models of the first subtree of machine learning models, or training data for training the machine learning models of the first subtree of machine learning models. In some cases, a user (e.g., user 414) is allowed access to the machine learning models of the first subtree of machine learning models and the additional models of the second subtree of machine learning models.
[0144] In some examples, the AI system can receive a query. The AI system can obtain context data associated with the query. The AI system can identify a hierarchical layer of the hierarchical model structure based on the context data. The machine learning model at the hierarchical layer can be used to generate digital components.
[0145] Figure 7 FIG. 700 is a block diagram of an example computer system 700 that may be used to perform the described operations according to an implementation of the present disclosure. The system 700 includes a processor 710, a memory 720, a storage device 730, and an input / output device 740. Each of the components 710, 720, 730, and 740 may be interconnected, for example, using a system bus 750. The processor 710 is capable of processing instructions for execution within the system 700. In one implementation, the processor 710 is a single-threaded processor. In another implementation, the processor 710 is a multi-threaded processor. The processor 710 is capable of processing instructions stored in the memory 720 or on the storage device 730.
[0146] The memory 720 stores information within the system 700. In one implementation, the memory 720 is a computer-readable medium. In one implementation, the memory 720 is a volatile memory unit. In another implementation, the memory 720 is a non-volatile memory unit.
[0147] The storage device 730 is capable of providing mass storage for the system 700. In one implementation, the storage device 730 is a computer-readable medium. In various different implementations, the storage device 730 may include, for example, a hard disk device, an optical disk device, a storage device shared by multiple computing devices over a network (e.g., a cloud storage device), or some other mass storage device.
[0148] The input / output device 740 provides input / output operations for the system 700. In one implementation, the input / output device 740 may include one or more network interface devices (e.g., an Ethernet card), a serial communication device (e.g., and an RS-232 port), and / or a wireless interface device (e.g., and an 802.11 card). In another implementation, the input / output device may include a driver device configured to receive input data and send output data to other devices (e.g., a keyboard, a printer, a display, and other peripheral devices 760). However, other implementations may also be used, such as mobile computing devices, mobile communication devices, and set-top box TV client devices.
[0149] Although an example processing system has been described in Figure 7 , implementations of the subject matter and functional operations described in this specification may be implemented in other types of digital electronic circuitry (circuitry) or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents) or in a combination of one or more of them.
[0150] An electronic document (referred to simply as a document for brevity) need not correspond to a file. A document can be stored as part of a file that holds other documents, in a single file dedicated to the document in question, or in a collection of coordinated files.
[0151] In cases where the systems discussed here collect and / or use personal information about users, users may be provided with an opportunity to enable / disable or control programs or features that can collect and / or use personal information (e.g., information about a user's social network, social actions or activities, a user's preferences, or a user's current location). Additionally, certain data can be disposed of in one or more ways before it is stored or used, so that personally identifiable information associated with the user is removed. For example, a user's identity can be anonymized so that the user's personally identifiable information cannot be determined, or a user's geographic location can be generalized (such as to the city, ZIP code, or state level) if location information is obtained, so that the user's specific location cannot be determined.
[0152] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents) or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a computer storage medium for execution by, or to control the operation of, data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, although a computer storage medium is not a propagated signal, a computer storage medium can be the source or destination of computer program instructions encoded in an artificially generated propagated signal. A computer storage medium can also be one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices), or be included in one or more separate physical components or media.
[0153] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0154] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines for processing data, including, for example, programmable processors, computers, system-on-chips, or multiple of the foregoing items or combinations thereof. The apparatus may include dedicated logic circuitry, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer programs involved, for example, code that constitutes processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations of one or more of them. The apparatus and the execution environment may implement various different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.
[0155] This document relates to service apparatuses. As used herein, a service apparatus is one or more data processing apparatuses that perform operations to facilitate the distribution of content over a network. A service apparatus is depicted as a single block in a block diagram. However, while a service apparatus may be a single device or a single set of devices, the present disclosure contemplates that a service apparatus may also be a group of devices, or even multiple different systems that communicate to provide various content to client devices. For example, a service apparatus may encompass one or more of a search system, a video streaming service, an audio streaming service, an email service, a navigation service, an advertising service, a gaming service, or any other service.
[0156] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language (including compiled or interpreted languages, declarative or procedural languages), and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. The program may be stored in a part of a file that holds other programs or data (such as one or more scripts in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (such as files that store one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0157] The processes and logical flows described in this specification may be performed by one or more programmable processors that execute one or more computer programs to perform actions by operating on input data and generating output. The processes and logical flows may also be performed by dedicated logic circuitry, and the apparatus may also be implemented as dedicated logic circuitry, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits).
[0158] For example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read only memory or a random access memory (RAM) or both. The basic elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing the instructions and data. In general, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from or transfer data to or both from the one or more mass storage devices. However, a computer need not have such devices. In addition, a computer may be embedded in another device (e.g., a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few). Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0159] To provide for interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user may provide input to the computer. Other kinds of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user may be in any form, including sound, speech, or tactile input. Additionally, a computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page in response to a request received from a web browser to a web browser on a client device of the user.
[0160] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0161] The computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, the server sends data (e.g., an HTML page) to the client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., the result of a user interaction) can be received at the server from the client device.
[0162] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination can be deleted from the combination, and the claimed combination may refer to a sub-combination or variation of a sub-combination.
[0163] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0164] Accordingly, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures need not be in the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.
Claims
1. A computer-implemented method, comprising: obtaining, by an artificial intelligence (AI) system, first model parameters associated with a first machine learning model at a first hierarchical layer; generating, by the AI system, a second machine learning model including the first model parameters, the second machine learning model being at a second hierarchical layer below the first hierarchical layer; and refining, by the AI system, the second machine learning model to obtain an updated second machine learning model including second model parameters, wherein the first model parameters and the second model parameters have different values for at least one model parameter, and the second model parameters cannot be shared with the first machine learning model.
2. The computer-implemented method according to claim 1, wherein the first model parameters include at least one of a model type of a machine learning model, model weights of a machine learning model, a loss function of a machine learning model, or an optimization algorithm of a machine learning model.
3. The computer-implemented method according to claim 1, comprising: providing, by the AI system, a first data sharing option and a second data sharing option, wherein the first data sharing option indicates that data associated with the second machine learning model cannot be shared with the first machine learning model, and wherein the second data sharing option indicates that at least a portion of the data associated with the second machine learning model can be shared with the first machine learning model.
4. The computer-implemented method according to claim 3, comprising: receiving, by the AI system, an indication that the first data sharing option is selected.
5. The computer-implemented method according to claim 3, wherein the data associated with the second machine learning model includes at least one of model parameters associated with the second machine learning model, a query input to the second machine learning model, an output of the second machine learning model, or training data used to train the second machine learning model.
6. The computer-implemented method according to claim 1, comprising: receiving, by the AI system, a query; generating, by the AI system, a digital component using the second machine learning model; and generating, by the AI system, training data based on the digital component, wherein the query, the digital component, and the training data cannot be shared with the first machine learning model.
7. The computer-implemented method according to claim 6, comprising: obtaining, by the AI system, additional query data including data different from the query, wherein the additional query data restricts the digital component generated by the second machine learning model and cannot be shared with the first machine learning model; generating, by the AI system, a prompt including the query and the additional query data; inputting, by the AI system, the prompt into the second machine learning model; and generating, by the second machine learning model, the digital component.
8. The computer-implemented method according to claim 6, wherein: the second machine learning model is a supervised machine learning model; and Generating the digital component by the AI system using the second machine learning model includes: Including the query as a feature of the training data; And Including at least one of the digital component or the algorithm for generating the digital component as a label of the training data.
9. The computer-implemented method according to claim 6, wherein: The second machine learning model is trained using a reinforcement learning (RL) algorithm; And Generating the digital component by the AI system using the second machine learning model includes: Including at least one of the digital component, the algorithm for generating the digital component, or the reward of the digital component in the training data.
10. The computer-implemented method according to claim 1, wherein a third machine learning model is at a third hierarchical layer below the second hierarchical layer, a fourth machine learning model is at a fourth hierarchical layer below the third hierarchical layer, and data associated with the third machine learning model cannot be shared with the second machine learning model but can be shared with the fourth machine learning model.
11. The computer-implemented method according to claim 1, wherein the training data for training the first machine learning model can be shared with the second machine learning model.
12. The computer-implemented method according to claim 1, wherein each of the first machine learning model and the second machine learning model is a generative model.
13. A computer-implemented artificial intelligence (AI) system, comprising: One or more processors; And One or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: Obtaining, by an artificial intelligence (AI) system, first model parameters associated with a first machine learning model at a first hierarchical layer; Generating, by the AI system, a second machine learning model including the first model parameters, the second machine learning model being at a second hierarchical layer below the first hierarchical layer; and Refining, by the AI system, the second machine learning model to obtain an updated second machine learning model including second model parameters, wherein the first model parameters and the second model parameters have different values for at least one model parameter, and the second model parameters cannot be shared with the first machine learning model.
14. The computer-implemented AI system according to claim 13, wherein the first model parameters include at least one of a model type of a machine learning model, model weights of a machine learning model, a loss function of a machine learning model, or an optimization algorithm of a machine learning model.
15. The computer-implemented AI system according to claim 13, the operations including: The AI system provides a first data sharing option and a second data sharing option, where the first data sharing option indicates that data associated with the second machine learning model cannot be shared with the first machine learning model, and where the second data sharing option indicates that at least a portion of the data associated with the second machine learning model can be shared with the first machine learning model.
16. The computer-implemented AI system of claim 15, the operations comprising: Receiving, by the AI system, an indication that the first data sharing option has been selected.
17. The computer-implemented AI system of claim 15, where the data associated with the second machine learning model includes at least one of model parameters associated with the second machine learning model, queries input to the second machine learning model, outputs of the second machine learning model, or training data used to train the second machine learning model.
18. The computer-implemented AI system of claim 13, the operations comprising: Receiving, by the AI system, a query; Generating, by the AI system, digital components using the second machine learning model; And Generating, by the AI system, training data based on the digital components, where the query, the digital components, and the training data cannot be shared with the first machine learning model.
19. The computer-implemented AI system of claim 18, the operations comprising: Obtaining, by the AI system, additional query data including data different from the query, where the additional query data restricts the digital components generated by the second machine learning model and cannot be shared with the first machine learning model; Generating, by the AI system, a prompt including the query and the additional query data; Inputting, by the AI system, the prompt into the second machine learning model; And Generating, by the second machine learning model, the digital components.
20. One or more non-transitory computer-readable media storing instructions that, when executed by a computer-implemented artificial intelligence (AI) system, cause the computer-implemented AI system to perform operations, the operations comprising: Obtaining, by the artificial intelligence (AI) system, first model parameters associated with a first machine learning model at a first hierarchical level; Generating, by the AI system, a second machine learning model including the first model parameters, the second machine learning model being at a second hierarchical level below the first hierarchical level; and Refining, by the AI system, the second machine learning model to obtain an updated second machine learning model including second model parameters, where the first model parameters and the second model parameters have different values for at least one model parameter, and the second model parameters cannot be shared with the first machine learning model.