Robust model performance across heterogeneous subgroups within the same group
By adding error regularization terms or divergence minimization terms to the loss function, the problem of large performance changes in machine learning models when training on unbalanced user data sets is solved, achieving more accurate user characteristic prediction and more consistent user experience.
Patent Information
- Application Number
- CN202080019637.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-09-30
AI Technical Summary
When existing machine learning models are trained on unbalanced user data sets, they are prone to cause significant changes in model performance between different user subgroups, affecting the accuracy of prediction and user experience.
By modifying the loss function, adding error regularization terms or divergence minimization terms, reducing the difference in performance measurements between different user subgroups, thus training a more robust model.
The improved model can more accurately predict user characteristics, improve prediction consistency across subgroups of users, enhance user experience, and reduce the amount of sensitive data transmission, protect user privacy.
Smart Images

Figure CN114600125B_ABST
Abstract
Description
Technical Field
[0001] This specification involves data processing and machine learning models. Background Art
[0002] The client device can use an application (e.g., a web browser, a native application) to access a content platform (e.g., a search platform, a social media platform, or another platform that hosts content). The content platform can display a digital component (a discrete unit of digital content or digital information, such as, for example, a video clip, an audio clip, a multimedia clip, an image, text, or another unit of content) within an application launched on the client device that may be provided by one or more content sources / platforms. Summary of the invention
[0003] Generally, an innovative aspect of the subject matter described in this specification can be embodied in a method comprising the following operations: identifying a loss function for a model to be trained, the loss function generating a loss representing a performance measure that the model seeks to optimize during training; modifying the loss function, including adding an additional term to the loss function, the additional term reducing the difference in performance measures across different user subgroups all represented by the same user group identifier, wherein each of the different user subgroups has characteristics that are different from characteristics of other subgroups among the different user subgroups; training the model using the modified loss function; receiving a request for a digital component from a client device, the request including a given user group identifier for a particular user group among the different user groups; generating one or more user characteristics not included in the request by applying the trained model to information included in the request; selecting one or more digital components based on the one or more user characteristics generated by the trained model; and sending the selected one or more digital components to the client device.
[0004] Other implementations of this aspect include corresponding apparatus, systems, and computer programs configured to execute aspects of the method encoded on a computer storage device.These and other implementations can each optionally include one or more of the following features.
[0005] In some aspects, modifying the loss function includes adding an error regularization term to the loss function. In some aspects, modifying the loss function includes adding a divergence minimization term to the loss function.
[0006] In some aspects, adding an error regularization term to the loss function includes adding a loss variance term to the loss function, wherein the loss variance function characterizes the squared difference between an average loss of a model within a group of users based on a certain attribute and an average loss of the model across all users, wherein the difference can be calculated separately across users based on different attributes.
[0007] In some aspects, adding an error regularization term to the loss function includes adding a maximum weighted loss difference term to the loss function, wherein the maximum weighted loss difference term is the maximum weighted difference between the loss in the user group and all users in all different user groups, and the method also includes using a function that quantifies the importance of each user group.
[0008] In some aspects, adding the error regularization term to the loss function includes adding a coarse loss variance term to the loss function, wherein the coarse loss variance term is an average of squared differences between losses for the first user group and the second user group conditioned on individual user attributes.
[0009] In some aspects, adding an error regularization term to the loss function includes adding a HSIC regularization term to the loss function, wherein the HSIC term characterizes differences in losses across different user groups in a non-parametric manner and independently of the distribution of users in the different user groups.
[0010] In some aspects, adding a divergence minimization term to the loss function includes adding one of a mutual information term or a Kullback-Leibler divergence term to the loss function, wherein the mutual information term characterizes the similarity of distributions of model predictions across multiple user groups, and wherein the Kullback-Leibler divergence term characterizes the difference in distributions of model predictions across multiple user groups.
[0011] Specific embodiments of the subject matter described in this specification can be implemented to achieve one or more of the following advantages. Machine learning models can be trained to predict user characteristics rather than collecting specific information from users, for example via third-party cookies, thereby respecting user privacy issues. However, implementing such methods requires training machine learning models on unbalanced data sets obtained from real-world users, resulting in more frequently observed subgroups having a higher degree of influence on the parameters of the model relative to less frequently observed subgroups. This may result in considerable variation in model performance across subgroups. Modifying the loss function of the machine learning model enables the machine learning model to learn and distinguish complex patterns of the training data set, thereby reducing errors in predictions about user characteristics and increasing the consistency of the prediction accuracy of the machine learning model across subgroups. Such an implementation allows the delivery of digital components that are carefully selected based on predicted user characteristics to users, thereby improving user experience and maintaining user privacy.
[0012] Embodiments of the subject matter described herein can reduce the amount of potentially sensitive data transmitted between a client device and the rest of a network. For example, in the event that a client device sends a request for a digital component, where a group identifier for a user indicates a less-observed subgroup, the methods described herein can avoid the need to send more information about the user in order to appropriately customize the content provided to the user. The reduction in additional data about the user provided by the client device reduces the bandwidth requirements of the client device that sends requests to users of the less-observed subgroup. This reduction may be more significant for subgroups that are less-observed due to factors consistent with sensitive bandwidth requirements (e.g., less-observed subgroups may be discovered more often in situations where local network connectivity is poor). In addition, by avoiding the transmission of additional information, the less-observed subgroup can be protected from identification by third parties.
[0013] The details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a block diagram of an example environment in which digital components are distributed.
[0015] Figure 2 is a block diagram of an example machine learning model implemented by a user evaluation device.
[0016] Figure 3 is a flowchart of an example process for distributing digital components using a modified loss function.
[0017] Figure 4 is a block diagram of an example computer system that can be used to perform the described operations. DETAILED DESCRIPTION
[0018] This document discloses methods, systems, apparatus, and computer-readable media for modifying a machine learning model to ensure that model performance varies as little as possible across different subgroups of users within the same user group.
[0019] Typically, digital components can be provided to users who are connected to the Internet via a client device. In such a scenario, the digital component provider may wish to provide the digital component based on the user's online activities and the user's browsing history. However, more and more users are choosing not to allow the collection and use of certain information, and third-party cookies are being blocked and / or deprecated by some browsers, making it necessary to perform digital component selection without using third-party cookies (i.e., cookies from a domain different from the domain of the web page that is allowed to access the content of the cookie file).
[0020] New technologies have emerged for distributing digital components to users by assigning them to user groups when they access a particular resource or perform a particular action at a resource (e.g., interacting with a particular item presented on a web page or adding an item to a virtual shopping cart). These user groups are typically created in a way that each user group includes a sufficient number of users so that individual users cannot be identified. Demographic information about users remains important for providing users with a personalized online experience, such as by providing specific digital components that are relevant to the user. However, due to the unavailability of such information, machine learning models that can predict such user information and / or characteristics can be implemented.
[0021] Even if such techniques and methods are state-of-the-art, machine learning models can still suffer greatly from imbalanced datasets. Since standard techniques for training machine learning models (such as empirical risk minimization, which seeks to minimize the average loss over the training data) are agnostic about the class distribution of the training dataset, learning from such datasets leads to degradation of model performance. For example, a machine learning model designed to predict user attributes and trained on an imbalanced dataset may perform poorly and result in a large mismatch between the user's true attributes and the model's predictions.
[0022] In contrast, the techniques and methods explained in this document outperform conventional techniques by modifying the loss function in a manner that achieves high accuracy in predicting user characteristics, thereby enabling the selection of relevant digital components for a user using the output of a machine learning model. More specifically, the loss function of the machine learning model to be trained is modified to include an additional term that reduces the differences in performance measures within or across different user subgroups and / or user subsets that are all represented by the same user group identifier (i.e., all included in the same larger user group), despite the fact that each of these different user subgroups has characteristics that are different from those of other user subgroups in the different user subgroups within the user group. This added term ensures the robustness of the model's performance across different user subgroups. Reference Figure 1-Figure 4 The techniques and methods are further explained.
[0023] Figure 1 1 is a block diagram of an example environment 100 in which digital components are distributed for presentation with electronic documents. Example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. Network 102 connects content server 104, client device 106, digital component server 108, and digital component distribution system 110 (also referred to as component distribution system (CDS)).
[0024] The client device 106 is an electronic device capable of requesting and receiving resources over the network 102. Example client devices 106 include personal computers, mobile communication devices, wearable devices, personal digital assistants, and other devices capable of sending and receiving data over the network 102. The client device 106 typically includes a user application 112, such as a web browser, to facilitate sending and receiving data over the network 102, but native applications executed by the client device 106 can also facilitate sending and receiving data over the network 102. The client device 106, particularly the personal digital assistant, can include hardware and / or software that enables voice interaction with the client device 106. For example, the client device 106 can include a microphone through which a user can submit audio (e.g., voice) input, such as commands, search queries, browsing instructions, smart home instructions, and / or other information. In addition, the client device 106 can include a speaker through which audio (e.g., voice) output can be provided to the user. The personal digital assistant can be implemented in any client device 106 , examples of which include a wearable device, a smart speaker, a home appliance, a car, a tablet device, or other client device 106 .
[0025] An electronic document is data that presents a set of content at a client device 106. Examples of electronic documents include web pages, word processing documents, portable document format (PDF) documents, images, videos, search result pages, and feeds. Native applications (e.g., "apps") such as applications installed on mobile, tablet, or desktop computing devices are also examples of electronic documents. Electronic documents can be provided to a client device 106 by a content server 104. For example, a content server 104 can include a server that hosts a publisher's website. In the example, a client device 106 can initiate a request for a given publisher's web page, and the content server 104 that hosts the given publisher's web page can respond to the request by sending machine-executable instructions that initiate presentation of the given web page at the client device 106.
[0026] In another example, the content server 104 can include an application server from which the client device 106 can download the application. In this example, the client device 106 can download the files required to install the app at the client device 106 and then execute the downloaded app locally. The downloaded app can be configured to present a combination of native content as part of the application itself and one or more digital components (e.g., content created / distributed by a third party) obtained from the digital component server 108 and inserted into the app while the app is being executed at the client device 106.
[0027] An electronic document can include a variety of content. For example, an electronic document can include static content (e.g., text or other specified content) that is within the electronic document itself and / or does not change over time. An electronic document can also include dynamic content that can change over time or on a per-request basis. For example, a publisher of a given electronic document can maintain data sources for populating various portions of an electronic document. In this example, a given electronic document can include a tag or script that causes a client device 106 to request content from a data source when the given electronic document is processed (e.g., rendered or executed) by the client device 106. The client device 106 integrates the content obtained from the data source into a given electronic document to create a composite electronic document that includes the content obtained from the data source.
[0028] In some cases, a given electronic document can include a digital component tag or digital component script that references the digital component distribution system 110. In these cases, when the given electronic document is processed by the client device 106, the digital component tag or digital component script is executed by the client device 106. Execution of the digital component tag or digital component script configures the client device 106 to generate a request for a digital component 112 (referred to as a "component request"), which is sent to the digital component distribution system 110 via the network 102. For example, the digital component tag or digital component script can enable the client device 106 to generate a packetized data request including header and payload data. The digital component request 112 can include event data specifying characteristics, such as the name (or network location) of the server from which the media is being requested, the name (or network location) of the requesting device (e.g., the client device 106), and / or information that the digital component distribution system 110 can use to select one or more digital components to provide in response to the request. The client device 106 sends the component request 112 to a server of the digital component distribution system 110 via the network 102 (e.g., a telecommunications network).
[0029] The digital component request 112 can include event data that specifies other event characteristics, such as characteristics of the electronic document being requested and the location of the electronic document in which the digital component can be presented. For example, event data that specifies the following can be provided to the digital component distribution system 110: a reference (e.g., a uniform resource locator (URL)) to the electronic document (e.g., a web page or application) in which the digital component will be presented, an available location of the electronic document that can be used to present the digital component, the size of the available location, and / or the media type that is eligible to be presented in the location. Similarly, event data that specifies keywords associated with the electronic document ("document keywords") or entities referenced by the electronic document (e.g., people, places, or things) can also be included in the component request 112 (e.g., as payload data) and provided to the digital component distribution system 110 to facilitate identifying digital components that are eligible to be presented with the electronic document. The event data can also include a search query submitted from the client device 106 to obtain a search results page and / or data specifying search results and / or text, audible, or other visual content included in the search results.
[0030] The component request 112 can also include event data related to other information, such as information that has been provided by a user of the client device, geographic information indicating the state or region in which the component request was submitted, or other information that provides context for the environment in which the digital component will be displayed (e.g., the time of day of the component request, the day of the week of the component request, the type of device that will display the digital component, such as a mobile device or a tablet device). The component request 112 can be sent, for example, over a packetized network, and the component request 112 itself can be formatted as packetized data having a header and payload data. The header can specify a destination for the packet, and the payload data can include any of the information discussed above.
[0031] A digital component distribution system 110, including one or more digital component distribution servers, selects a digital component to be presented with a given electronic document in response to receiving a component request 112 and / or using information included in the component request 112. In some implementations, the digital component is selected in less than one second to avoid errors that may be caused by delayed selection of the digital component. For example, a delay in providing a digital component in response to the component request 112 may result in page loading errors at the client device 106, or may result in portions of the electronic document remaining unpopulated even after other portions of the electronic document are presented at the client device 106. In addition, as the delay in providing the digital component to the client device 106 increases, it is more likely that the electronic document will no longer be presented at the client device 106 when the digital component is delivered to the client device 106, thereby negatively affecting the user's experience of the electronic document. In addition, for example, if the electronic document is no longer presented at the client device 106 when the digital component is provided, the delay in providing the digital component may result in a failure in the delivery of the digital component.
[0032] To facilitate searching for electronic documents, environment 100 can include a search system 150 that identifies electronic documents by crawling and indexing them (e.g., indexing based on the content of the crawled electronic documents). Data about the electronic documents can be indexed based on the electronic documents associated with the data. Indexed and (optionally) cached copies of the electronic documents are stored in a search index 152 (e.g., a hardware memory device). Data associated with an electronic document is data representing content included in the electronic document and / or metadata about the electronic document.
[0033] The client device 106 can submit a search query to the search system 150 via the network 102. In response, the search system 150 accesses the search index 152 to identify electronic documents related to the search query. The search system 150 identifies the electronic document in the form of search results and returns the search results to the client device 106 in a search results page. Search results are data generated by the search system 150 that identify electronic documents that are responsive to (e.g., related to) a specific search query and include active links (e.g., hypertext links) that cause the client device to request data from a specified location in response to a user's interaction with the search results. An example search result can include a web page title, a portion of a text snippet or image extracted from a web page, and a URL of a web page. Another example search result can include a title of a downloadable application, a text snippet describing the downloadable application, an image depicting a user interface of the downloadable application, and / or a URL to a location from which the application can be downloaded to the client device 106. Another example search result can include a title of a streaming media, a text snippet describing the streaming media, an image depicting the content of the streaming media, and / or a URL to a location from which the streaming media can be downloaded to the client device 106. Like other electronic documents, a search results page may include one or more slots in which digital components (eg, advertisements, video clips, audio clips, images, or other digital components) can be presented.
[0034] In some implementations, digital component distribution system 110 is implemented in a distributed computing system that includes, for example, a server and a collection of multiple computing devices 114 that are interconnected and identify and distribute digital components in response to component requests 112. The collection of multiple computing devices 114 operate together to identify a collection of digital components that are eligible for presentation in an electronic document from a corpus of potentially millions of available digital components.
[0035] In some implementations, the digital component distribution system 110 implements different techniques for selecting and distributing digital components. For example, a digital component can include corresponding distribution parameters that facilitate (e.g., condition or restrict) the selection / distribution / transmission of the corresponding digital component. For example, the distribution parameters can facilitate the transmission of the digital component by requiring that a component request include at least one criterion that matches (e.g., exactly or with some pre-specified level of similarity) one of the distribution parameters of the digital component.
[0036] In another example, the distribution parameters for a particular digital component can include distribution keywords that must be matched (e.g., according to an electronic document, document keyword, or term specified in the component request 112) in order for the digital component to be eligible for presentation. The distribution parameters can also require that the component request 112 include information specifying a particular geographic region (e.g., a country or state) and / or information specifying that the component request 112 originates from a particular type of client device 106 (e.g., a mobile device or a tablet device) in order for the component item to be eligible for presentation. As discussed in more detail below, the distribution parameters can also specify a qualification value (e.g., a ranking, a score, or some other specified value) for evaluating the qualification of the component item for selection / distribution / transmission (e.g., among other available digital components). In some cases, the qualification value can be based on an amount that will be submitted when a particular event is attributed to the digital component item (e.g., presentation of the digital component).
[0037] The identification of eligible digital components can be split into a plurality of tasks 117a-117c, which are then distributed among computing devices within the set of the plurality of computing devices 114. For example, different computing devices in the set 114 can each analyze different digital components to identify various digital components having distribution parameters that match the information included in the component request 112. In some implementations, each given computing device in the set 114 can analyze different data dimensions (or sets of dimensions) and communicate (e.g., send) the results of the analysis (Res 1-Res 3) 118a-118c back to the digital component distribution system 110. For example, the results 118a-118c provided by each computing device in the set 114 can identify a subset of digital component items that are eligible for distribution in response to the component request and / or a subset of digital components having specific distribution parameters. The identification of the subset of digital components can include, for example, comparing the event data with the distribution parameters, and identifying a subset of digital components having distribution parameters that match at least some features of the event data.
[0038] The digital component distribution system 110 aggregates the results 118a-118c received from the set of multiple computing devices 114 and uses information associated with the aggregated results to select one or more digital components to be provided in response to the component request 112. For example, the digital component distribution system 110 can select a set of winning digital components (one or more digital components) based on the results of one or more digital component evaluation processes. In turn, the digital component distribution system 110 can generate and send response data 120 (e.g., digital data representing the response) over the network 102, the response data 120 enabling the client device 106 to integrate the set of winning digital components into a given electronic document so that the set of winning digital components and the content of the electronic document are presented together at a display of the client device 106.
[0039] In some implementations, the client device 106 executes instructions included in the response data 120 that configure and enable the client device 106 to obtain a set of winning digital components from one or more digital component servers 108. For example, the instructions in the response data 120 can include a network location (e.g., a URL) and a script that causes the client device 106 to send a server request (SR) 121 to the digital component server 108 to obtain a given winning digital component from the digital component server 108. In response to the server request 121, the digital component server 108 will identify the given winning digital component specified in the server request 121 and send to the client device 106 digital component data 122 (DI data) that presents the given winning digital component in an electronic document at the client device 106.
[0040] In some implementations, the distribution parameters for the distribution of the digital component may include user characteristics, such as demographic information, user interests, and / or other information that can be used to personalize the user's online experience. In some cases, these characteristics and / or information about the user of the client device 106 are readily available. For example, a content platform such as the content server 104 or the search system 150 may allow a user to register with the content platform by providing such user information. In another example, the content platform can use cookies to identify the client device, which can store information about the user's online activities and / or user characteristics.
[0041] In the effort to protect user privacy, these and other methods of identifying user characteristics are becoming less popular. In order to protect user privacy, users can be assigned to one or more user groups based on the digital content that users visit during a single browsing session. For example, when a user visits a specific website and interacts with a specific item presented on the website or adds a commodity to a virtual shopping cart, the user can be assigned to a group of users who have visited the same website or other websites with similar contexts, or are interested in the same commodity. For example, if a user of client device 106 searches for shoes and visits multiple web pages of different shoe manufacturers, the user can be assigned to user group "shoes", which can include identifiers of all users who have visited websites related to shoes. Therefore, user groups can represent the interests of users who gather, without identifying individual users and without making any individual users identifiable. For example, user groups can be identified by user group identifiers for each user in the group. As an example, if a user adds shoes to a shopping cart of an online retailer, the user can be added to a shoe user group with a specific identifier, and the specific identifier is assigned to each user in the group.
[0042] In some implementations, for example, through a browser-based application, a user's group membership can be maintained at the user's client device 106, rather than by the digital component provider or by the content platform, or by another party. A user group can be specified by a corresponding user group identifier. The user group identifier of a user group can describe the group (e.g., a gardening group) or be a code (e.g., a non-descriptive alphanumeric sequence) representing the group.
[0043] In the event that user characteristics are not available (e.g., beyond the group identifier), the digital component distribution system 110 can include a user evaluation device 170 that predicts user characteristics based on the available information. In some implementations, the user evaluation device 170 implements one or more machine learning models that predict one or more user characteristics based on information included in the component request 112 (e.g., the group identifier).
[0044] For example, if a user of client device 106 uses browser-based application 107 to load a website that includes one or more digital component slots, browser-based application 107 can generate and send component requests 112 for each of the one or more digital component slots. Component request 112 includes a user group identifier (or multiple) corresponding to the user group (or multiple), the user group identifier including an identifier of client device 106, other information such as geographic information indicating the state or region in which component request 112 is submitted, or other information that provides context for the environment in which digital component 112 will be displayed (e.g., time of day of component request, day of week of component request, type of client device 106 that will display the digital component, such as a mobile device or tablet device, etc.).
[0045] After receiving component request 112, digital component distribution system 110 provides the information included in component request 112 as input to the machine learning model. After processing the input, the machine learning model generates an output including a prediction of one or more user characteristics not included in component request 112. These one or more user characteristics along with other information included in the component request can be used to retrieve the digital component from digital component server 108. Figure 2 Further interpretation is given to generate the predicted output of user features.
[0046] Figure 2is a block diagram of an example machine learning model implemented within the user evaluation device 170. In general, the machine learning model can be any technology that is considered suitable for a particular implementation, such as an artificial neural network (ANN), a support vector machine (SVM), a random forest (RF), etc., which includes multiple trainable parameters. During the training process, multiple training parameters are adjusted while iterating over multiple samples of the training data set (a process called optimization) based on the error measure generated by the loss function. The loss function compares the predicted attribute value of the machine learning model with the true value of the sample in the training set to generate an error measure, wherein the model is trained to minimize the error measure.
[0047] In some implementations, a machine learning model may include multiple sub-machine learning models (also referred to as "sub-models") such that each sub-model predicts a specific user characteristic (e.g., user demographics, user interests, or some other characteristic). For example, the user evaluation device 170 includes three sub-models: (i) characteristic 1 model 220, (ii) characteristic 2 model 230, and (iii) characteristic 3 model 240. Each of these sub-models predicts the likelihood that a user has a different characteristic (e.g., demographics or user interests). Other implementations may include more or fewer separate sub-models to predict a system (or administrator) defined number of user characteristics.
[0048] In some implementations, the machine learning model can accept as input information included in the component request 112. As previously described, the component request 112 can include a user group identifier corresponding to a user group that includes the client device 106 and / or other information, such as geographic information indicating the state or region in which the component request 112 was submitted, or other information that provides context for the environment in which the digital component 112 will be displayed (e.g., the time of day of the component request, the day of the week of the component request, the type of client device 106 that will display the digital component, such as a mobile device or tablet device). For example, the input 205 includes information included in the component request 112, namely, the user group identifier (user group ID) 202 and other information 204.
[0049] In some implementations, the machine learning model (or sub-model) implemented within the user evaluation apparatus 170 can be further extended to accept additional inputs 210 related to the user's current online activity. For example, the additional inputs can include a list of websites previously visited by the user in the current session, prior interactions with digital components presented in previously visited websites, or predictions of other sub-models. The additional inputs can also include information related to the user group of which the user of the client device 106 is a member. For example, similarity measures of web content accessed by users in the same user group, average predictions of user characteristics of users belonging to the same group, distribution of attributes across users in the same group, etc.
[0050] Depending on the particular implementation, the machine learning model (or each sub-model) can use one or more input features to generate an output including a prediction of one or more user characteristics. For example, feature 1 model 220 can predict predicted feature 1 252 (e.g., predicted gender) of the user of client device 106 as an output. Similarly, feature 2 model 230 can be a regression model that processes inputs 205 and 210 and generates predicted feature 2 254 (e.g., predicted age range) of the user of client device 106 as an output. In the same manner, feature model 3 240 generates predicted feature 3 256 of the user as an output.
[0051] These predicted user characteristics are used together with the input features 205 and 210 to select digital components provided by the digital component provider and / or server 108. However, the implementation of the machine learning model (or each sub-model) by the user evaluation device 170 to predict user characteristics may suffer from data class imbalance. Data class imbalance is a problem of predictive classification that occurs when there is an uneven distribution of categories in the training data set. Since the samples of the training data set are obtained from the real world, the distribution of the samples in the training data set will reflect the distribution of the categories according to the sampling process used to collect the data, and may not evenly represent all categories. In other words, within a single user group (e.g., a user group including users interested in "shoes"), there will be user subgroups with different characteristics, also referred to as different user groups (e.g., a subgroup of male users interested in shoes and a second subgroup of female users interested in shoes), and when there are different numbers of users in each subgroup in the training sample (e.g., more females than males, and vice versa), imbalance may result.
[0052] Training a machine learning model on such disproportionate data results in a tendency for the model to achieve better predictions for over-represented groups, due to the fact that the machine learning model is class-indifferent. For example, assume that the number of male users associated with the user group "cosmetics" is less than the number of female users associated with the same user group. In another example, assume that the number of users in the 19-29 age group associated with the user group "video games" is significantly greater than the number of users in the 59-69 age group. In such a scenario, the machine learning model implemented in the user evaluation device 170 is trained to predict user characteristics without considering the distribution and / or proportion of the training samples, which may lead to degradation of model performance for user groups that are under-represented in the training data set. In other words, the predictions output by the model may be skewed toward the characteristics of the user subgroup with the most members in the training data set. For example, assume that the ratio of the true number of male users in the user group "cosmetics" to the true number of female users in the same user group is 1 / 10, that is, 10% of the users in the user group "cosmetics" are male. A machine learning model trained on a training dataset that shows the same distribution of males to females as the true ratio can classify all users in a user group as females and still achieve 90% accuracy.
[0053] Training a machine learning model on disproportionate data may also result in lower performance across different user groups. For example, assume that a machine learning model implemented within user evaluation apparatus 170 is trained to predict user characteristics for users belonging to user groups 1 and 2. Also assume that 95% of the training samples represent users associated with user group 1. In such a scenario, the machine learning model can predict user characteristics for users associated with user group 2 while exhibiting a bias toward user characteristics that are generally common to user group 1.
[0054] To address the problem of class imbalance, a machine learning model (or sub-model) is trained to optimize a loss function that is modified by adding additional terms to the loss function, which reduces the difference in performance measurement across different user groups that are all represented by the same group identifier. In other words, the additional terms in the loss function help reduce the number of false positives and / or false negatives output by the model due to insufficient representation of any user subgroup within a particular user group. For example, the additional terms added to the modified loss function penalize the machine learning model to update its learnable parameters during the training process, so that even if the training data set has a disproportionate number of male and female users and / or samples, the model is able to distinguish between male and female users in the user group "cosmetics", thereby providing a more accurate output. In a scenario where the loss function is modified by adding additional terms to provide consistent performance across different user groups, the additional terms measure the variability of model performance within each user group, thereby prompting the model to achieve more consistent performance within each user group.
[0055] In some implementations, the additional term added to the loss function is an error regularization term that penalizes the machine learning model according to its error across multiple subgroups within a particular user group and / or across multiple user groups during the training process, so that the machine learning model generalizes over the distribution of the training data set. In other implementations, the additional term added to the loss function is a divergence minimization term that penalizes the machine learning model if there are differences between the distributions of the predicted outputs of the machine learning model across different groups. Error regularization terms and divergence minimization terms are further explained below.
[0056] In some implementations, the additional term added to the loss function is a loss variance term. In some implementations, the loss variance term is the squared difference between the loss of the first user in the first subgroup and the average loss of all users in all subgroups within the user group. For example, such an implementation would calculate the variance of the average loss across different interest groups (such as "makeup" and "shoes"). In such an example, the loss variance term takes into account the variability of the loss across different interest groups (e.g., "makeup" and "shoes"). When the variance across different interest groups is higher, the term will take a higher value, thereby encouraging more uniform performance across subgroups (measured by lower variance). The loss variance term is as follows:
[0057] L(X)+Var(L)
[0058] in,
[0059] Var(L)=E x~X ((L(x)-E[L(x)]) 2 )
[0060] And, where L(X) is the loss function across the data distribution X. Other variations of this loss variance formula are also possible, where the variance is conditioned on the true label of each observation or user attribute.
[0061] In some implementations, the additional term added to the loss function is a maximum weighted loss difference term. In some implementations, the maximum weighted loss difference term is the maximum weighted difference in user loss between all users in a subgroup within a particular user group and all users in all subgroups within the particular user group. In some implementations, the subgroups can be weighted based on importance within the user group. The maximum weighted loss difference term is as follows:
[0062] L(X)+w(g)|E[(g(x)=1)-E(L(x))]|
[0063] in,
[0064] g(x) is a function indicating that a specific element is in the subgroup g, and w(g) indicates a weighting function of the group g.
[0065] In some implementations, the additional term added to the loss function is a coarse loss variance term. In some implementations, the coarse loss variance term is the average of the squared differences between the losses of subgroups within the user group. The coarse loss variance term is as follows:
[0066] L(X)+Var(E x~X [L(x)|A])
[0067] in,
[0068] Var(E x~X [A])=E x~X [(A)-E(L(x)) 2 ]
[0069] Where A is a subgroup of the user group.
[0070] In some implementations, the additional term added to the loss function is a Hilbert-Schmidt independence criterion (HSIC) regularization term. In some implementations, HSIC measures the degree of statistical independence between the distribution of predictions of user attributes and the members of examples in a certain group (e.g., interest group, demographic group). The HSIC regularization term is as follows:
[0071] HSIC(p xy ,F,G)=||C xy ||HS 2
[0072] Where F and G are the joint measurements p xy The reproducing kernel Hilbert space, C xy is the cross-covariance operator, and ||·|| HS is the Hilbert-Schmidt matrix norm (in this example, X and Y would be the model predictions and group memberships for a set of examples, respectively).
[0073] In some implementations, the additional term added to the loss function is a mutual information (MI) term. In some implementations, the mutual information term is a measure of the similarity of the distribution of the model's predictions over different groups (such as interest groups) (e.g., the MI term can measure the degree of similarity between the overall distributions of the model's predictions for users in the shoe interest group and the cosmetics interest group). The mutual information is shown below:
[0074] L(X)+MI(logits T ,membership T )
[0075] In some implementations, the additional term added to the loss function is a Kullback-Leibler divergence term. In some implementations, the Kullback-Leibler divergence term is a measure of the difference in the distribution of machine learning model predictions across multiple subgroups of users belonging to the same user group. In some implementations, the Kullback-Leibler divergence term can be expressed as:
[0076] L(X)+XL(logits A ||logits B )
[0077] Here, the function KL is the continuous Kullback-Leibler divergence of the two subgroups A and B within the user group, and the function KL of the two distributions p and q can be written as KL(p||q)=∑ x p(x)log(p(x) / q(x)).
[0078] After modifying the loss function for each of the sub-models 220, 230, and 240, the sub-model is trained on the training data set. Depending on the specific implementation, the training process of the sub-machine learning model can be supervised, unsupervised, or semi-supervised, and may also include adjusting multiple hyperparameters associated with the model (a process known as hyperparameter tuning).
[0079] Once the machine learning model (or sub-model) is trained, the digital component distribution system 110 is able to select digital components based on one or more user characteristics predicted by the user evaluation device 170 (or the machine learning model implemented in the user evaluation device 170). For example, assume that a male user belonging to the subgroup "cosmetics" provides a search query "face cream" through the client device 106 to obtain a search result page and / or data specifying the search result and / or text, audible content or other visual content related to the search query. Assume that the search result page includes a slot for a digital component provided by an entity other than the entity that generates and provides the search result page. The browser-based application 107 running on the client device 106 generates a component request 112 for the digital component slot. After receiving the component request 112, the digital component distribution system 110 provides the information included in the component request 112 as input to the machine learning model implemented by the user evaluation device 170. The machine learning model generates a prediction of one or more user characteristics as output after processing the input. For example, the sub-machine learning model 220 correctly predicts the user of the client device 106 as a male based on the parameters learned by optimizing the modified loss function. Thus, the digital component provider 110 is able to select a digital component associated with a facial cream designated for distribution to men. After selection, the selected digital component is sent to the client device 106 to be presented in the search results page with the search results.
[0080] Figure 3 is a flow chart of an example process 300 for distributing digital components based on a modified loss function. The operation of process 300 is described below as Figure 1 and Figure 2 The operations of process 300 are performed by components of the system described and depicted in the description. The operations of process 300 are described below for illustration purposes only. The operations of process 300 can be performed by any suitable device or system (e.g., any suitable data processing device). The operations of process 300 can also be implemented as instructions stored on a computer-readable medium that can be non-transitory. The execution of the instructions causes one or more data processing devices to perform the operations of process 300.
[0081] A loss function is identified that generates a loss that represents a measure of performance that the model seeks to optimize during training (310). In some implementations, such as reference Figure 1 and Figure 2As described, the digital component distribution system 110 can include a user evaluation apparatus 170 that can implement a machine learning model (or multiple machine learning models referred to as sub-machine learning models) to predict user characteristics based on information included in the component request 112. For example, the component request 112 includes a user group identifier corresponding to a user group associated with the client device 106, in addition to other information such as geographic information indicating the country or region in which the component request 112 is submitted, or other information providing context for the environment in which the digital component 112 will be displayed (e.g., time of day of the component request, day of the week of the component request, type of client device 106 that will display the digital component, such as a mobile device or tablet device).
[0082] The loss function is modified (320). In some implementations, the loss function is modified by adding an additional term to the loss function, the additional term reducing the difference in performance measurement across different user groups all represented by the same group identifier. To address the class imbalance problem, the machine learning model (or sub-machine learning model) is trained to optimize the loss function modified by adding an additional term to the loss function, which reduces the difference in performance measurement across different user groups all represented by the same group identifier. The relative weighting between the two loss terms (standard loss and additional term) can be adjusted by applying a multiplier to one or both terms, respectively.
[0083] In some implementations, the additional term added to the loss function is an error regularization term that penalizes the machine learning model for each incorrect prediction during the training process across multiple subgroups within a particular user group so that the machine learning model generalizes over the distribution of the training data set. In other implementations, the additional term added to the loss function is a divergence minimization term that penalizes the machine learning model if there is a difference between the distribution of the predicted output of the machine learning model and the correct annotation (ground truth). The error regularization term and the divergence minimization term are further explained below.
[0084] For example, the additional term can be a loss variance term, which is the square difference between the loss of the first user in the first subgroup and the average loss of all users in all subgroups within the user group. In another example, the additional term can be a maximum weighted loss difference term, which is the maximum weighted difference between the user losses of all users in a subgroup within a specific user group and all users in all subgroups within a specific user group. In another example, the additional term can be a rough loss variance term, which is the average of the squared differences between the losses of the subgroups within the user group. In another example, the additional term can be a HSIC regularization term, which is a loss difference across different user subgroups in a non-parametric manner and independent of the distribution of users in different user subgroups. Similarly, other additional terms can be mutual information terms and Kullback-Leibler divergence terms, which are measures of the differences in the distribution of multiple user subgroups belonging to the same user group predicted by the machine learning model.
[0085] The model is trained using the modified loss function (330). For example, after modifying the loss function of each of the sub-machine learning models 220, 230, and 240, the sub-machine learning model is trained on the training data set. Depending on the specific implementation, the training process of the sub-machine learning model can be supervised, unsupervised, or semi-supervised, and can also include adjusting multiple hyperparameters associated with the model (a process known as hyperparameter tuning).
[0086] A request for a digital component is received (340). In some implementations, the request includes a given group identifier for a particular group of different user groups. For example, if a user of client device 106 uses browser-based application 107 to load a website that includes one or more digital component slots, browser-based application 107 can generate and send component requests 112 for each of the one or more digital component slots, which can be received by digital component distribution system 110.
[0087] The trained model is applied to the information included in the request to generate one or more user characteristics not included in the request (350). For example, after receiving the component request 112, the digital component distribution system 110 provides the information included in the component request 112 as input to the machine learning model implemented by the user evaluation device 170. After processing the input, the machine learning model generates a prediction of one or more user characteristics as an output.
[0088] One or more digital components are selected based on one or more user characteristics generated by the trained model (360). For example, assume that a male user belonging to the subgroup "cosmetics" provides a search query "face cream" through the client device 106 to obtain a search result page and / or data specifying the search result and / or text, audible content, or other visual content related to the search query. Assume that the search result page includes a slot for a digital component. The browser-based application 107 running on the client device 106 generates a component request 112 for the digital component slot. After receiving the component request 112, the digital component distribution system 110 provides the information included in the component request 112 as input to the machine learning model implemented by the user evaluation device 170. After processing the input, the machine learning model generates a prediction of one or more user characteristics as an output. For example, the sub-machine learning model 220 correctly predicts the user of the client device 106 as a male based on the parameters learned by optimizing the modified loss function. Therefore, the digital component provider 110 is able to select a digital component related to a face cream having a distribution criterion indicating that the digital component should be distributed to men.
[0089] The selected one or more digital components are sent to the client device (370). For example, after the digital components are selected by the digital component distribution system 110 based on the predicted user characteristics, the selected digital components are sent to the client device 106 for presentation.
[0090] In addition to the above description, controls may be provided to the user that allow the user to select whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., information about the user's social network, social actions or activities, occupation, the user's preferences, or the user's current location) and whether to send content or communications from a server to the user. In addition, specific data may be processed in one or more ways before it is stored or used so that personally identifiable information is removed. For example, the user's identity may be processed so that personally identifiable information cannot be determined for the user, or the user's geographic location may be generalized (such as to a city, zip code, or state level) when location information is obtained so that the user's specific location cannot be determined. Thus, the user can control what information is collected about the user, how the information is used, and what information is provided to the user.
[0091] Figure 44 is a block diagram of an example computer system 400 that can be used to perform the above operations. System 400 includes a processor 410, a memory 420, a storage device 430, and an input / output device 440. Each of components 410, 420, 430, and 440 can be interconnected, for example, using a system bus 450. Processor 410 is capable of processing instructions for execution within system 400. In one implementation, processor 410 is a single-threaded processor. In another implementation, processor 410 is a multi-threaded processor. Processor 410 is capable of processing instructions stored in memory 420 or on storage device 430.
[0092] The memory 420 stores information within the system 400. In one implementation, the memory 420 is a computer readable medium. In one implementation, the memory 420 is a volatile memory unit. In another implementation, the memory 420 is a non-volatile memory unit.
[0093] The storage device 430 can provide mass storage for the system 400. In one implementation, the storage device 430 is a computer-readable medium. In various implementations, the storage device 430 can include, for example, a hard disk device, an optical disk device, a storage device shared by multiple computing devices over a network (e.g., a cloud storage device), or some other mass storage device.
[0094] The input / output device 440 provides input / output operations for the system 400. In one implementation, the input / output device 440 can include one or more of a network interface device (e.g., an Ethernet card), a serial communication device (e.g., an RS-232 port), and / or a wireless interface device (e.g., an 802.11 card). In another implementation, the input / output device can include a driver device that is configured to receive input data and send output data to other input / output devices, such as a keyboard, a printer, and a display device 370. However, other implementations can also be used, such as mobile computing devices, mobile communication devices, set-top TV client devices, etc.
[0095] Although already Figure 4 An example processing system is described in the specification, but the subject matter and implementation of the functional operations described in this specification can be implemented in other types of digital electronic circuits, or in computer software, firmware or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them.
[0096] An electronic document (referred to simply as a document for brevity) does not necessarily correspond to a file. A document may be stored as part of a file that holds other documents, in a single file dedicated to the document in question, or in multiple collaborative files.
[0097] The subject matter and the embodiments of the operation described in this specification can be implemented in digital electronic circuits, or in computer software, firmware or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more thereof. The embodiments of the subject matter described in this specification can be implemented as one or more computer programs, that is, one or more modules of computer program instructions, which are encoded on a computer storage medium for execution by a data processing device or control of the operation of the data processing device. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagation signal, for example, a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device for execution by a data processing device. The computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more thereof, or is included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more thereof. In addition, although the computer storage medium is not a propagation signal, the computer storage medium can be the source or destination of the computer program instructions encoded in the artificially generated propagation signal. The computer storage media may also be, or be included in, one or more separate physical components or media (eg, multiple CDs, disks, or other storage devices).
[0098] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0099] The term "data processing apparatus" includes all types of apparatus, devices and machines for processing data, including, for example, a programmable processor, a computer, a system on a chip, or a plurality or combination of the foregoing. The apparatus can include dedicated logic circuits, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the apparatus can also include code that creates an execution environment for the computer program in question, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.
[0100] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed on multiple sites and interconnected by a communication network.
[0101] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuits (e.g., FPGAs (field programmable gate arrays) or ASICs (application-specific integrated circuits)), and the apparatus can also be implemented as special purpose logic circuits (e.g., FPGAs (field programmable gate arrays) or ASICs (application-specific integrated circuits)).
[0102] As an example, processors suitable for executing computer programs include both general-purpose microprocessors and special-purpose microprocessors. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as a magnetic disk, a magneto-optical disk, or an optical disk, or operably coupled to receive data from it or transfer data to it or both. However, a computer does not need to have such a device. In addition, a computer can be embedded in another device, such as, to name a few, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive). Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0103] To provide interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, voice, or tactile input. In addition, the computer can interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's client device in response to a request received from the web browser.
[0104] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification), or includes any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), interconnected networks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0105] A computing system can include a client and a server. The client and the server are usually remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by means of computer programs running on each computer and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to the client device (e.g., for the purpose of displaying data to a user interacting with the client device and receiving user input from the user). Data generated at the client device (e.g., the result of a user interaction) can be received from the client device at the server.
[0106] Although this specification contains many specific implementation details, these should not be interpreted as limitations on any invention or the scope that may be claimed, but as descriptions of features specific to a particular embodiment of a particular invention. Specific features described in the context of a separate embodiment in this specification can also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination. In addition, although the features may be described above as acting in a specific combination and even initially claimed as such, one or more features from the claimed combination can be deleted from the combination in some cases, and the claimed combination may involve a sub-combination or a variant of a sub-combination.
[0107] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that the operations be performed in the particular order shown or in sequence, or that all of the operations shown be performed, to achieve the desired results. In certain circumstances, multitasking and parallel processing may be advantageous. In addition, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0108] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.
Claims
1. A computer-implemented method comprising: For a model to be trained, identifying a loss function that generates a loss representing a performance measure that the model seeks to optimize during training; modifying a loss function, including adding an additional term to the loss function, the additional term measuring variability of model performance within a user group, the user group including different user subgroups all represented by the same user group identifier, wherein each user subgroup of the different user subgroups has characteristics that are different from characteristics of other subgroups of the different user subgroups; Training the model using a modified loss function; receiving a request for a digital component from a client device, the request including a given user group identifier for a particular user group among the different user groups; generating one or more user features not included in the request by applying the trained model to information included in the request; selecting one or more digital components based on the one or more user characteristics generated by the trained model; and The selected one or more digital components are sent to the client device.
2. The computer-implemented method of claim 1, wherein: Modifying the loss function involves adding an error regularization term to the loss function.
3. The computer-implemented method of claim 2, wherein: Adding an error regularization term to the loss function includes adding a loss variance term to the loss function, wherein the loss variance function is the squared difference between an average loss of the model within the subgroup and an average loss of the model across all users, and wherein the difference is calculated separately across users based on different attributes.
4. The computer-implemented method of claim 2, wherein: Adding the error regularization term to the loss function includes adding a maximum weighted loss difference term to the loss function, wherein the maximum weighted loss difference term is the maximum weighted difference between the loss in the user group and the loss for all users in all different user groups, and wherein the subgroups are weighted based on importance within the user group.
5. The computer-implemented method of claim 2, wherein: Adding an error regularization term to the loss function includes adding a coarse loss variance term to the loss function, wherein the coarse loss variance term is an average of squared differences between losses of subgroups within a user group conditioned on individual user attributes.
6. The computer-implemented method of claim 2, wherein: Adding an error regularization term to the loss function includes adding a HSIC regularization term to the loss function, wherein the HSIC term characterizes the difference in loss across different user groups in a non-parametric manner and independently of the distribution of users in the different user groups.
7. The computer-implemented method of claim 1, wherein: Modifying the loss function involves adding a divergence minimization term to the loss function.
8. The computer-implemented method of claim 7, wherein: Adding a divergence minimization term to the loss function includes adding one of a mutual information term or a Kullback-Leibler divergence term to the loss function, wherein the mutual information term characterizes the similarity of distributions of model predictions across multiple user subgroups, and wherein the Kullback-Leibler divergence term characterizes the difference in distributions of model predictions across multiple user subgroups.
9. A system for distributing a digital component, comprising: For a model to be trained, identifying a loss function that generates a loss representing a performance measure that the model seeks to optimize during training; modifying the loss function, including adding an additional term to the loss function, the additional term measuring variability of model performance within a user group, the user group including different user subgroups all represented by the same user group identifier, wherein each user subgroup of the different user subgroups has characteristics that are different from characteristics of other subgroups of the different user subgroups; Training the model using a modified loss function; receiving a request for a digital component from a client device, the request including a given user group identifier for a particular user group among the different user groups; generating one or more user features not included in the request by applying the trained model to information included in the request; selecting one or more digital components based on the one or more user characteristics generated by the trained model; and The selected one or more digital components are sent to the client device.
10. The system of claim 9, wherein: Modifying the loss function involves adding an error regularization term to the loss function.
11. The system of claim 10, wherein: Adding an error regularization term to the loss function includes adding a loss variance term to the loss function, wherein the loss variance function is the squared difference between the average loss of the model within the subgroup and the average loss of the model across all users, and wherein the difference is calculated separately across users based on different attributes.
12. The system of claim 10, wherein: Adding the error regularization term to the loss function includes adding a maximum weighted loss difference term to the loss function, wherein the maximum weighted loss difference term is the maximum weighted difference between the loss in the user group and the loss for all users in all different user groups, and wherein the subgroups are weighted based on importance within the user group.
13. The system of claim 10, wherein: Adding an error regularization term to the loss function includes adding a coarse loss variance term to the loss function, wherein the coarse loss variance term is an average of squared differences between losses of subgroups within a user group conditioned on individual user attributes.
14. The system of claim 10, wherein: Adding an error regularization term to the loss function includes adding a HSIC regularization term to the loss function, wherein the HSIC term characterizes the difference in loss across different user groups in a non-parametric manner and independently of the distribution of users in the different user groups.
15. The system of claim 9, wherein: Modifying the loss function involves adding a divergence minimization term to the loss function.
16. The system of claim 15, wherein: Adding a divergence minimization term to the loss function includes adding one of a mutual information term or a Kullback-Leibler divergence term to the loss function, wherein the mutual information term characterizes the similarity of distributions of model predictions across multiple user subgroups, and wherein the Kullback-Leibler divergence term characterizes the difference in distributions of model predictions across multiple user subgroups.
17. A non-transitory computer-readable medium storing instructions that, when executed by one or more data processing devices, cause the one or more data processing devices to perform operations comprising: For a model to be trained, identifying a loss function that generates a loss representing a performance measure that the model seeks to optimize during training; modifying the loss function, including adding an additional term to the loss function, the additional term measuring variability of model performance within a user group, the user group including different user subgroups all represented by the same user group identifier, wherein each user subgroup of the different user subgroups has characteristics that are different from characteristics of other subgroups of the different user subgroups; Training the model using a modified loss function; receiving a request for a digital component from a client device, the request including a given user group identifier for a particular user group among the different user groups; generating one or more user features not included in the request by applying the trained model to information included in the request; selecting one or more digital components based on the one or more user characteristics generated by the trained model; and The selected one or more digital components are sent to the client device.
18. The non-transitory computer readable medium of claim 17, wherein: Modifying the loss function involves adding an error regularization term to the loss function.
19. The non-transitory computer readable medium of claim 18, wherein: Adding an error regularization term to the loss function includes adding a loss variance term to the loss function, wherein the loss variance function is the squared difference between the average loss of the model within the subgroup and the average loss of the model across all users, wherein the difference is calculated separately across users based on different attributes.
20. The non-transitory computer readable medium of claim 18, wherein: Adding the error regularization term to the loss function includes adding a maximum weighted loss difference term to the loss function, wherein the maximum weighted loss difference term is the maximum weighted difference between the loss in the user group and the loss for all users in all different user groups, and wherein the subgroups are weighted based on importance within the user group.
21. The non-transitory computer readable medium of claim 18, wherein: Adding an error regularization term to the loss function includes adding a coarse loss variance term to the loss function, wherein the coarse loss variance term is an average of squared differences between losses of subgroups within a user group conditioned on individual user attributes.
22. The non-transitory computer readable medium of claim 18, wherein: Adding an error regularization term to the loss function includes adding a HSIC regularization term to the loss function, wherein the HSIC term characterizes the difference in loss across different user groups in a non-parametric manner and independently of the distribution of users in the different user groups.
23. The non-transitory computer readable medium of claim 17, wherein: Modifying the loss function involves adding a divergence minimization term to the loss function.
24. The non-transitory computer readable medium of claim 21, wherein: Adding a divergence minimization term to the loss function includes adding one of a mutual information term or a Kullback-Leibler divergence term to the loss function, wherein the mutual information term characterizes the similarity of distributions of model predictions across multiple user subgroups, and wherein the Kullback-Leibler divergence term characterizes the difference in distributions of model predictions across multiple user subgroups.
Citation Information
Patent Citations
Determining a number of cluster groups associated with content identifying users eligible to receive the content
US20160232575A1