Privacy-preserving machine learning for content distribution and analytics

By training machine learning models using secure multi-party computation technology and maintaining user profiles on client devices, the privacy leakage problem caused by third-party cookies is solved. This enables user group expansion and privacy protection without relying on third-party cookies, improving data transmission security and reporting accuracy.

CN115335825BActive Publication Date: 2026-03-06GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180025475.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-03
Publication Date
2026-03-06
Estimated Expiration
2041-02-03

AI Technical Summary

Technical Problem

Existing technologies rely on third-party cookies in the distribution of demographic-based digital components, leading to user privacy leaks and insecure data transmission. Furthermore, browsers lose functionality when blocking the use of third-party cookies, making it impossible to effectively expand user groups.

Method used

Secure multi-party computation (MPC) technology is used to train machine learning models, user profiles are maintained on client devices, and secure data processing is performed through an MPC cluster to generate user group identifiers, avoiding plaintext data transmission and achieving user group expansion and privacy protection.

Benefits of technology

Without using third-party cookies, accurately identify user interests and expand user groups, protect user privacy, improve data transmission security and storage efficiency, and generate more accurate demographic-based user group reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115335825B_ABST
    Figure CN115335825B_ABST
Patent Text Reader

Abstract

This disclosure relates to systems and techniques that can be implemented by a content platform to optimize: (a) demographic-based digital component distribution that appropriately targets a user to classify each user into a specific demographic for the purpose of maximizing the effectiveness of the digital component shown to that user; and (b) demographic reporting that reports the effectiveness of the digital component to the digital component provider.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to a privacy-preserving machine learning platform that uses secure multi-party computation to train and use machine learning models. Background Technology

[0002] Some machine learning models are trained on data collected from multiple sources (e.g., across multiple websites and / or native applications). However, this data may include private or sensitive data that should not be shared or disclosed to other parties. Summary of the Invention

[0003] Digital component providers often benefit from the ability to limit the audience to which their digital components are shown. Such limitations are beneficial to users, as they are shown more relevant content. These limitations can be based on audience demographics. Content platforms give digital component providers the ability to distribute digital components of their activities to specific user groups (e.g., users in a specific demographic group). Content platforms often do this by associating a user's cookie (e.g., a third-party cookie) with one of the demographic groups, which can be done even when the user is not logged into the website displaying the digital component. Content platforms can facilitate the display of digital components on the content provider's website or other websites (e.g., third-party websites hosted by entities other than the content provider's platform). To keep track of user browsing preferences, content platforms use the same cookie to track a user's browsing history across different websites. The content platform uses the data extracted from this cookie to (a) categorize each user into a specific demographic so as to appropriately target that user for the purpose of maximizing the effectiveness of the digital component shown to that user, and (b) report the effectiveness of the digital component to the digital component provider. However, the use of cookies is disadvantageous because other websites can access them, and thus user preferences across different websites can be accessed by each of those websites. Such access to user behavior across many websites can sometimes be considered an invasion of user privacy. To avoid this intrusion, there is a need to optimize the distribution and reporting of demographic-based digital components performed by content platforms without using such cookies. Furthermore, some browsers block the use of third-party cookies, making browsing information typically collected across multiple websites unavailable. Therefore, any functionality relying on information collected through third-party cookies (e.g., cookies from domains different from the domain of the web page the user is currently viewing), such as regulating the distribution of data to client devices, is no longer available.

[0004] This disclosure relates to a privacy-preserving machine learning platform that uses secure multi-party computation to train and utilize machine learning models. In particular, this disclosure describes systems and techniques for classifying users into specific demographic groups without using third-party cookies and for reporting the effectiveness of activities involving the distribution of digital components to user groups that are at least partially based on demographic information.

[0005] In one aspect, an application on a client device is capable of receiving data from one or more computers that identifies inferred demographic characteristics of the application's user. The application is capable of displaying digital content, which includes computer-readable code for reporting events related to the digital content, data specifying a set of permitted demographic-based user group identifiers, and activity identifiers for the digital content. The application is capable of determining whether a given inferred characteristic matches a given permitted demographic-based user group identifier. In response to determining that a given inferred demographic characteristic matches a given permitted demographic-based user group identifier, computer-readable code is capable of generating and sending a request to update one or more event counts for the digital content and the given permitted demographic-based user group identifiers.

[0006] In some implementations, one or more of the following can be additionally or alternatively implemented in any suitable combination. Receiving data identifying the inferred demographic characteristics of a user of the application can include receiving, at the client device, an inferred demographic user group identifier for which the user is added. Sending an inference request, including the user's profile, to one or more computers is possible. Receiving the inferred demographic user group identifier from one or more computers in response to the inference request is possible. The one or more computers can include multiple multi-party computation (MPC) servers. Sending the inference request to one or more computers can include sending a corresponding secret share of the user profile to each of the multiple MPC servers. The multiple MPC servers can use one or more machine learning models to perform a secure MPC process to generate a secret share of the inferred demographic user group identifier and transmit the secret share of the inferred demographic user group identifier to the application.

[0007] An application can send an inference request to one or more computers for inferred demographic characteristics of a user of the application. The inference request can include a user profile and a context signal associated with at least one of: (i) a digital content slot in which digital content is displayed; or (ii) the digital content, wherein an inferred demographic user group identifier is received from one or more computers in response to the inference request; or (iii) a Uniform Resource Locator (URL) of the resource containing the digital content, the location of the client device, and the application's spoken language setting. Generating and sending a request using computer-readable code to update one or more event counts for the digital content and a given permitted demographic user group identifier can include: generating an aggregation key including an activity identifier and a given permitted demographic user group identifier; and transmitting the aggregation key along with the request. Generating and sending a request using computer-readable code to update one or more event counts for the digital content and a given permitted demographic user group identifier can include invoking the application's application programming interface (API) to send the request.

[0008] Methods, systems, apparatuses, computer-programmable products, etc., that implement the features described above are also within the scope of this disclosure. For example, in some aspects, a system is described comprising: at least one programmable processor; and a machine-readable medium storing instructions that, when executed by the at least one programmable processor, cause the at least one programmable processor to perform operations including: receiving, by an application of a client device, data from one or more computers identifying inferred demographic characteristics of a user of the application; displaying, by the application, digital content including computer-readable code for reporting events related to the digital content, data specifying a set of permitted demographic-based user group identifiers, and an activity identifier for the digital content; determining, by the application, that a given inferred characteristic matches a given permitted demographic-based user group identifier; and, in response to determining that the given inferred demographic characteristic matches the given permitted demographic-based user group identifier, using computer-readable code to generate and send a request to update one or more event counts for the digital content and the given permitted demographic-based user group identifier.

[0009] Similarly, in some aspects, a non-transitory computer program product is described that is capable of storing instructions that, when executed by at least one programmable processor, cause at least one programmable processor to perform operations, which can include: receiving data from one or more computers by an application of a client device identifying inferred demographic characteristics of a user of the application; displaying digital content by the application, the digital content including computer-readable code for reporting events related to the digital content, data specifying a set of permitted demographic-based user group identifiers, and an activity identifier for the digital content; determining by the application that a given inferred characteristic matches a given permitted demographic-based user group identifier; and, in response to determining that a given inferred demographic characteristic matches a given permitted demographic-based user group identifier, using computer-readable code to generate and send a request to update one or more event counts for the digital content and the given permitted demographic-based user group identifiers.

[0010] In some aspects, one or more processors are capable of performing operations including: receiving data identifying digital content and a first set of one or more demographic categories for which a demographic report is to be performed from a first browser of a digital content provider; associating a user of a client device displaying the digital content thereon with a second set of one or more demographic categories; transmitting browsing events and at least one common demographic category entered on the client device to an API if the first set of one or more demographic categories and the second set of one or more demographic categories have at least one common demographic category, wherein the API combines the browsing events and at least one common demographic category with at least one browsing event from other users and at least one associated demographic category that is a demographic category in the first set of one or more demographic categories to generate aggregated data; receiving the aggregated data from the API; and generating a report including the aggregated data.

[0011] Associating a user with a second set of one or more demographic categories can include: receiving self-identification of the user with the second set of one or more demographic categories on a first browser; and mapping the user to the second set of one or more demographic categories. Associating a user with a second set of one or more demographic categories can include: transmitting the user's browsing history to a multi-party computing cluster; receiving inferences from a machine learning model within the multi-party computing cluster, these inferences including the second set of demographic categories; and mapping the user to the second set of one or more demographic categories. Generating a report can include: arranging aggregated data in tables; generating analysis based on the aggregated data; and providing the tables and analysis in the report. Reports can be generated in response to a request for a report or automatically at preset time intervals. Requests for reports can be generated by one or more of the digital content providers, content platforms, secure multi-party computing clusters, or publishers that develop and provide the first browser. Requests for reports can be generated automatically at preset time intervals or after the event count exceeds a preset threshold.

[0012] Data identifying digital content and a first set of one or more demographic categories can be included in the aggregation key. Data identifying digital content can include Uniform Resource Locators (URLs) of the resources displaying the digital content. These operations can further include: determining a third set of demographic categories associated with the user; and restricting a first browser on the client device to display digital content indicated by the appropriate digital content provider to be associated with categories in the third set of demographic categories. Determining the third set of one or more demographic categories can include receiving self-identification by the user on the first browser of the third set of one or more demographic categories associated with the user. Determining the third set of one or more demographic categories can include: transmitting the user's browsing history to a multi-party computing cluster; and receiving data from the multi-party computing cluster identifying the third set of demographic categories.

[0013] Methods, systems, apparatuses, computer-programmable products, etc., that implement the features described above are also within the scope of this disclosure. For example, in some aspects, a system is described that can include: at least one programmable processor; and a machine-readable medium storing instructions that, when executed by the at least one programmable processor, cause the at least one programmable processor to perform operations including: receiving data identifying digital content and a first set of one or more demographic categories for which a demographic report is to be performed from a first browser of a digital content provider; associating a user of a client device displaying digital content thereon with a second set of one or more demographic categories; if the first set of one or more demographic categories and the second set of one or more demographic categories have at least one common demographic category, transmitting browsing events and at least one common demographic category entered on the client device to an API, wherein the API combines the browsing events and at least one common demographic category with at least one browsing event of other users and at least one associated demographic category as one of the demographic categories in the first set of one or more demographic categories to generate aggregate data; receiving the aggregate data from the API; and generating a report including the aggregate data.

[0014] Similarly, in another aspect, a non-transitory computer program product is described, capable of storing instructions that, when executed by at least one programmable processor, cause the programmable processor to perform operations including: receiving data identifying digital content and a first set of one or more demographic categories for which a demographic report is to be performed from a first browser of a digital content provider; associating a user of a client device displaying the digital content thereon with a second set of one or more demographic categories; if the first set of one or more demographic categories and the second set of one or more demographic categories have at least one common demographic category, transmitting browsing events and at least one common demographic category entered on the client device to an API, wherein the API combines the browsing events and at least one common demographic category with at least one browsing event from other users and at least one associated demographic category that is a demographic category in the first set of one or more demographic categories to generate aggregated data; receiving the aggregated data from the API; and generating a report including the aggregated data.

[0015] The subject matter described herein can be implemented in specific embodiments to achieve one or more of the following advantages. The digital component distribution technology described herein can identify users with similar interests and expand user group memberships while protecting user privacy, for example, without disclosing users' online activities to any computing system. This protects user privacy and safeguards data from corruption during transmission to or from the platform. Cryptographic techniques such as secure multi-party computation (MPC) enable the expansion of user groups based on similarity in user profiles without the use of third-party cookies. This protects user privacy without negatively impacting the ability to expand user groups and, in some cases, provides better user group expansion based on more complete profiles than is achievable using third-party cookies. MPC technology ensures that user data cannot be obtained in plaintext by any computing system or another party, provided that one computing system in the MPC cluster is honest. Therefore, the method described herein allows for the secure identification, grouping, and transmission of user data without the need for third-party cookies to determine any relationships between user data. This is a radical departure from traditional methods that typically require third-party cookies to determine relationships between data. By grouping user data in this way, the efficiency of delivering data to user devices is improved because there is no need to transmit data content unrelated to specific users. In particular, third-party cookies are not required, thus avoiding their storage and improving memory utilization. The exponential decay technique can be used to build user profiles at the client device, reducing the size of the original data required to build the profile and thereby lowering data storage requirements.

[0016] Demographic reporting technology involves generating reports for digital component providers based on user group memberships to indicate how their digital components perform for specific user groups. By using secure MPC machine learning techniques to predict user demographics, the reports can more accurately reflect the performance of demographic-based user groups without sacrificing user privacy. The process used to generate such reports does not disclose individual user data to any entity, thus protecting user privacy. The reporting framework is designed not to use cookies and prevents the leakage of user data.

[0017] The various features and advantages of the foregoing subject matter are described below with reference to the figures. Additional features and advantages will be apparent from the subject matter and claims described herein. Attached Figure Description

[0018] Figure 1 This is a block diagram of an environment where a secure MPC cluster trains a machine learning model and the machine learning model is used to extend the user group.

[0019] Figure 2 This is a swimlane diagram of an example process for training a machine learning model and using the machine learning model to add users to a user group.

[0020] Figure 3 This is a flowchart illustrating an example process for generating user profiles and sending a share of user profiles to the MPC cluster.

[0021] Figure 4 This is a flowchart illustrating an example process for generating a machine learning model.

[0022] Figure 5 This is a flowchart illustrating an example process for adding a user to a user group using a machine learning model.

[0023] Figure 6 This is a conceptual diagram of an exemplary framework for generating inference results for a user profile.

[0024] Figure 7 This is a conceptual diagram of an exemplary framework for generating inference results for user profiles with improved performance.

[0025] Figure 8 This is a flowchart illustrating an example process for generating inference results for a user profile at an MPC cluster with improved performance.

[0026] Figure 9 This is a flowchart illustrating an example process for preparing and training a second machine learning model at an MPC cluster to improve inference performance.

[0027] Figure 10 This is a conceptual diagram of an exemplary framework for evaluating the performance of a first machine learning model.

[0028] Figure 11 This is a flowchart illustrating an example process for evaluating the performance of a first machine learning model at an MPC cluster.

[0029] Figure 12 This is a flowchart illustrating an example process for generating inference results for a user profile at the computing system of an MPC cluster with improved performance.

[0030] Figure 13 This is a flowchart illustrating a demographic report, which is used to report the effectiveness of digital components to digital component providers.

[0031] Figure 14 This is a flowchart illustrating an example process executed by a client device to facilitate the distribution of demographic-based digital components and demographic reporting.

[0032] Figure 15This is a block diagram of an example computer system.

[0033] In the various figures, the same reference numerals and names indicate the same elements. Detailed Implementation

[0034] Generally, systems and techniques are described that allow users to be categorized into specific demographic groups without the use of third-party cookies and report the effectiveness of activities used to distribute digital components to user groups that are at least partially based on demographic information.

[0035] Demographic-based user group expansion

[0036] The machine learning platforms described in this document enable content platforms to generate and expand user groups based at least in part on demographic information such as age range, geographic location, spoken language, etc. These user groups allow digital component providers to reach specific sets of users for whom they created these user groups. For example, a digital component provider operating a fitness studio specifically for women could benefit from distributing digital components to members of a female user group, which helps avoid showing digital components more relevant to men. To achieve this targeted content distribution without requiring or using cookies, user groups are formed and subsequently expanded using privacy-preserving machine learning systems and techniques. Such systems and techniques are used to train and expand user group memberships using machine learning models while protecting user privacy and ensuring data security.

[0037] Generally, user profiles are not created and maintained on the computing systems of other entities such as content platforms, but rather on the user's client device. To train a machine learning model, a user's client device can optionally send their encrypted user profile (e.g., as secret shares of the user profile) along with other data to multiple computing systems within a secure multi-party computation (MPC) cluster via the content platform. For example, each client device can generate two or more secret shares of the user profile and send the corresponding secret shares to each computing system. The computing systems of the MPC cluster can use MPC techniques to train the machine learning model in a way that protects user privacy by suggesting user groups based on the user's profile, preventing any computing system in the MPC cluster (or any other party other than the user) from obtaining any user's profile in plaintext (which can also be referred to as plain text). The machine learning model can be a k-nearest neighbor (k-NN) model. In some implementations, the machine learning model can be a k-NN augmented (e.g., via gradient boosting) with other machine learning models (e.g., neural network models or gradient boosting decision trees).

[0038] After training the machine learning model, it can be used to suggest one or more user groups for each user based on their profile. For example, a user's client device can query the MPC cluster for suggested user groups for that user or determine whether the user should be added to a specific user group. Various inference techniques, such as binary classification, regression (e.g., using the arithmetic mean or root mean square), and / or multi-class classification, can be used to identify user groups. User group memberships can be used to deliver content to users in a privacy-preserving and secure manner.

[0039] Example systems for generating and using machine learning models

[0040] Figure 1 This is a block diagram of an environment 100 in which a secure MPC cluster 130 trains a machine learning model, and the machine learning model is used to expand a user group. Example environment 100 includes a data communication network 105, such as a local area network (LAN), a wide area network (WAN), the Internet, a mobile network, or a combination thereof. Network 105 connects client devices 110, the secure MPC cluster 130, publishers 140, websites 142, and content platforms 150. Example environment 100 may include many different client devices 110, secure MPC clusters 130, publishers 140, websites 142, content platforms 150, digital component providers 160, and aggregation systems 180.

[0041] Client device 110 is an electronic device capable of communicating via network 105. Example client device 110 includes a personal computer, a mobile communication device (e.g., a smartphone), and other devices capable of sending and receiving data via network 105. Client device 110 may also include a digital assistant device that accepts audio input via a microphone and outputs audio via a speaker. When the digital assistant detects a “hot word” or “hot phrase” that activates the microphone to accept audio input, it can be put into listening mode (e.g., ready to accept audio input). The digital assistant device may also include a camera and / or display to capture images and visually present information. Digital assistants can be implemented in various forms of hardware devices, including wearable devices (e.g., watches or glasses), smartphones, speaker devices, tablet devices, or other hardware devices. Client device 110 may also include a digital media device, such as a streaming device that plugs into a television or other display to stream video to the television.

[0042] Client device 110 typically includes applications 112, such as web browsers and / or native applications, to facilitate the sending and receiving of data over network 105. Native applications are applications developed for a specific platform or device (e.g., a mobile device with a specific operating system). Publisher 140 is able to develop and provide native applications to client device 110, for example, by making native applications available for download. Web browsers are able to request resources 145 from a web server hosting website 142 of publisher 140, for example, in response to a user of client device 110 entering the resource address of resource 145 in the web browser's address bar or selecting a link referencing the resource address. Similarly, native applications are able to request application content from a publisher's remote server.

[0043] Some resources, application pages, or other application content can include digital component slots for presenting digital components together with resource 145 or application pages. As used throughout this document, the phrase "digital component" refers to a discrete unit of digital content or digital information (e.g., a video clip, audio clip, multimedia clip, image, text, or another unit of content). Digital components can be stored electronically on a physical storage device as a single file or as a collection of files, and digital components can take the form of video files, audio files, multimedia files, image files, or text files and include advertising information, making an advertisement a digital component. For example, a digital component can be content designed to complement the content of a web page or other resource presented by application 112. More specifically, a digital component can include digital content related to the resource content (e.g., a digital component can address the same topic as the web page content, or relate to related topics). The availability of digital components can thus complement and often enhance the web page or application content.

[0044] When application 112 loads a resource (or application content) that includes one or more digital component slots, application 112 can request digital components for each slot. In some implementations, the digital component slot can include code (e.g., a script) that enables application 112 to request digital components from a digital component distribution system that selects and provides the digital components to application 112 for presentation to the user of client device 110.

[0045] Content platform 150 can include a supply-side platform (SSP) and a demand-side platform (DSP). Generally, content platform 150 manages the selection and distribution of digital components on behalf of publisher 140 and digital component provider 160.

[0046] Some publishers 140 use an SSP to manage the process of obtaining digital components for their resources and / or applications. An SSP is a technology platform implemented in hardware and / or software that automates the process of obtaining digital components for resources and / or applications. Each publisher 140 can have one or more SSPs. Some publishers 140 may use the same SSP.

[0047] Digital component provider 160 is capable of creating (or otherwise publishing) digital components that are rendered in digital component slots within a publisher's resources and applications. Digital component provider 160 can use a DSP to manage the supply of its digital components for rendering in digital component slots. A DSP is a technology platform implemented in hardware and / or software that automates the process of distributing digital components for rendering with resources and / or applications. The DSP is capable of interacting on behalf of digital component provider 160 with multiple supply-side platforms (SSPs) to provide digital components for rendering with resources and / or applications from multiple different publishers 140. Generally, the DSP is capable of receiving requests for digital components (e.g., from SSPs), generating (or selecting) selection parameters for one or more digital components created by one or more digital component providers based on the request, and providing data related to the digital component (e.g., the digital component itself) and the selection parameters to the SSP. The SSP is then able to select the digital component for rendering at client device 110 and provide client device 110 with data that enables client device 110 to render the digital component.

[0048] In some cases, it is beneficial for users to receive digital components associated with web pages, application pages, or other electronic resources that they have previously visited and / or interacted with. To distribute such digital components to users, users can be assigned to user groups, such as user interest groups, groups of similar users, or other group types involving similar user data, when they access a specific resource or perform a specific action at that resource (e.g., interacting with a specific item presented on a web page or adding the item to a virtual shopping cart). User groups can be generated by digital component provider 160. That is, each digital component provider 160 can assign users to their respective user groups when users access electronic resources provided by digital component provider 160. In some implementations, publisher 140, content platform 150, or MPC cluster 130 can also assign users to selected user groups.

[0049] To protect user privacy, user membership can be maintained at the user's client device 110, for example, by one of the applications 112 or the operating system of the client device 110, rather than by the digital component provider, content platform or other party. In a specific example, a trusted program (e.g., a web browser or operating system) can maintain a list of user group identifiers (“user group list”) for users using that web browser or another application. The user group list can include a group identifier for each user group to which the user has been added. The digital component provider 160 or content platform (e.g., a DSP) that creates the user groups can assign user group identifiers to their user groups. The user group identifier for a user group can describe the group (e.g., a female group) or represent the group using a code (e.g., a non-descriptive alphanumeric sequence). Generally, user groups can be demographic-based groups. Demographically based user groups can be associated with one or more demographic characteristics and include users who have one or more demographic characteristics (e.g., based on self-reporting) or who are predicted and / or classified as having demographic characteristics using a privacy-preserving machine learning model as members. Such user groups can also have other non-demographic characteristics, such as those related to an activity or product. The user's user group list can be stored in a secure storage device at the client device 110 and / or can be encrypted during storage to prevent access to the list by others.

[0050] When application 112 presents resources or application content related to digital component provider 160, or web pages on website 142, the resources may request application 112 to add one or more user group identifiers to the user group list. In response, application 112 may add one or more user group identifiers to the user group list and securely store the user group list.

[0051] Content platform 150 can use a user's user group membership to select digital components or other content that may be of interest to the user or may otherwise benefit the user / user device. For example, such digital components or other content may include data that improves user experience, enhances the operation of the user device, or otherwise benefits the user or user device. However, user privacy is protected when using user group membership data to select digital components by preventing content platform 150 from providing a list of user group identifiers in a way that associates user group identifiers with specific users. In some implementations, security and privacy protections can be even stronger. For example, computing systems MPC1 and MPC2 can be independent, and only application 112 can be permitted to view user group identifiers in plaintext.

[0052] Application 112 can provide user group identifiers from a list of user groups to a trusted computing system that interacts with content platform 150 to select digital components for presentation at client device 110 in a manner based on user group membership to prevent content platform 150 or any other entity other than the user itself from knowing any user's user group membership.

[0053] In some cases, it is beneficial for both the user and the digital component provider to expand user groups to include users with similar interests or other similar data (e.g., similar demographics) to users who are already members of the user group. Usefully, this can be achieved without using third-party cookies. For example, a first user might be interested in skiing and could be a member of a user group for a specific ski resort. A second user might also be interested in skiing but be unaware of this ski resort and not be a member of that specific ski resort's user group. If the two users have similar interests or data, such as similar user profiles, the second user can be added to the user group for that specific ski resort, allowing the second user to receive content related to that ski resort and that might be of interest to or otherwise beneficial to the second user or their user device, such as digital components. In other words, user groups can be expanded to include other users with similar user data. In a specific example, it is possible to expand demographic-based user groups to include other users with the same or similar demographics associated with the demographic-based user group.

[0054] The secure MPC cluster 130 is capable of training a machine learning model that suggests user groups to a user (or its application 112) based on a user's profile, or can be used to generate user group suggestions to a user (or its application 112) based on a user's profile. The secure MPC cluster 130 includes two computing systems, MPC1 and MPC2, that perform secure MPC techniques to train the machine learning model. Although the example MPC cluster 130 includes two computing systems, more computing systems can be used, as long as the MPC cluster 130 includes more than one computing system. For example, the MPC cluster 130 can include three computing systems, four computing systems, or another suitable number of computing systems. Using more computing systems in the MPC cluster 130 can provide greater security and fault tolerance, but may also increase the complexity of the MPC process.

[0055] Computing systems MPC1 and MPC2 can be operated by different entities. In this way, each entity may not be able to access user profiles in plaintext. Plaintext is text that is not computationally marked, specially formatted, or written in code or data (including binary files) in a form that can be viewed or used without keys or other decryption devices or other decryption processes. For example, one of the computing systems MPC1 or MPC2 can be operated by a trusted party different from the user, publisher 140, content platform 150, and digital component provider 160. For example, an industry group, government group, or browser developer may maintain and operate one of the computing systems MPC1 and MPC2. Another computing system may be operated by different groups among these groups, such that different trusted parties operate each of the computing systems MPC1 and MPC2. Preferably, the different parties operating the different computing systems MPC1 and MPC2 do not have an incentive to collude to compromise user privacy. In some implementations, the computing systems MPC1 and MPC2 are architecturally separate and monitored to not communicate with each other outside of performing the secure MPC processes described in this document.

[0056] In some implementations, the MPC cluster 130 trains one or more k-NN models for each content platform 150 and / or for each digital component provider 160. For example, each content platform 150 can manage the distribution of digital components for one or more digital component providers 160. A content platform 150 can request the MPC cluster 130 to train a k-NN model for the one or more digital component providers 160 it manages to distribute digital components. Generally, a k-NN model represents the distance between user profiles (and optionally additional information) of a set of users. Each k-NN model of a content platform can have a unique model identifier. Figure 4 The diagram below illustrates and describes an example process for training a k-NN model.

[0057] After training a k-NN model for content platform 150, content platform 150 can query, or application 112 on client device 110 can query, the k-NN model to identify one or more user groups for users of client device 110. For example, content platform 150 can query the k-NN model to determine whether the "k" user profiles closest to a threshold number of users are members of a specific user group. If so, content platform 150 can add the user to that user group. If a user group has been identified for a user, content platform 150 or MPC cluster 130 can request application 112 to add the user to the user group. If approved by the user and / or application 112, application 112 can add a user group identifier for the user group to the user group list stored at client device 110.

[0058] In some implementations, application 112 may provide a user interface that enables users to manage the user groups to which they are assigned. For example, the user interface may allow users to remove user group identifiers, preventing all or specific resources 145, publishers 140, content platforms 150, digital component providers 160, and / or MPC clusters 130 from adding users to user groups (e.g., preventing entities from adding user group identifiers to a list of user group identifiers maintained by application 112). This provides users with greater transparency, choice / consent, and control.

[0059] In addition to the descriptions throughout this document, users may be provided with controls (e.g., user interface elements that users can interact with) allowing them to choose whether and when a system, program, or feature described herein can collect user information (e.g., information about a user's social networks, social actions or activities, occupation, user preferences, or current location) and whether to send content or communications to the user from a server. Furthermore, some data may be processed in one or more ways before it is stored or used, causing personally identifiable information to be removed. For example, a user's identity may be processed to the point that personally identifiable information about the user cannot be determined, or, if location information is available, the user's geographic location may be generalized (e.g., to the city, zip code, or state level), making it impossible to determine the user's specific location. Therefore, users have control over what information is collected about them, how that information is used, and what information is provided to them.

[0060] In some implementations, example environment 100 can also facilitate reporting of data such as browsing events, such as impressions, clicks, and / or conversions, by various users in corresponding demographic categories. In such cases, content platform 150 (e.g., DSP or SSP, and in some implementations, a separate reporting platform) can implement or communicate with a reporting API that can communicate with an aggregation system 180 that combines browsing events (e.g., impressions, clicks, and / or conversions, and / or their absence / absence) and demographic categories to generate aggregated data. Aggregation system 180 can deliver the aggregated data to content platform 150 (or the separate reporting platform) in the form of a report or in the form of data that can be easily combined and inserted into the report.

[0061] Example process for generating and using machine learning models

[0062] Figure 2This is a swimlane diagram of an example process 200 for training a machine learning model and using the machine learning model to add users to a user group. The operation of process 200 can be implemented, for example, by client device 110, computing systems MPC1 and MPC2 of MPC cluster 130, and content platform 150. The operation of process 200 can also be implemented as instructions stored on one or more computer-readable media that may be non-transitory, and execution of the instructions by one or more data processing devices can cause one or more data processing devices to perform the operation of process 200. Although process 200 and the following other processes are described in relation to two computing systems MPC cluster 130, MPC clusters with more than two computing systems can also be used to execute similar processes.

[0063] Content platform 150 can initiate the training and / or updating of one of its machine learning models by requesting application 112 running on client device 110 to generate user profiles for its respective users and uploading a secret-shared and / or encrypted version of the user profile to MPC cluster 130. For the purposes of this document, the secret share of the user profile can be considered as an encrypted version of the user profile, since the secret share is not plaintext. Generally, each application 112 can store user profile data and generate updated user profiles in response to requests received from content platform 150. Since the content of the user profiles and the machine learning models differ for different content platforms 150, application 112 running on a user's client device 110 can maintain data for multiple user profiles and generate multiple user profiles, each user profile specific to a specific content platform or a specific model owned by a specific content platform.

[0064] Application 112 running on client device 110 constructs user profiles for users of client device 110 (202). A user's profile may include data related to events initiated by the user and / or events that may have been initiated by the user relative to electronic resources (e.g., web pages or application content). Events may include views of electronic resources, views of digital components, user interactions with electronic resources or digital components (e.g., selection of such interactions), the absence of user interactions with electronic resources or digital components (e.g., selection of such interactions), transformations that occur after a user interacts with an electronic resource, and / or other appropriate events related to the user and the electronic resource.

[0065] A user's profile can be specific to content platform 150 or a machine learning model chosen by content platform 150. For example, see the reference below. Figure 3 In more detail, each content platform 150 can request application 112 to generate or update user profiles specific to that content platform 150.

[0066] A user's profile can be in the form of a feature vector. For example, a user profile can be an n-dimensional feature vector. Each of the n dimensions can correspond to a specific feature, and the value of each dimension can be a value specific to that feature. For example, one dimension could be whether a particular digital component is presented to the user (or interacted with by the user). In this example, the value of that feature could be "1" if the digital component is presented to the user (or interacted with by the user), or "0" if the digital component has not yet been presented to the user (or interacted with by the user). Figure 3 The diagram below illustrates and describes an example process for generating user profiles for users.

[0067] In some implementations, content platform 150 may want to train a machine learning model based on additional signals, such as context signals, signals related to a specific digital component, or signals related to the user that application 112 may not know or may not be able to access, such as the current weather at the user's location. For example, content platform 150 may want to train a machine learning model to predict whether a user will interact with a specific digital component when it is presented to the user in a specific context. In this example, for each presentation of a digital component to the user, the context signals may include the current geographic location of the client device 110 (if the user grants permission), signals describing the content of the electronic resource presented with the digital component, and signals describing the digital component, such as the content of the digital component, the type of the digital component, and the location of the digital component on the electronic resource. In another example, one dimension could be whether the digital component presented to the user has a specific type. In this example, the values ​​could be 1 for travel, 2 for food, 3 for movies, etc. For ease of subsequent description, P i This will represent both the user profile and the additional signals associated with the i-th user profile (e.g., field signals and / or digital component level signals).

[0068] Application 112 generates user profiles for users. i The share (204). In this example, application 112 generates user profile P. i The application generates two shares, one for each compute system in the MPC cluster 130. Note that each share can be a random variable that does not reveal anything about the user profile. The two shares will need to be combined to obtain the user profile. If the MPC cluster 130 includes more compute systems participating in the training of machine learning models, application 112 will generate more shares, one for each compute system. In some implementations, to protect user privacy, application 112 may use a pseudo-random function to generate the user profile P. iIt is split into shares. In other words, application 112 can use the pseudo-random function PRF(P i To generate two shares {[P]} i,1 ],[P i,2 The exact split can depend on the secret-sharing algorithm and cryptographic library used by application 112.

[0069] In some implementations, application 112 can also provide one or more labels to MPC cluster 130. While labels may not be used when training a machine learning model with a certain architecture (e.g., k-NN), labels can be used to fine-tune hyperparameters controlling the model training process (e.g., the value of k), evaluate the quality of the trained machine learning model, or make predictions, i.e., determine whether to suggest user groups to a user. Labels can include one or more user group identifiers, for example, for a user and accessible to content platform 150. That is, labels can include user group identifiers for user groups managed by or accessible to content platform 150. In some implementations, a single label includes multiple user group identifiers for a user. In some implementations, user-specific labels can be heterogeneous and include all user groups to which the user is a member, along with additional information, such as whether the user interacts with a given digital component. This allows the k-NN model to be used to predict whether another user will interact with a given digital component. Labels for each user profile can indicate the user group membership of the user corresponding to the user profile.

[0070] The tags for a user profile predict the user groups to which the user corresponding to the input profile will or should be added. For example, the tags corresponding to the k nearest neighbor user profiles of the input user profile can predict the user groups to which the user corresponding to the input user profile will or should be added, based on the similarity between the user profiles. These predicted tags can be used to suggest user groups to users or request the application to add users to user groups corresponding to the tags.

[0071] If labels are included, application 112 can also assign each label to a specific label. i Break it down into shares, for example [label] i,1 ] and [label i,2 In this way, without communication between computing systems MPC1 and MPC2, neither MPC1 nor MPC2 can obtain [P]. i,1 ] or [P i,2 Reconstruct P i Or from [label] i,1 ] or [label] i,2 Reconstruct the label i .

[0072] Application 112 pairs of user profiles P i share [P] i,1 ] or [P i,2 ] and / or each label i share [label] i,1 ] or [label] i,2 Encryption is performed (206). In some implementations, application 112 generates a user profile P. i The first share [P] i,1 ] and label i First share [label] i,1 The compound message is encrypted using the encryption key of the computing system MPC1. Similarly, application 112 generates a user profile P. i The second share [P] i,2 ] and label i Second share [label] i,2 The function `PubKeyEncrypt([P]` represents a composite message and encrypts it using the encryption key of the MPC2 computing system. These functions can be represented as `PubKeyEncrypt([P]`. i,1 ]||[label i,1 ],MPC1) and PubKeyEncrypt([P i,2 ]||[label i,2 ], MPC2), where PubKeyEncrypt represents the public-key encryption algorithm using the corresponding public key of MPC1 or MPC2. The symbol "||" represents a reversible method for composing complex messages from multiple simple messages, such as JavaScript Object Representation (JSON), Concise Binary Object Representation (CBOR), or protocol buffers.

[0073] Application 112 provides encrypted shares to content platform 150 (208). For example, application 112 can transmit encrypted shares of user profiles and tags to content platform 150. Because each share is encrypted using the encryption key of computing system MPC1 or MPC2, content platform 150 cannot access the user's user profile or tags.

[0074] Content platform 150 can receive user profile shares and tag shares from multiple client devices. Content platform 150 can initiate training of a machine learning model by uploading user profile shares to computing systems MPC1 and MPC2. Although tags may not be used during training, content platform 150 can upload tag shares to computing systems MPC1 and MPC2 for use in optimizing the training process (e.g., hyperparameter tuning), evaluating model quality, or later when querying the model.

[0075] Content platform 150 will receive the first encrypted share (e.g., PubKeyEncrypt([P...)) from each client device 110. i,1 ]||[label i,1 The content platform 150 uploads the second encrypted share (e.g., PubKeyEncrypt([P, MPC1)) to the computing system MPC1(210). Similarly, the content platform 150 uploads the second encrypted share (e.g., PubKeyEncrypt([P, MPC1)) to the computing system MPC1(210). i,2 ]||[label i,2 The data is uploaded to the computing system MPC2(212). Both uploads can be performed in batches and can include encrypted portions of user profiles and tags received during a specific time period used to train the machine learning model.

[0076] In some implementations, the order in which content platform 150 uploads the first encrypted share to computing system MPC1 must match the order in which content platform 150 uploads the second encrypted share to computing system MPC2. This allows computing systems MPC1 and MPC2 to be properly matched with two shares of the same secret (e.g., two shares of the same user profile).

[0077] In some implementations, content platform 150 may explicitly assign the same pseudo-randomly or sequentially generated identifiers to shares of the same secret to facilitate matching. While some MPC techniques can rely on random shuffling of inputs or intermediate results, the MPC techniques described in this document may not include such random shuffling and may instead rely on upload order for matching.

[0078] In some implementations, operations 208, 210, and 212 can be directly applied by 112 to [P] i,1 ]||[label i,1 Upload [P] to MPC1 and [P] i,2 ]||[label i,2 The upload process to MPC2 is replaced. This replacement process reduces the infrastructure costs for content platform 150 to support operations 208, 210, and 212, and reduces the latency of starting to train or update machine learning models in MPC1 and MPC2.

[0079] Computing systems MPC1 and MPC2 generate machine learning models (214). Each generation of a new machine learning model based on user profile data can be referred to as a training session. Computing systems MPC1 and MPC2 can train machine learning models based on encrypted portions of user profiles received from client device 110. For example, computing systems MPC1 and MPC2 can use MPC technology to train a k-NN model based on portions of user profiles.

[0080] To minimize or at least reduce cryptographic computation, and thus minimize or at least reduce the computational burden placed on computing systems MPC1 and MPC2 to protect user privacy and data during both model training and inference, MPC cluster 130 is able to use random projection techniques, such as SimHash, to quickly, securely, and probabilistically quantize two user profiles P. i and P j The similarity between them. This can be determined by identifying the similarity between two user profiles P. i and P j The Hamming distance between two bit vectors resulting from SimHash determines the two user profiles P. i and P j The similarity between them. This Hamming distance is inversely proportional to the cosine distance between the two user profiles with a high probability.

[0081] Conceptually, for each training session, m random projected hyperplanes U = {U1, U2, ..., U...} can be generated. m The random projection hyperplane can also be called the random projection plane. One objective of the multi-step computation between computation systems MPC1 and MPC2 is to compute for each user profile P used in the training of the k-NN model. i Create a bit vector B of length m i In this bit vector B i In, each bit B i,j U represents one of the projection planes j and user profile P i The sign of the dot product, that is, B for all j∈[1, m]. i,j =sign(U j ⊙P i ), where ⊙ represents the dot product of two vectors of equal length. That is, each bit represents the user profile P. i Located in plane U j Which side. A bit value of one represents a positive sign, while a bit value of zero represents a negative sign. In some implementations, to protect user privacy, the above SimHash algorithm can be performed on encrypted (e.g., in the form of secret shares) user profiles and / or projection hyperplanes, so that neither MPC1 nor MPC2 can access the user profiles and / or projection matrices in plaintext.

[0082] In some implementations, before applying random projection during training, MPC1 and MPC2 collaboratively compute the average (also known as the mean), or mean_P, of all user profiles in the training dataset. For privacy protection, the computation of mean_P can be done in a secret share. MPC1 and MPC2 then compute the mean_P from each user profile P. iSubtract mean_P from the middle, then randomly project the result, i.e., P. i -mean_P. The subtraction step can be referred to as "zero mean". For privacy protection, the zero mean step can also be performed on the secret share. If zero mean is applied during training, MPC1 and MPC2 will also apply zero mean at prediction time. That is, for the user profile P in the request to be predicted, MPC1 and MPC2 will compute P-mean_P, i.e., the same mean_P computed during training, and then apply a random projection to P-mean_P (using the same random projection matrix in the secret share selected during training).

[0083] At the end of each multi-step computation, each of the two computation systems MPC1 and MPC generates an intermediate result that includes a bit vector of each user profile in plaintext, a share of each user profile, and a share of a label for each user profile. For example, the intermediate result of computation system MPC1 could be the data shown in Table 1 below. Computation system MPC2 would have a similar intermediate result, but with different shares for each user profile and each label. To add additional privacy protection, each of the two servers in MPC cluster 130 is only able to obtain half of the m-dimensional bit vector in plaintext; for example, computation system MPC1 obtains the first m / 2 dimension of all m-dimensional bit vectors, and computation system MPC2 obtains the second m / 2 dimension of all m-dimensional bit vectors. Plaintext is text that is not computationally marked, specially formatted, or written in code or data (including binary files) in a form that can be viewed or used without keys or other decryption devices or other decryption processes.

[0084] plaintext bit vector <![CDATA[For the MPC1 share of P i > <![CDATA[For the MPC1 share of label i > … … … <![CDATA[B i ]]> … … <![CDATA[B i+1 ]]> … … … … …

[0085] Table 1

[0086] Given two arbitrary user profile vectors P of unit length i ≠ j i and P j It has been shown that there are two user profile vectors P i and P j Bit vector B i and B j The Hamming distance between them has a high probability of being related to the user profile vector P. i and P j The cosine distance between them is proportional, assuming the number of random projections m is large enough.

[0087] Based on the intermediate results shown above and because of the bit vector B iIt is plaintext, so each computing system MPC1 and MPC2 can, for example, independently create corresponding k-NN models using the k-NN algorithm through training. Computing systems MPC1 and MPC2 can use the same or different k-NN algorithms. Figure 4 The diagram below illustrates and describes an example process for training a k-NN model. Once the k-NN model is trained, application 112 can query the k-NN model to determine whether to add the user to the user group.

[0088] Application 112 submits an inference request (216) to MPC cluster 130. In this example, application 112 forwards the inference request to computing system MPC1. In other examples, application 112 may forward the inference request to computing system MPC2. Application 112 may submit an inference request in response to a request to submit an inference request from content platform 150. For example, content platform 150 may request application 112 to query the k-NN model to determine whether the user of client device 110 should be added to a specific user group, such as a demographic-based user group. This request may be referred to as an inference request to infer whether the user should be added to the user group.

[0089] In order to initiate an inference request, content platform 150 can send an inference request token M to application 112. infer Inferring request token M infer This enables the servers in MPC cluster 130 to verify that application 112 is authorized to query a specific machine learning model owned by a specific domain. If model access control is optional, the request token M is inferred. infer It is optional. Infer request token M infer It can have the following items shown and described in Table 2 below.

[0090]

[0091] Table 2

[0092] In this example, the request token M is inferred. infer The digital signature is generated based on seven items and a private key from the content platform 150. The eTLD+1 is the valid top-level domain (eTLD) plus one level above the public suffix. An example eTLD+1 is "example.com", where ".com" is the top-level domain.

[0093] In order to request inferences about a specific user, content platform 150 can generate an inference request token M. infer The token is then sent to application 112 running on the user's client device 110. In some implementations, content platform 150 uses the public key of application 112 to infer the request token M. inferEncryption is performed so that only application 112 can use its secret private key corresponding to the public key to infer the request token M. infer Decryption is performed. In other words, the content platform can send PubKeyEnc(M) to application 112. infer ,application_public_key).

[0094] Application 112 can infer the request token M infer Decryption and verification are performed. Application 112 is able to use its private key to decrypt and verify the encrypted inference request token M. infer Decryption is performed. Application 112 can verify the deduced request token M through the following steps. infer (i) Verify the digital signature using the public key of content platform 150 corresponding to the private key used to generate the digital signature, and (ii) ensure that the token creation timestamp is not stale, for example, that the time indicated by the timestamp is within a threshold time amount of the current time at which verification is in progress. If the request token M is inferred... infer If valid, application 112 can query MPC cluster 130.

[0095] Conceptually, an inference request can include the model identifier of the machine learning model, the current user profile, and the P... i , k (the number of nearest neighbors to retrieve), optionally additional signals (e.g., field signals or digital component signals), aggregation functions, and aggregation function parameters. However, to prevent the user profile P from being... i To protect user privacy, the application 112 can disclose the user profile P in plaintext to computing systems MPC1 or MPC2. i Split into two shares, one for MPC1 and one for MPC2 [P] i,1 ] and [P i,2 Application 112 can then select, for example, randomly or pseudo-randomly, one of two computing systems, MPC1 or MPC2, for querying. If application 112 selects computing system MPC1, then application 112 can query the system with the first share [P]. i,1 The second share of the encrypted version (e.g., PubKeyEncrypt([P i,2 The computing system MPC1 sends a single request to the second share [P]. In this example, application 112 uses the public key of computing system MPC2 to send a single request to the second share [P]. i,2 Encryption is used to prevent the computing system MPC1 from accessing [P] i,2 This will enable the computing system MPC1 to access [P] i,1 ] and [P i,2 Reconstructing User Profiles (P) i .

[0096] As described in more detail below, computing systems MPC1 and MPC2 collaboratively compute to user profile P. i The computational systems MPC1 and MPC2 can then use one of several possible machine learning techniques (e.g., binary classification, multi-class classification, regression, etc.) to determine whether to add a user to a user group based on the user profile of the k nearest neighbors. For example, an aggregation function can identify machine learning techniques (e.g., binary, multi-class, regression) and the aggregation function parameters can be based on the aggregation function.

[0097] In some implementations, the aggregation function parameter can include the user group identifier that the content platform 150 is querying for the k-NN model of the user group. For example, the content platform 150 might want to know whether to add the user to a user group that is related to hiking and has the user group identifier "hiking". In this example, the aggregation function parameter can include the "hiking" user group identifier. Generally, computing systems MPC1 and MPC2 can determine whether to add the user to the user group based on the number of the k nearest neighbors that are members of the user group, such as based on their labels.

[0098] MPC cluster 130 provides inference results (218) to application 112. In this example, computing system MPC1, which receives the query, sends the inference results to application 112. The inference results can indicate to application 112 whether the user should be added to zero or more user groups. For example, the user group result can specify a user group identifier for the user group. However, in this example, computing system MPC1 will know the user group. To prevent this, computing system MPC1 can calculate a share of the inference result and computing system MPC2 can calculate another share of the same inference result. Computing system MPC2 can provide computing system MPC1 with an encrypted version of its share, where the share is encrypted using application 112's public key. Computing system MPC1 can provide application 112 with a share of its inference result and an encrypted version of computing system MPC2's share of the user group result. Application 112 can decrypt computing system MPC2's share and calculate the inference result from both shares. Figure 5 The diagram below illustrates and describes an example process for querying a k-NN model to determine whether to add a user to a user group. In some implementations, to prevent computing system MPC1 from forging the results of computing system MPC2, computing system MPC2 digitally signs its results before or after encrypting them using the public key of application 112. Application 112 uses the public key of MPC2 to verify the digital signature of computing system MPC2.

[0099] Application 112 updates the list of user groups for the user (220). For example, if the inference is that the user should be added to a specific user group, application 112 can add the user to the user group. In some implementations, application 112 can prompt the user for permission to add the user to the user group.

[0100] Application 112 transmits a request for content (222). For example, application 112 may transmit a request for a digital component to content platform 150 in response to loading an electronic resource with a digital component slot. In some implementations, the request may include one or more user group identifiers that include the user as a member of a user group. For example, application 112 may obtain one or more user group identifiers from a list of user groups and provide the user group identifiers upon request. In some implementations, techniques may be used to prevent the content platform from associating the user group identifier with the user from whom the request is received, application 112, and / or client device 112.

[0101] Content platform 150 delivers content (224) to application 112. For example, content platform 150 may select a digital component based on a user group identifier and provide that digital component to application 112. In some implementations, content platform 150 collaborates with application 112 to select a digital component based on a user group identifier without disclosing the user group identifier from application 112.

[0102] Application 112 displays or otherwise implements the received content (226). For example, application 112 can display the received digital component in the digital component slot of the electronic resource.

[0103] Example process for generating user profiles

[0104] Figure 3 This is a flowchart illustrating an example process 300 for generating user profiles and sending portions of those profiles to the MPC cluster. The operation of process 300 can be performed, for example, by... Figure 1 The operation of process 300 can also be implemented as instructions stored in one or more computer-readable media that may be non-transitory, and execution of the instructions by one or more data processing devices can cause the one or more data processing devices to perform the operation of process 300.

[0105] Application 112, running on the user's client device 110, receives data (302) related to an event. An event can be, for example, the presentation of an electronic resource on the client device 110, the presentation of a digital component on the client device 110, user interaction with an electronic resource or digital component on the client device 110, a conversion of a digital component, or the absence of user interaction with or conversion of a presented electronic resource or digital component. When an event occurs, content platform 150, publisher 140, or digital component provider 160 can provide event-related data to application 112 for use in generating the user's profile.

[0106] Application 112 can generate different user profiles for each content platform 150. That is, a user's profile specific to a particular content platform 150 can include only event data received from that specific content platform 150. This protects user privacy by not sharing event-related data with other content platforms. In some implementations, application 112 can generate different user profiles for each machine learning model owned by content platform 150, based on a request from content platform 150. Depending on the design goals, different machine learning models may require different training data. For example, a first model might be used to determine whether to add a user to a user group. A second model might be used to predict whether a user will interact with a digital component. In this example, the user profile for the second model can include additional data not present in the user profile for the first model, such as whether the user interacted with the digital component.

[0107] Content platform 150 can update token M with a brief profile. update Event data is sent in the form of a brief update token M. update It has the following items shown and described in Table 3 below.

[0108]

[0109]

[0110] Table 3

[0111] The model identifier identifies the machine learning model, such as the k-NN model, for which the user profile will be used for training or for user group inference. The profile record is an n-dimensional feature vector that includes event-specific data such as the type of event, electronic resource or digital component, the time the event occurred, and / or other appropriate event data that the content platform 150 wants to use in training the machine learning model and for user group inference. The digital signature is generated using the content platform 150's private key based on seven items.

[0112] In some implementations, in order to protect the update token M during transmission... update Content platform 150 will update token M update Update token M before sending to application 112 update Encryption is performed. For example, content platform 150 can use the application's public key to encrypt the update token M. update Encryption is performed, for example, PubKeyEnc(M update ,application_public_key).

[0113] In some implementations, content platform 150 can update event data with a profile update token M. update Event data and update requests are sent to application 112 in a format that does not encode the event data or update requests. For example, a script originating from content platform 150 running within application 112 can send event data and update requests directly to application 112 via a script API, where application 112 relies on a security model based on the World Wide Web Consortium (W3C) source and / or HTTPS (Secure Hypertext Transfer Protocol) to protect the event data and update requests from forgery or disclosure.

[0114] Application 112 stores data for the event (304). If the event data is encrypted, application 112 can decrypt the event data using its private key corresponding to the public key used to encrypt the event data. If an update token M is used... update If the event data is sent in the form of [method 112], then application 112 can verify the update token M before storing the event data. update Application 112 can verify the update token M through the following steps. update (i) Verify the digital signature using the public key of content platform 150 corresponding to the private key used to generate the digital signature, and (ii) ensure that the token creation timestamp is not stale, for example, that the time indicated by the timestamp is within a threshold time amount of the current time at which verification is in progress. If the token M is updated... update If valid, application 112 can store the event data, for example, by storing an n-dimensional profile record. If any validation fails, application 112 can ignore the update request, for example, by not storing the event data.

[0115] For each machine learning model, for example, for each unique model identifier, application 112 can store event data for that model. For example, application 112 can maintain a data structure for each unique model identifier that includes a set of n-dimensional feature vectors (e.g., a profile record of an update token), and maintain an expiration time for each feature vector. An example data structure for a model identifier is shown in Table 4 below.

[0116] Feature vector Expired n-dimensional eigenvectors Expiration time … …

[0117] Table 4

[0118] Upon receiving a valid update token M update Then, application 112 can update the token M. update The feature vector and expiration time are added to the data structure to update the token M. update The data structure for model identifiers. Periodically, application 112 can remove expired feature vectors from the data structure to reduce storage size.

[0119] Application 112 determines whether to generate a user profile (306). For example, application 112 may generate a user profile for a specific machine learning model in response to a request from content platform 150. The request may be to generate a user profile and return a share of the user profile to content platform 150. In some implementations, application 112 may upload the generated user profiles directly to MPC cluster 130, for example, instead of sending them to content platform 150. To ensure the security of the request to generate and return a share of the user profile, content platform 150 may send an upload token M to application 112. upload .

[0120] Upload Token M upload Able to have update token M update Similar structure, but with different operations (e.g., "update server" instead of "accumulate user profiles"). Upload token M upload It can also include additional items for operational delays. Operational delays can instruct application 112 to delay the calculation and uploading of user profiles, while application 112 accumulates more event data, such as more feature vectors. This allows machine learning models to capture user event data immediately before and after key events (e.g., joining a user group). Operational delays can specify a delay period. In this example, a digital signature can be generated using the content platform's private key based on the other seven items in Table 3 and the operational delay. Content platform 150 can then use the application's public key to update the token M. update A similar method is used for uploading token M upload Encryption is performed, for example, PubKeyEnc(M upload (application_public_key) to protect the upload token M during transmission. upload .

[0121] Application 112 can receive upload token M upload In the case that it is encrypted, the upload token M upload Decrypt and verify the uploaded token M. uploadThis verification is similar to that used to verify the update token M. update The method is as follows. Application 112 can verify the upload token M through the following steps. upload (i) Verify the digital signature using the public key of content platform 150 corresponding to the private key used to generate the digital signature, and (ii) ensure that the token creation timestamp is not stale, for example, that the time indicated by the timestamp is within a threshold time amount of the current time at which verification is in progress. If the uploaded token M upload If valid, application 112 can generate a user profile. If any verification fails, application 112 can ignore the upload request, for example, by not generating a user profile.

[0122] In some implementations, content platform 150 can request application 112 to upload token M in a file. upload User profiles can be uploaded in a format that does not encode the upload request. For example, a script originating from content platform 150 running within application 112 can send the upload request directly to application 112 via a script API, where application 112 relies on a W3C-based security model and / or HTTPS to protect the upload request from forgery or disclosure.

[0123] If a decision is made not to generate a user profile, process 300 can return to operation 302 and wait for additional event data from content platform 150. If a decision is made to generate a user profile, application 112 generates the user profile (308).

[0124] Application 112 can generate user profiles based on stored event data (e.g., data stored in the data structure shown in Table 4). Application 112 can also generate user profiles based on model identifiers included in the request (e.g., upload token M). upload The first item (the content platform eTLD+1 domain) and the second item (the model identifier) ​​are used to access the appropriate data structure.

[0125] Application 112 can compute a user profile by aggregating n-dimensional feature vectors from a data structure within learning periods that have not yet expired. For example, a user profile could be the average of n-dimensional feature vectors from a data structure within learning periods that have not yet expired. The result is an n-dimensional feature vector representing the user in the profile space. Optionally, application 112 can, for example, use L2 normalization to normalize the n-dimensional feature vectors to unit length. Content platform 150 can specify optional learning periods.

[0126] In some implementations, the decay rate can be used to compute user profiles. Since there can be many content platforms 150 using MPC cluster 130 to train machine learning models, and each content platform 150 can have multiple machine learning models, storing user feature vector data can create significant data storage requirements. Using decay techniques can substantially reduce the amount of data stored at each client device 110 for the purpose of generating user profiles for training machine learning models.

[0127] Suppose that for a given machine learning model, there exist k feature vectors {F1, F2, ..., Fk}. k} and its corresponding age (record_age_in_seconds) i Each feature vector is an n-dimensional vector. Application 112 can use the following relation 1 to calculate the user profile:

[0128] Relation 1:

[0129] In this relation, the parameter record_age_in_seconds i It is a brief record F i The amount of time in seconds has already been stored at client device 110, and the parameter decay_rate_in_seconds is the decay rate recorded in the profile in seconds (e.g., in the update token M). update (Received in item 6). In this way, newer feature vectors carry greater weight. This also allows application 112 to avoid storing feature vectors and only store profile records with constant storage. Instead of storing multiple individual feature vectors for each model identifier, application 112 only needs to store an n-dimensional vector P and a timestamp user_profile_time for each model identifier.

[0130] To initialize the n-dimensional vector user profile P and timestamp, the application can set the vector P as an n-dimensional vector, where each dimension has a value of zero, and set user_profile_time to the epoch. To use the new feature vector F at any time... x After updating the user profile P, application 112 can use the following relation 2:

[0131] Relation 2:

[0132] When updating a user profile using relation 2, application 112 can also update the user profile time to the current time (current_time). Note that if application 112 calculates the user profile using the decay rate algorithm described above, operations 304 and 308 are omitted.

[0133] Application 112 generates a share (310) of the user profile. Application 112 can use a pseudo-random function to generate the user profile P. i (For example, an n-dimensional vector P) can be split into shares. That is, application 112 can use the pseudo-random function PRF(P) i To generate a user profile P i The two shares {[Pi,1],[P]} i,2 The exact split can depend on the secret-sharing algorithm and cryptographic library used by application 112. In some implementations, the application uses the Shamir secret-sharing scheme. In some implementations, the application uses an additive secret-sharing scheme. If shares of one or more tags are being provided, application 112 can also generate shares of the tags in the same way.

[0134] Application 112 pairs of user profiles P i share {[P i,1 ],[P i,2 Encryption is performed (312). For example, as described above, application 112 can generate a composite message including a share of user profile and tags and encrypt the composite message to obtain the encryption result PubKeyEncrypt([P i,1 ]||[label i,1 ],MPC1) and PubKeyEncrypt([P i,2 ]||[label i,2 [, MPC2]. The encryption key of MPC cluster 130 is used to encrypt the share to prevent content platform 150 from accessing the user profile in plaintext. Application 112 transmits the encrypted share to the content platform (314). Note that if application 112 transmits the secret share directly to computing systems MPC1 and MPC2, operation 314 is omitted.

[0135] Example process for generating and using machine learning models

[0136] Figure 4 This is a flowchart illustrating an example process 400 for generating a machine learning model. The operation of process 400 can be performed, for example, by... Figure 1 The MPC cluster 130 is implemented. The operation of process 400 can also be implemented as instructions stored on one or more computer-readable media that may be non-transitory, and the execution of the instructions by one or more data processing devices can cause the one or more data processing devices to perform the operation of process 400.

[0137] MPC cluster 130 obtains a share of the user profile (402). Content platform 150 can request MPC cluster 130 to train a machine learning model by sending a share of the user profile to MPC cluster 130. Content platform 150 can access encrypted shares received from client device 110 for a machine learning model within a given time period and upload those shares to MPC cluster 130.

[0138] For example, content platform 150 can transmit an encrypted first share of a user profile and an encrypted first share of its tags to computing system MPC1 (e.g., for each user profile P). i PubKeyEncrypt([P i,1 ]||[label i,1 The tags described herein and other flowcharts in this disclosure can be or include user group identifiers or demographic features. Similarly, content platform 150 can transmit an encrypted second share of a user profile and an encrypted second share of its tags (e.g., for each user profile P) to computing system MPC2. i PubKeyEncrypt([P i,2 ]||[label i,2 ],MPC2)).

[0139] In some implementations of application 112 directly sending a secret share of a user profile to MPC cluster 130, content platform 150 can request MPC cluster 130 to train a machine learning model by sending a training request to MPC cluster 130.

[0140] Computational systems MPC1 and MPC2 create random projection planes (404). Computational systems MPC1 and MPC2 are able to collaboratively create m random projection planes U = {U1, U2, ..., U...}. m These random projection planes should be kept as a secret share between the two computing systems, MPC1 and MPC2. In some implementations, computing systems MPC1 and MPC2 use a Diffie-Hellman key exchange technique to create the random projection planes and maintain their confidentiality.

[0141] As described in more detail below, computation systems MPC1 and MPC2 project their shares of each user profile onto each random projection plane and, for each random projection plane, determine whether the user profile's shares lie on one side of the random projection plane. Each computation system MPC1 and MPC2 is then able to construct a bit vector in the secret shares from the secret shares of the user profile based on the result of each random projection. The user has partial knowledge of the bit vector, for example, the user profile P. i Is it on the projection plane U?k On one side, the computing system MPC1 or MPC2 is allowed to obtain information about P. i This is some knowledge about the distribution of user profile P. i The growth of prior knowledge has a unit length. To prevent computing systems MPC1 and MPC2 from gaining access to this information (e.g., in implementations where this is necessary or preferred for user privacy and / or data security), in some implementations, the random projection plane is in a secret share, so neither computing systems MPC1 nor MPC2 can access the random projection plane in plaintext. In other implementations, a secret-sharing algorithm can be used to apply a random bit-flipping pattern to the random projection result, as described in optional operations 406-408.

[0142] To demonstrate how bit flipping occurs via secret shares, assume two secrets, x and y, whose values ​​are either zero or one with equal probability. The equality operation [x] == [y] will flip the bits of x if y == 0, and will leave the bits of x unchanged if y == 1. In this example, the operation will flip the bit x randomly with a 50% probability. This operation requires a remote procedure call (RPC) between two computing systems, MPC1 and MPC2, and the number of rounds depends on the chosen data size and secret-sharing algorithm.

[0143] Each computing system MPC1 and MPC2 creates a secret m-dimensional vector (406). Computing system MPC1 is capable of creating a secret m-dimensional vector {S1, S2, ..., S...} m}, where each element S i Each has a value of zero or one with equal probability. The computational system MPC1 splits its m-dimensional vector into two shares, namely the first share {[S 1,1 ],[S 2,1 ],…[S m,1 ]} and the second share {[S 1,2 ],[S 2,2 ],…[S m,2 The computational system MPC1 can keep the first share secret and provide the second share to the computational system MPC2. The computational system MPC1 can then discard the m-dimensional vector {S1, S2, ... S}. m}

[0144] The computational system MPC2 can create secret m-dimensional vectors {T1, T2, ..., T}. m}, where each element T i It has a value of zero or one. The computational system MPC2 splits its m-dimensional vector into two shares, namely the first share {[T 1,1 ],[T 2,1 ],…[T m,1 ]} and the second share {[T1,2 ],[T 2,2 ],…[T m,2 The computational system MPC2 can keep the first share secret and provide the second share to the computational system MPC1. The computational system MPC2 can then discard the m-dimensional vector {T1, T2, ..., T}. m}

[0145] Two computing systems, MPC1 and MPC2, use secure MPC technology to calculate the share of the bit-flipping pattern (408). Computing systems MPC1 and MPC2 can calculate the share of the bit-flipping pattern using a secret share MPC equality test involving multiple round trips between computing systems MPC1 and MPC2. The bit-flipping pattern can be based on the above operation [x] == [y]. That is, the bit-flipping pattern can be {S1 == T1, S2 == T2, ... S...} m ==T m Let each ST i =(S i ==T i Each ST i It has a value of zero or one. After the MPC operation is completed, the computing system MPC1 has the first share of the bit-flipping mode {[ST 1,1 [ST] 2,1 ],…[ST m,1 ]} and compute system MPC2 has a second share {[ST] with bit-flipping mode. 1,2 [ST] 2,2 ],…[ST m,2 Each ST i The share allows the two computing systems MPC1 and MPC2 to flip bits in the bit vector in a way that is opaque to either of the two computing systems MPC1 and MPC2.

[0146] Each computing system MPC1 and MPC2 projects the share of each user profile onto each random projection plane (410). That is, for each user profile whose share is received by computing system MPC1, computing system MPC1 is able to project the share [P] onto each random projection plane. i,1 Projected onto each projection plane U j Above. For each share of the user profile and for each random projection plane U j Performing this operation yields a z x m dimension matrix R, where z is the number of available user profiles and m is the number of random projection planes. This can be achieved by computing the projection planes U. j With share [P] i,1 The dot product between [] determines each element R in matrix R. i,j For example, R i,j=U j ⊙[P i,1 The operation ⊙ represents the dot product of two vectors of equal length.

[0147] If bit flipping is used, computing system MPC1 can modify one or more elements R in the matrix using a bit flipping pattern secretly shared between computing systems MPC1 and MPC2. i,j The value of . For each element R of matrix R. i,j The computing system MPC1 can calculate [ST] j,1 ]==sign(R i,j ) as element R i,j The value of R. Therefore, element R i,j The symbol will be in its bit flip mode [ST] j,1 The corresponding bit in the [] is flipped if it has a zero value. This calculation can be performed by multiple RPCs on the computing system MPC2.

[0148] Similarly, for each user profile for which the computing system MPC2 receives a share, the computing system MPC2 is able to allocate the share [P] i,2 Projected onto each projection plane U j Above. For each share of the user profile and for each random projection plane U j Performing this operation produces a z x m dimension matrix R', where z is the number of available user profiles and m is the number of random projection planes. This can be achieved by computing the projection planes U. j With share [P] i,2 The dot product between [] determines each element R in matrix R'. i,j ', for example, R i, ' j =U j ⊙[P i,2 The operation ⊙ represents the dot product of two vectors of equal length.

[0149] If bit flipping is used, computing system MPC2 can modify one or more elements R in the matrix using a bit flipping pattern secretly shared between computing systems MPC1 and MPC2. i,j The value of '. For each element R in matrix R. i,j The computing system MPC2 can calculate [ST] j,2 ]==sign(R i,j ') as element R i,j The value of '. Therefore, element R i,j The sign of ' will be in its bit ST in bit flip mode jThe corresponding bit in the calculation is flipped if it has a zero value. This calculation can be performed by multiple RPCs on the computing system MPC1.

[0150] Computation systems MPC1 and MPC2 reconstruct the bit vector (412). Computation systems MPC1 and MPC2 can reconstruct the bit vector for the user profile based on matrices R and R' of exactly the same size. For example, computation system MPC1 can send a portion of the columns of matrix R to MPC2, and computation system MPC2 can send the remaining portion of the columns of matrix R' to MPC1. In a particular example, computation system MPC1 can send the first half of the columns of matrix R to computation system MPC2, and computation system MPC2 can send the second half of the columns of matrix R' to MPC1. Although horizontal reconstruction is performed using columns in this example, and columns are preferred to protect user privacy, vertical reconstruction using rows is possible in other examples.

[0151] In this example, computing system MPC2 can combine the first half of the columns of matrix R' with the first half of the columns of matrix R received from computing system MPC1 to reconstruct the first half of the bit vector (i.e., m / 2 dimension) in plaintext. Similarly, computing system MPC1 can combine the second half of the columns of matrix R with the second half of the columns of matrix R' received from computing system MPC2 to reconstruct the second half of the bit vector (i.e., m / 2 dimension) in plaintext. Conceptually, computing systems MPC1 and MPC2 have now combined corresponding shares of the two matrices R and R' to reconstruct the bit matrix B in plaintext. This bit matrix B will include a bit vector representing the projection result (projected onto each projection plane) of each user profile received by the machine learning model from content platform 150. Each of the two servers in MPC cluster 130 possesses half of the bit matrix B in plaintext.

[0152] However, if bit flipping is used, computation systems MPC1 and MPC2 have already flipped the bits of the elements in matrices R and R' according to a fixed random pattern for the machine learning model. This random bit flipping pattern is opaque to either of the two computation systems MPC1 and MPC2, making it impossible for either MPC1 or MPC2 to infer the original user profile from the bit vector of the projection result. The cryptographic design further prevents MPC1 or MPC2 from inferring the original user profile by horizontally splitting the bit vector; that is, computation system MPC1 preserves the latter half of the bit vector of the projection result in plaintext, and computation system MPC2 preserves the first half of the bit vector of the projection result in plaintext.

[0153] Computational systems MPC1 and MPC2 generate machine learning models (414). Computational system MPC1 can generate a k-NN model using the latter half of a bit vector. Similarly, computational system MPC2 can generate a k-NN model using the first half of a bit vector. The use of bit flipping and horizontal partitioning of the matrix to generate the model applies the principle of defense in depth to protect the confidentiality of the user profiles used to generate the model.

[0154] Generally, each k-NN model represents the cosine similarity (or distance) between user profiles of a group of users. The k-NN model generated by computing system MPC1 represents the similarity between the latter half of the bit vectors, while the k-NN model generated by computing system MPC2 represents the similarity between the first half of the bit vectors. For example, each k-NN model can define the cosine similarity between half of its bit vectors.

[0155] The two k-NN models generated by computing systems MPC1 and MPC2 are referred to as k-NN models, each possessing a unique model identifier as described above. Computing systems MPC1 and MPC2 are able to store their models along with a share of tags for each user profile used to generate the models. Content platform 150 is then able to query the models to make inferences about user groups against users. In some implementations, to protect user privacy, the tags are encrypted, for example, in the form of secret shares.

[0156] Example process for using machine learning models to infer user groups

[0157] Figure 5 This is a flowchart illustrating an example process 500 for adding a user to a user group using a machine learning model. The operation of process 500 can be performed, for example, by... Figure 1 The process 500 is implemented using an MPC cluster 130 and a client device 110 (e.g., an application 112 running on the client device 110). The operation of process 500 can also be implemented as instructions stored on one or more computer-readable media that may be non-transitory, and execution of the instructions by one or more data processing devices can cause those devices to perform the operation of process 500.

[0158] MPC cluster 130 receives an inference request (502) for a given user profile. Application 112 running on the user's client device 110 can, for example, transmit the inference request to MPC cluster 130 in response to a request from content platform 150. For example, content platform 150 can transmit an inference request token M to application 112. infer Application 112 submits an inference request to MPC cluster 130. The inference request can be an inquiry into whether a user should be added to any number of user groups.

[0159] Inferring Request Token M infer It can include a share of a given user profile, a model identifier of a machine learning model (e.g., a k-NN model) and the owner domain to be used for inference, the number k of the nearest neighbors of the given user profile to be used for inference, additional signals (e.g., context or digital component signals), the aggregation function to be used for inference and any aggregation function parameters to be used for inference, and a signature on all of the above information created by the owner domain using the owner domain's confidential privacy key.

[0160] As mentioned above, in order to prevent a given user profile P from being... i To protect user privacy, the application 112 can disclose a given user profile P in plaintext to computing systems MPC1 or MPC2. i Split into two shares, one for MPC1 and one for MPC2 [P] i,1 ] and [P i,2 Application 112 can then be used with the first share of a given user profile [P]. i,1 ] and a second share of the encrypted version of the given user profile (e.g., PubKeyEncrypt([P i,2 Together, MPC1 and MPC2 send a single inference request to the computing system MPC1. The inference request may also include an inference request token M. infer This enables the MPC cluster 130 to authenticate inference requests. By sending an inference request that includes a first share and an encrypted second share, the number of outgoing requests sent by application 112 is reduced, resulting in savings in computation, bandwidth, and battery power at client device 110.

[0161] In other implementations, application 112 is able to transfer the first share [P] of a given user profile. i,1 ]Sent to computing system MPC1 and the second share of the given user profile [P i,2 ]Sent to computing system MPC2. By sending a second copy of the given user profile [P] without going through computing system MPC1. i,2 The second share is sent to computing system MPC2; it does not need to be encrypted to prevent computing system MPC1 from accessing the second share of a given user profile. i,2 ].

[0162] Each computing system, MPC1 and MPC2, identifies the k nearest neighbors (504) of a given user profile in the secret share representation. Computing system MPC1 is able to use the first share [P] of a given user profile. i,1 To calculate half of the bit vector of a given user profile, the computing system MPC1 can use [the following]. Figure 4Operations 410 and 412 of process 400. That is, the computing system MPC1 can project the share of a given user profile using a random projection vector generated for the k-NN model [P]. i,1 And create a secret share of the bit vector for a given user profile. If bit flipping is used to generate the k-NN model, the computation system MPC1 can then use the first share {[ST] of the bit flipping pattern used to generate the k-NN model. 1,1 [ST] 2,1 ],…[ST m,1 The `}` variable is used to modify elements of the secret share of a bit vector for a given user profile.

[0163] Similarly, computing system MPC1 can provide computing system MPC2 with a second encrypted share of PubKeyEncrypt([P i,2 The computing system MPC2 is able to use its private key to access the second share of a given user profile [P]. i,2 Decrypt and use the second share of the given user profile [P] i,2 This is used to compute half of the bit vector for a given user profile. In other words, the computation system MPC2 can project a share of a given user profile using a random projection vector generated for a k-NN model. i,2 And create a bit vector for a given user profile. If bit flipping is used to generate a k-NN model, the computation system MPC2 can then use the second share {[ST] of the bit flipping pattern used to generate the k-NN model. 1,2 [ST] 2,2 ],…[ST m,2 The system modifies the elements of the bit vector for a given user profile. The computation systems MPC1 and MPC2 then reconstruct the bit vector using horizontal partitioning, as follows: Figure 4 As described in operation 412. After reconstruction is complete, computing system MPC1 has the first half of the overall bit vector for a given user profile and computing system MPC2 has the second half of the overall bit vector for a given user profile.

[0164] Each computing system, MPC1 and MPC2, uses half of the bit vector for a given user profile and its k-NN model to identify k' nearest neighbor user profiles, where k' = α × k, and α is empirically determined based on actual production data and statistical analysis. For example, α = 3 or another suitable number. Computing system MPC1 is able to calculate the Hamming distance between the first half of the overall bit vector and the bit vector for each user profile using the k-NN model. Computing system MPC1 then identifies k' nearest neighbors based on the calculated Hamming distances, for example, the k' user profiles with the lowest Hamming distances. In other words, computing system MPC1 identifies a set of nearest neighbor user profiles based on the share of a given user profile and a k-nearest neighbor model trained using multiple user profiles. Example results in tabular form are shown in Table 5 below.

[0165] Line ID Hamming distance (plaintext) User profile share Share of the label i <![CDATA[d i,1 ]]> <![CDATA[[P i,1 ]]]> <![CDATA[[label i,1 ]]]> … … … …

[0166] Table 5

[0167] In Table 5, each row represents a specific nearest-neighbor user profile and includes the Hamming distance between the first half of the bit vector for each user profile and the bit vector for a given user profile computed by the computing system MPC1. The row for a specific nearest-neighbor user profile also includes a first share of that user profile and a first share of the tag associated with that user profile. As described herein, the tag can be or include user group identifiers or demographic features.

[0168] Similarly, the computation system MPC2 can compute the Hamming distance between the latter half of the overall bit vector and the bit vector of each user profile for the k-NN model. The computation system MPC2 then identifies k' nearest neighbors based on the computed Hamming distances, for example, the k' user profiles with the lowest Hamming distances. Example results in tabular form are shown in Table 5 below.

[0169] Line ID Hamming distance (plaintext) User profile share Share of the label j <![CDATA[d j,2 ]]> <![CDATA[[P j,2 ]]]> <![CDATA[[label j,2 ]]]> … … … …

[0170] Table 6

[0171] In Table 6, each row is for a specific nearest neighbor user profile and includes the Hamming distance between that user profile and a given user profile calculated by the computing system MPC2. The row for a specific nearest neighbor user profile also includes a second share of that user profile and a second share of the label associated with that user profile.

[0172] Computational systems MPC1 and MPC2 can exchange lists of row identifiers (row IDs) and Hamming distance pairs with each other. Subsequently, each computational system MPC1 and MPC2 can independently select k nearest neighbors using the same algorithm and input data. For example, computational system MPC1 can find common row identifiers for partial query results from both computational systems MPC1 and MPC2. For each i in the common row identifiers, computational system MPC1 calculates the combined Hamming distance d from the two partial Hamming distances. i For example, d i =d i,1 +d i,2 The computational system MPC1 can then be based on the combined Hamming distance d. i We sort the common row identifiers and select the k nearest neighbors. The row identifiers used for the k nearest neighbors can be represented as ID = {id1, ..., id2}. k It can be proven that if α is large enough, the k nearest neighbors determined in the above algorithm are, with a high probability, the true k nearest neighbors. However, a large value of α leads to high computational cost. In some implementations, MPC1 and MPC2 participate in a Private Set Intersection (PSI) algorithm to determine the common row identifier for the partial query results from both computer systems MPC1 and MPC2. Furthermore, in some implementations, MPC1 and MPC2 participate in an Enhanced Private Set Intersection (PSI) algorithm to compute d for the common row identifier for the partial query results from both computer systems MPC1 and MPC2. i =d i,1 +d i,2 And it does not reveal any content to MPC1 or MPC2 but instead uses d i The first k nearest neighbors are determined.

[0173] Make a determination (506) on whether to add the user to the user group. This determination can be made based on the k nearest neighbor profiles and their associated labels. The determination is also based on the aggregation function used and any aggregation parameters used for that aggregation function. The aggregation function can be selected based on the nature of the machine learning problem, such as binary classification, regression (e.g., using arithmetic mean or root mean square), multi-class classification, and weighted k-NN. Each method of determining whether to add a user to the user group can include different interactions between the MPC cluster 130 and the application 112 running on the client 110, as described in more detail below. Adding different users to a common user group can advantageously ensure that users with similar demographics are grouped together.

[0174] If a determination is made not to add the user to the user group, application 112 may not add the user to the user group (508). If a determination is made to add the user to the user group, application 112 may add the user to the user group, for example, by updating the list of user groups stored at client device 110 to include the user group identifier of the user group (510).

[0175] As noted above, application 112 can submit an inference request to MPC cluster 130, which in response can query a machine learning model (e.g., a k-NN model) to send inference results that indicate whether application 112 should add the user to zero or more user groups. In some implementations, if the inference results indicate that application 112 should add the user to multiple user groups, MPC cluster 130 can prevent the user from being added to the opposite group (i.e., classified in the opposite group) before generating or transmitting the inference results (e.g., if the user is self-declared or identified as 30-35 years old, the user can be excluded from the user group for 50-60 years old).

[0176] Example of binary classification inference technique

[0177] For binary classification, inference requests can include threshold, L true and L false As an aggregate function parameter, the label value is a boolean, either true or false. The threshold parameter can be used to indicate whether a user is added to user group L. true The threshold percentage must be the k nearest neighbor profiles of the label that have a true value. Otherwise, the user will be added to user group L. false In one approach, if the number of nearest-neighbor user profiles with true label values ​​is greater than the product of threshold and k, then MPC cluster 130 can instruct application 112 to add the user to user group L. true (otherwise L) false However, the computing system MPC1 will learn to infer outcomes, such as the user groups a user should join.

[0178] To protect user privacy, the inference request can include a threshold of plaintext, the first share [L] for the computing system MPC1. true,1 ] and [L false,1 ], and the second share of encryption for the computing system MPC2, PubKeyEncrypt([L true,2 ]||[L false,2 ]||application_public_key,MPC2). In this example, application 112 can obtain from [L true,2 ]、[L fasle,2The public key of application 112 is used to generate a composite message, as represented by the symbol ||, and this composite message is encrypted using the public key of computing system MPC2. The inference response from computing system MPC1 to application 112 can include a first share [L] of the inference result determined by computing system MPC1. result,1 ] and the second share of the inference result determined by the computing system MPC2 [L result,2 ].

[0179] To prevent the second share from being accessed by computing system MPC1 and thus enable computing system MPC1 to obtain the inference result in plaintext, computing system MPC2 can send the second share of the inference result [L] to computing system MPC1. result,2 An encrypted (and optionally digitally signed) version of ], for example, PubKeySign(PubKeyEncrypt([L result,2 [,application_public_key),MPC2), to be included in the inference response sent to application 112. In this example, application 112 is able to verify the digital signature using the public key of computing system MPC2 corresponding to the private key of computing system MPC2 used to generate the digital signature, and using the second share [L] used for the inference result. result,2 The public key (application_public_key) used for encryption corresponds to the private key of application 112 to infer the second share of the result [L]. result,2 Decryption is performed.

[0180] Application 112 can then be used to obtain the first share [L] result,1 ] and second share [L result,2 Reconstructing the inference result L result Using digital signatures enables application 112 to detect forgery of results from computing system MPC2, for example, by computing system MPC1. Depending on the desired security level, the computing system operating MPC cluster 130, and the assumed security model, digital signatures may not be required.

[0181] Computational systems MPC1 and MPC2 are able to use MPC technology to determine the share of binary classification results [L] result,1 ] and [L result,2 In binary classification, the value of label1 used for a user profile is either zero (false) or one (true). Assume the selected k nearest neighbors are identified by the identifiers {id1, ..., id1}. k The identification and computation systems MPC1 and MPC2 can calculate the sum of labels (sum_of_labels) for the profiles of the k nearest neighbors, where the sum is represented by the following relation 3:

[0182] Relationship 3: sum_of_labels = ∑ i∈{id1,...idk} label i

[0183] To determine the sum, the computation system MPC1 will use ID (i.e., {id1, ... id1}). k The data is sent to the computing system MPC2. MPC2 can verify that the number of row identifiers in the ID is greater than a threshold to enforce k-anonymity. MPC2 can then use the following relation 4 to calculate the second share of the sum of labels [sum_of_labels2]:

[0184] Relation 4: [sum_of_labels2]=∑i∈ {id1,...iak} [label i,2 ]

[0185] The computing system MPC1 can also use the following relation 5 to calculate the first share of the sum of labels [sum_of_labels1]:

[0186] Relation 5: [sum_of_labels1] = ∑ i∈{id1,...idk} [label i,1 ]

[0187] If the sum of labels, sum_of_labels, is confidential information that computation systems MPC1 and MPC2 should know as little as possible, then MPC1 and MPC2 can execute a cryptographic protocol to calculate whether sum_of_labels1 < threshold × k. That is, computation system MPC1 can calculate whether the first share of the sum of labels, [sum_of_labels1], is below the threshold, for example, [below_threshold1] = [sum_of_labels1] < threshold × k. Similarly, computation system MPC2 can calculate whether the second share of the sum of labels, [sum_of_labels2], is below the threshold, for example, [below_threshold2] = [sum_of_labels2] < threshold × k. Computation system MPC1 can then continue by multiplying [below_threshold1] by [L... false,1 ]+(1-[below_threshold1])×[L true,1 To calculate the inference result [L] result,1 Similarly, the computing system MPC2 can achieve [below_threshold2]×[L] false,2 ]+(1-[below_threshold2])×[L true,2 To calculate [L] result,2].

[0188] If the sum of labels, sum_of_labels, is not confidential information, computation systems MPC1 and MPC2 can reconstruct sum_of_labels from [sum_of_labels1] and [sum_of_labels2]. Computation systems MPC1 and MPC2 can then set the parameter below_threshold to sum_of_labels < threshold × k, for example, a value of one if it is below the threshold or zero if it is not below the threshold.

[0189] After calculating the parameter below_threshold, computational systems MPC1 and MPC2 can continue to determine the inference result L. result For example, the computing system MPC2 can determine [L] based on the value of below_threshold. result,2 ] set to [L true,2 ] or [L false,2 For example, the computing system MPC2 can calculate [L] if the sum of the labels is not lower than a threshold. result,2 ] set to [L true,2 Or, if the sum of the labels is below a threshold, [L] result,2 ] set to [L false,2 The computing system MPC2 is then able to encrypt the second share of the inference result (PubKeyEncrypt([L...)). result,2 The result, or a digitally signed version thereof, is returned to the computing system MPC1.

[0190] Similarly, the computing system MPC1 can determine [L] based on the value of below_threshold. result,1 ] set to [L true,1 ] or [L false,1 For example, the computing system MPC1 can calculate [L] if the sum of the labels is not lower than a threshold. result,1 ] set to [L true,1 Or, if the sum of the labels is below a threshold, [L] result,1 ] set to [L false,1 The computing system MPC1 is able to process the first share of the inference result [L]. result,1 ] and the encrypted second share of the inference result [L result,2 The inference response is transmitted to application 112. Application 112 can then calculate the inference result based on these two shares as described above.

[0191] Example of multi-class classification inference techniques

[0192] For multi-category classification, the tags associated with each user profile can be classification features. The content platform 150 can specify a lookup table that maps any possible category values ​​to corresponding user group identifiers. The lookup table can be one of the aggregate function parameters included in the inference request.

[0193] Within the k nearest neighbors found, MPC cluster 130 identifies the most frequent tag value. MPC cluster 130 is then able to find the user group identifier corresponding to the most frequent tag value in a lookup table and request application 112 to add the user to the user group corresponding to the user group identifier, for example, by adding the user group identifier to the user group list stored at client device 110.

[0194] Similar to binary classification, the hidden inference result L from the computational systems MPC1 and MPC2 can be optimized. result To do this, application 112 or content platform 150 can create their own mappings that map classification values ​​to inference results L. result Two lookup tables for the corresponding shares. For example, an application can create a lookup table that maps categorical values ​​to the first share [L]. result1 The first lookup table and mapping of categorical values ​​to the second share [L] result2 The second lookup table is used. An inference request from the application to computing system MPC1 can include a plaintext first lookup table for computing system MPC1 and an encrypted version of a second lookup table for computing system MPC2. The second lookup table can be encrypted using the public key of computing system MPC2. For example, a composite message including the second lookup table and the application's public key can be encrypted using the public key of computing system MPC2, e.g., PubKeyEncrypt(lookuptable2||application_public_key,MPC2).

[0195] The inference response sent by the computing system MPC1 can include the first share [L] of the inference result generated by the computing system MPC1. result1 Similar to binary classification, to prevent the second share from being accessed by computational system MPC1 and thus enable computational system MPC1 to obtain the inference result in plaintext, computational system MPC2 can send the second share of the inference result [L] to computational system MPC1. result,2 The encrypted (and optionally digitally signed) version of ] (e.g., PubKeySign(PubKeyEncrypt([L result,2 [,application_public_key),MPC2)) to be included in the inference results sent to application 112. Application 112 can obtain from [L result1 ] and [L result2 Reconstructing the inference result Lresult .

[0196] Suppose that for a multi-class classification problem there exist w valid labels {l1, l2, ... ln}. w In order to determine the inference result L in multi-class classification. result share [L] result1 ] and [L result2 The computing system MPC1 will use ID (i.e., {id1, ... id1}) to calculate ID. k}) is sent to the computing system MPC2. The computing system MPC2 can verify that the number of row identifiers in the ID is greater than a threshold to enforce k-anonymity. In general, k in k-NN can be significantly larger than k in k-anonymity. The computing system MPC2 can then compute the j-th label [l] defined using the following relation 6. j,2 The second frequency component j,2 ].

[0197] Relation 6:

[0198] Similarly, the computational system MPC1 computes the j-th label defined using the following relation 7 [l] j,1 The first frequency share j,1 ].

[0199] Relation 7:

[0200] Assume the frequency of the tag within the k nearest neighbors. i The frequency is insensitive, and the computing systems MPC1 and MPC2 are able to obtain the two shares [frequency] for this label. i,1 ] and [frequency] i,2 Reconstructing frequency i The computational systems MPC1 and MPC2 are then able to determine the frequency. index The index parameter with the maximum value, for example, index = argmax. i (frequency i ).

[0201] The computing system MPC2 can then look up the share [L] corresponding to the tag with the highest frequency in its lookup table. result,2 And PubKeyEncrypt([L result,2 The application_public_key is returned to the computing system MPC1. The computing system MPC1 can then similarly look up the share [L] corresponding to the tag with the highest frequency in its lookup table.result,1 The computing system MPC1 is then able to send to application 112 a message including two shares (e.g., [L]). result,1 ] and PubKeyEncrypt([L result,2 The inferred response of application_public_key). As described above, the second share can be digitally signed by MPC2 to prevent computing system MPC1 from forging the response of computing system MPC2. Application 112 can then calculate the inferred result based on these two shares as described above, and add the user to the user group identified by the inferred result.

[0202] Example Regression Inference Techniques

[0203] For regression, the labels associated with each user profile P must be numerical. Content platform 150 can specify an ordered list of thresholds, for example, (-∞ < t0 < t1 < ... < t...). n <∞), and a list of user group identifiers, for example, {L0, L1, ... L n L n+1 Additionally, the content platform 150 can specify aggregate functions, such as arithmetic mean or root mean square.

[0204] Within the k nearest neighbors found, MPC cluster 130 calculates the mean of the label values ​​(result) and then uses the result to look up the mapping to find the inference result L. result For example, MPC cluster 130 can use the following relation 8 to identify labels based on the mean of label values:

[0205] Relation 8:

[0206] If result ≤ t0, then L result ←L0;

[0207] If result > t n Then L result ←L n+1 ;

[0208] If t x <result≤t x+1 Then L result ←L x+1

[0209] In other words, if the result is less than or equal to the threshold t o The inference result L result It is L0. If the result is greater than the threshold t n The inference result L result It is L n+1 Otherwise, if the result is greater than the threshold tx And less than or equal to the threshold t x+1 The inference result L result It is L x+1 The computing system MPC1 then requests application 112 to add the user to the inference result L. result The corresponding user group, for example, by sending the inference result L to application 112 result The inferred response.

[0210] Similar to the other classification techniques mentioned above, it is possible to hide the inference result L from the computing systems MPC1 and MPC2. result To do this, the inference request from application 112 can include the first share [L] of the label for computing system MPC1. i,1 ] and the encrypted second share of the tag for the computing system MPC2 [L i,2 For example, PubKeyEncrypt([L 0,2 ||…||L n+1,2 ||application_public_key,MPC2)).

[0211] The inference results sent by the computing system MPC1 can include the first share [L] of the inference results generated by the computing system MPC1. result1 Similar to binary classification, to prevent the second share from being accessed by computational system MPC1 and thus enable computational system MPC1 to obtain the inference result in plaintext, computational system MPC2 can send the second share of the inference result [L] to computational system MPC1. result,2 The encrypted (and optionally digitally signed) version of ] (e.g., PubKeySign(PubKeyEncrypt([L result,2 [,application_public_key),MPC2)) to be included in the inference results sent to application 112. Application 112 can obtain from [L result,1 ] and [L result,2 Reconstructing the inference result L result .

[0212] When the aggregation function is the arithmetic mean, similar to binary classification, computational systems MPC1 and MPC2 calculate the sum of labels, sum_of_labels. If the sum of labels is insensitive, computational systems MPC1 and MPC2 can calculate two shares, [sum_of_labels1] and [sum_of_labels2], and then reconstruct sum_of_labels based on these two shares. Computational systems MPC1 and MPC2 can then calculate the average of the labels by dividing the sum of labels by the number of nearest neighbor labels (e.g., dividing by k).

[0213] The calculation system MPC1 can then use relation 8 to compare the average value with a threshold to identify the first share of the label corresponding to the average value and assign the first share [L] to the threshold. result,1 ] is set as the first share of the identified label. Similarly, the calculation system MPC2 is able to use relation 8 to compare the average value with a threshold to identify the second share of the label corresponding to the average value and set the second share [L] as the first share of the identified label. result,2 The second share is set as the identifier label. The computing system MPC2 can use the public key of application 112 to access the second share [L]. result,2 Encryption is performed using, for example, PubKeyEncrypt([L result,2 The application 112 sends the first share and the encrypted second share (which can optionally be digitally signed as described above) to the computing system MPC1. The application 112 can then add the user via a tag (e.g., a user group identifier). result Identified user groups.

[0214] If the sum of labels is sensitive, computation systems MPC1 and MPC2 may not be able to construct `sum_of_labels` in plaintext. Instead, computation system MPC1 can target the sum of labels. Calculate the mask [mask] i,1 = [sum_of_labels1] <t i ×k. This calculation may require one or more round trips between computing systems MPC1 and MPC2. Next, computing system MPC1 can calculate... Furthermore, the computing system MPC2 is capable of calculating... The equality test in this operation requires calculating multiple round trips between systems MPC1 and MPC2.

[0215] In addition, the computing system MPC1 is capable of calculating Furthermore, the computing system MPC2 is capable of calculating... for If and only if acc i When the value is 1, MPC cluster 130 will return L. i However, if use_default == 1, it will return L. n+1 This condition can be expressed in the following relation 9.

[0216] Relation 9:

[0217] The corresponding cryptographic implementation can be represented by the following relations 10 and 11.

[0218] Relation 10:

[0219] Relation 11:

[0220] These calculations are in L i In the case of plaintext, no round-trip computation between systems MPC1 and MPC2 is required, while in L... i In the case of a secret share, a round trip computation is required. The computation system MPC1 is able to process the results for both shares (e.g., [L]). result,1 ] and [L result,2 The second share is provided to application 112, wherein the second share is encrypted by MPC2 as described above and optionally digitally signed. In this way, application 112 is able to determine the inferred result L without the computational systems MPC1 or MPC2 learning anything about the intermediate or final result. result .

[0221] For the root mean square, the calculation system MPC1 will use ID (i.e., {id1, ... id1}). k The result is sent to the computation system MPC2. MPC2 is able to verify that the number of row identifiers in the ID is greater than a threshold to enforce k-anonymity. MPC2 uses the following relation 12 to calculate the second share of the sum_of_square_labels parameter (e.g., the sum of squares of the label values).

[0222] Relation 12:

[0223] Similarly, the computation system MPC1 can use the following relation 13 to calculate the first share of the sum_of_square_labels parameter.

[0224] Relation 13:

[0225] Assuming the sum_of_square_labels parameter is insensitive, computation systems MPC1 and MPC2 can reconstruct the sum_of_square_labels parameter from two shares, [sum_of_square_labels1] and [sum_of_square_labels2]. Computation systems MPC1 and MPC2 can calculate the root mean square of the labels by dividing sum_of_squares_labels by the number of nearest neighbor labels (e.g., by k) and then calculating the square root.

[0226] Regardless of whether the average is calculated via the arithmetic mean or the root mean square, the calculation system MPC1 can then use relation 8 to compare the average with a threshold to identify the label corresponding to the average and assign the first share [L]. result,1 ] is set as the identified label. Similarly, the computational system MPC2 is able to use relation 8 to compare the average with a threshold to identify the label (or the secret share of the label) corresponding to the average and set the second share [L] as the identified label. result,2 [Set as the identifier tag (secret share). The computing system MPC2 can use the public key of application 112 to access the second share [L] result,2 Encryption is performed using, for example, PubKeyEncrypt([L result,2 The application 112 sends the encrypted second share (which can optionally be digitally signed as described above) to the computing system MPC1. The computing system MPC1 can then provide the application 112 with the first share and the encrypted second share (which can optionally be digitally signed as described above) as the deduction result. The application 112 can then add the user via L... result The user group is identified by labels (e.g., user group identifiers). If the sum_of_square_labels parameter is sensitive, the computation systems MPC1 and MPC2 are able to perform cryptographic protocols similar to those used in the arithmetic mean example to calculate the share of the inferred result.

[0227] In the techniques described above used to infer results for classification and regression problems, all k nearest neighbors have equal influence on the final inference result, i.e., equal weight. For many classification and regression problems, if each of the k neighbors is assigned a value when the neighbor is associated with the query parameter P... i If the weights monotonically decrease as the Hamming distance between them increases, the model quality can be improved. The common kernel function with this property is the Epanechnikov (parabolic) kernel function. It can compute both the Hamming distance and the weights in plaintext.

[0228] Sparse Feature Vector User Profile

[0229] When features of electronic resources are included in user profiles and used to generate machine learning models, the resulting feature vectors can include high-cardinality categorical features such as domains, URLs, and IP addresses. These feature vectors are sparse, with most elements having zero values. Application 112 can split the feature vectors into two or more dense feature vectors, but this would consume too much upload bandwidth from client devices, making it impractical for machine learning platforms. To prevent this problem, the aforementioned systems and techniques can be adapted to better handle sparse feature vectors.

[0230] When providing a feature vector for an event to a client device, computer-readable code (e.g., a script) of the content platform 150 included in the electronic resource can invoke an application (e.g., a browser) API to specify the feature vector for the event. This code or the content platform 150 can determine whether the feature vector (or a portion thereof) is dense or sparse. If the feature vector (or a portion thereof) is dense, the code can be passed as a vector of values ​​to the API. If the feature vector (or a portion thereof) is sparse, the code can be passed a mapping, such as index key / value pairs for feature elements with non-zero feature values, where the key is the name or index of such feature element. If the feature vector (or a portion thereof) is sparse, and the non-zero feature values ​​are always the same, such as 1, the code can be passed a set whose elements are the names or indices of such feature elements.

[0231] When aggregating feature vectors to generate user profiles, application 112 can handle dense and sparse feature vectors differently. A user profile (or a portion thereof) computed from a dense vector remains a dense vector. A user profile (or a portion thereof) computed from a mapping remains a mapping until the fill rate is high enough that the mapping no longer saves storage costs. At that point, application 112 transforms the sparse vector representation into a dense vector representation.

[0232] If the aggregation function is a summation, then the user profile (or a portion thereof) computed from the set can be a mapping. For example, each feature vector can have a categorical feature "domains visited". The aggregation function, i.e., summation, will count the number of times the user visited the publisher domain. If the aggregation function is a logical OR, then the user profile (or a portion thereof) computed from the set can still be a set. For example, each feature vector can have a categorical feature "domains visited". The aggregation function, i.e., logical OR, will count all publisher domains visited by the user, regardless of the frequency of visits.

[0233] To send user profiles to MPC cluster 130 for ML training and prediction, application 112 can leverage any standard cryptographic library that supports secret sharing to split the dense portion of the user profile. To split the sparse portion of the user profile without significantly increasing upload bandwidth and computational costs on client devices, functional secret sharing (FSS) can be used. In this example, content platform 150 sequentially assigns unique indices, starting from 1, to each possible element in the sparse portion of the user profile. It is assumed that the valid range of the indices is inclusive within the range [1, N].

[0234] For a user profile with a non-zero value P calculated by the application i For the i-th element, 1≤i≤N, applying 112 can create two pseudo-random functions (PRFs) g with the following properties.i and h i :

[0235] For any j, g i (j)+h i (j) = 0, where 1 ≤ j ≤ N and j ≠ i

[0236] Otherwise g i (j)+h i (j)=P i .

[0237] Using FSS, g i or h i It can be concisely represented, for example, by log2(N) × size_of_tag bits, and it is impossible to represent it from g i or h i Inferring i or P i To prevent brute-force attacks, `size_of_tag` is typically 96 bits or more. In the N dimensions, it is assumed that there are n dimensions with non-zero values, where n << N. For each of the n dimensions, application 112 can construct two pseudo-random functions g and h as described above. Furthermore, application 112 can wrap the simplified representations of all n functions g into a vector G, and in the same order, wrap the simplified representations of the n functions h into another vector H.

[0238] Furthermore, application 112 can split the dense portion of the user profile P into two additive secret shares [P1] and [P2]. ​​Application 112 can then send [P1] and G to computing system MPC1 and [P2] and H to MPC2. Transmitting G requires log2(N) × size_of_tag = n × log2(N) × size_of_tag bits, which can be much smaller than the N bits required when n << N, in the case where application 112 transmits the sparse portion of the user profile in a dense vector.

[0239] When computational system MPC1 receives g1 and computational system MPC2 receives h1, the two computational systems MPC1 and MPC2 can independently create Shamir's secret share. For any j where 1 ≤ j ≤ N, computational system MPC1 is in the two-dimensional coordinates [1, 2×g...]. i A point is created on [j], and the system MPC2 is computed in two-dimensional coordinates [-1, 2×h]. i (j)] Create a point. If the two computing systems MPC1 and MPC2 cooperate to construct a line y = a0 + a1 × x through the two points, then relations 14 and 15 are formed.

[0240] Relation 14: 2×g i(j)=a0+a1

[0241] Relation 15: 2×h i (j)=a0-a1

[0242] If we add these two relations together, we get 2×g. i (j)+2×h i (j)=(a0+a1)+(a0-a1), which simplifies to a0=g i (j)+h i (j). Therefore, [1, 2×g i (j)] and [-1, 2×h i [j] is the i-th non-zero element in the sparse array (i.e., P). i Two secret shares.

[0243] During the random projection operation in the machine learning training process, the computational system MPC1 is able to independently assemble a vector of secret shares from both [P1] and G for the user profile. As described above, it is known that |G| = n, where n is the number of non-zero elements in the sparse part of the user profile. Furthermore, it is known that the sparse part of the user profile is N-dimensional, where n < 1 / 2. <N。

[0244] Assume G = {g1, ... g} n For the j-th dimension where 1≤j≤N and 1≤k≤n, let Similarly, let H = {h1,…h} n The MPC2 computing system can independently compute... It is easy to prove [SP] j,1 ] and [SP j,2 ] is SP j The additive secret share is the secret value of the j-th element in the original sparse part of the user profile.

[0245] Let [SP1] = {[SP 1,1 ],…[SP N,1 That is, the reconstructed secret share in the dense representation of the sparse part of the user profile. By cascading [P1] and [SP1], computational system MPC1 can reconstruct the complete secret share of the original user profile. Computational system MPC1 can then randomly project [P1]||[SP1]. Similarly, computational system MPC2 can randomly project [P2]||[SP2]. After projection, the above technique can be used to generate machine learning models in a similar manner.

[0246] Figure 6This is a conceptual diagram of an exemplary framework for generating inference results for a user profile in system 600. More specifically, the diagram depicts the stochastic projection logic 610, the first machine learning model 620, and the final result computation logic 640 that collectively constitute system 600. In some implementations, the functionality of system 600 can be provided in a secure and distributed manner through multiple computing systems in an MPC cluster. For example, the techniques described with reference to system 600 can be similar to those described above. Figure 2-5 The described technology. For example, the functionality associated with the random projection logic 610 can correspond to the above reference. Figure 2 and Figure 4 The functionality of one or more random projection techniques is described. Similarly, in some examples, the first machine learning model 620 may correspond to the above reference. Figure 2 , Figure 4 and Figure 5 One or more of the described machine learning models, such as those described above with reference to steps 214, 414, and 504. In some examples, an encrypted label dataset 626, which may be maintained and utilized by the first machine learning model 620 and stored in one or more memory units, can include at least one real label for each user profile used in the process of generating or training or evaluating the quality of training or fine-tuning the training of the first machine learning model 620, such as those described above with reference to... Figure 5 The encrypted label dataset 626 can include those associated with the k nearest neighbor profiles as described in step 506. That is, the encrypted label dataset 626 can include at least one true label for each of the n user profiles, where n is the total number of user profiles used to train the first machine learning model 620. For example, the encrypted label dataset 626 can include the j-th user profile (P) among the n user profiles. j At least one true label (L) j ), for the k-th user profile (P) out of n user profiles k At least one true label (L) k ), for the l-th user profile (P) among n user profiles l At least one true label (L) l ), where 1≤j, k, l≤n, and so on. Such real labels, if associated with the user profile used to generate or train the first machine learning model 620 and included as part of the encrypted label dataset 626, can be encrypted, for example, represented as secret shares. Additionally, in some examples, the final result computation logic 640 may correspond to performing one or more operations for generating the inference result (as referenced above). Figure 2The logic employed in one or more of the operations described in step 218. The first machine learning model 620 and the final result calculation logic 640 can be configured to employ one or more inference techniques, including binary classification, regression, and / or multi-class classification techniques.

[0247] exist Figure 6 In the example, system 600 is depicted as performing one or more operations at inference time. Random projection logic 610 can be used to apply user profile 609 (P) i Apply random projection transformation to obtain the transformed user profile 619(P) i The transformed user profile 619 obtained by employing random projection logic 610 can be in plaintext. For example, random projection logic 610 can be employed at least partially to obfuscate feature vectors with random noise, such as feature vectors included or indicated in user profile 609 and other user profiles, to protect user privacy.

[0248] Able to train and subsequently utilize a first machine learning model 620 to receive a transformed user profile 619 as input and generate at least one predicted label 629 in response to it. At least one predicted label 629 obtained using the first machine learning model 620 can be encrypted. In some implementations, the first machine learning model 620 includes a k-nearest neighbor (k-NN) model 622 and a label predictor 624. In such implementations, the k-NN model 622 can be employed by the first machine learning model 620 to identify k nearest neighbor user profiles considered most similar to the transformed user profile 619. In some examples, models other than the k-NN model, such as those rooted in one or more prototype methods, can be employed as model 622. The label predictor 624 is then able to identify the true label for each of the k nearest neighbor user profiles from the true labels included in the encrypted label dataset 626 and determine at least one predicted label 629 based on the identified label. In some implementations, the label predictor 624 can apply a softmax function to the data it receives and / or generates in determining at least one predicted label 629.

[0249] In an implementation where the first machine learning model 620 and the final result calculation logic 640 are configured to employ regression techniques, at least one predicted label 629 may correspond to, for example, a single label representing, as such, the sum of the true labels for the k nearest neighbor user profiles as determined by the label predictor 624. Such a sum of true labels for the k nearest neighbor user profiles, as determined by the label predictor 624, is essentially equivalent to, for example, the average of the true labels for the k nearest neighbor user profiles scaled by a factor of k. Similarly, in an implementation where the first machine learning model 620 and the final result calculation logic 640 are configured to employ binary classification techniques, at least one predicted label 629 may correspond to, for example, a single label representing, as such, an integer determined by the label predictor 624 at least partially based on such a sum. In the case of binary classification, each of the true labels of the k nearest neighbor user profiles can be a binary value of zero or one, such that the aforementioned average can be an integer value between zero and one (e.g., 0.3, 0.8, etc.), which, for example, effectively represents the predicted probability that the true label of the user profile (e.g., the transformed user profile 619) received as input by the first machine learning model 620 is equal to one. See below for reference. Figures 9 to 11 Additional details are provided regarding the nature of at least one predicted label 629 and the implementation in which the first machine learning model 620 and the final result calculation logic 640 are configured to employ regression techniques and the implementation in which the first machine learning model 620 and the final result calculation logic 640 are configured to employ binary classification techniques to determine at least one predicted label 629.

[0250] In the implementation of a multi-class classification technique, where the first machine learning model 620 and the final result calculation logic 640 are configured, at least one predicted label 629 can correspond to a vector or set of predicted labels as determined by the label predictor 624. Each predicted label in such a vector or set of predicted labels can correspond to a corresponding class and is determined by the label predictor 624 at least in part based on majority voting or the frequency of true labels in the vector or set corresponding to the true labels for user profiles in the k nearest neighbor user profiles, where the corresponding class is a first value (e.g., one) as determined by the label predictor 624. Similar to binary classification, in the case of multi-class classification, each true label in each vector or set of true labels for user profiles in the k nearest neighbor user profiles can be a binary value of zero or one. See below for reference. Figures 9 to 11 Additional details are provided regarding the nature of at least one predicted label 629 and the manner in which at least one predicted label 629 can be determined for an implementation in which the first machine learning model 620 and the final result computation logic 640 are configured to employ a multi-class classification technique.

[0251] The final result calculation logic 640 can be used to generate an inference result 649 based on at least one predicted label 629. i For example, the final result calculation logic 640 can be used to evaluate at least one predicted label 629 against one or more thresholds and determine an inference result 649 based on the evaluation results. In some examples, the inference result 649 may indicate whether a user associated with user profile 609 will be added to one or more user groups. In some implementations, at least one predicted label 629 may be included or otherwise indicated in the inference result 649.

[0252] In some implementations, such as Figure 6 The system 600 described herein is capable of representing, for example, systems such as Figure 1 The system is implemented using an MPC cluster of 130. Therefore, it should be understood that in at least some of these implementations, two or more computing systems within an MPC cluster can provide secure and distributed computing capabilities as described in this paper. Figure 6 The elements shown describe some or all of the functionality. For example, each of the two or more computing systems in an MPC cluster can provide a reference to this document. Figure 6 The corresponding functional share is described. In this example, two or more computing systems can operate in parallel to implement the chosen secret-sharing algorithm in order to collaboratively execute the algorithm referenced herein. Figure 6 The similar or equivalent operations described. In at least some of the foregoing implementations, user profile 609 may represent a share of the user profile. In such implementations, this document refers to... Figure 6 One or more of the other data or quantities described may also represent their secret share. It should be understood that, in providing this document for reference... Figure 6 The described functionality allows for additional operations to be performed by two or more computing systems for the purpose of protecting user privacy. See the example below for reference. Figure 12 Examples of one or more of the aforementioned implementations are described in further detail elsewhere in this document. In general, in at least some implementations, the term "share," as described below and elsewhere in this document, may correspond to a secret share.

[0253] While the training process for k-NN models such as k-NN model 622 can be relatively fast and simple because it does not require label knowledge, there is room for improvement in the quality of such models in some cases. Therefore, in some implementations, one or more systems and techniques, described in further detail below, can be used to improve the performance of the first machine learning model 620.

[0254] Figure 7 This is a conceptual diagram of an exemplary framework for generating inference results for user profiles in System 700 with improved performance. In some implementations, such as... Figure 7 One or more of the elements 609-629 described in the reference above can be respectively associated with the elements described in the reference above. Figure 6 One or more of the elements 609-629 described are similar or equivalent. Much like system 600, system 700 includes stochastic projection logic 610 and a first machine learning model 620, and is described as performing one or more operations at inference time.

[0255] However, unlike system 600, system 700 further includes a second machine learning model 730, which is trained and subsequently used to receive a transformed user profile 619 as input and generate prediction residual values ​​739 indicating the amount of prediction error in at least one prediction label 629. i The output is used to improve the performance of the first machine learning model 620. The predicted residual value 739 obtained using the second machine learning model 730 can be in plaintext. The final result calculation logic 740, included in system 700 instead of the final result calculation logic 640, can be used to generate an inference result 749 based on at least one predicted label 629 and further based on the predicted residual value 739. i Given that the prediction residual value 739 indicates the amount of prediction error in at least one prediction label 629, relying on at least one prediction label 629 and cooperating with the prediction residual value 739 can enable the final result calculation logic 740 to effectively cancel or neutralize at least some of the errors that can be expressed in at least one prediction label 629, thereby enhancing one or both of the accuracy and reliability of the inference result 749 generated by the system 700.

[0256] For example, the final result calculation logic 740 can be used to calculate the sum of at least one predicted label 629 and the predicted residual value 739. In some examples, the final result calculation logic 740 can be further used to evaluate such a calculated sum against one or more thresholds and to determine the inference result 749 based on the evaluation result. In some implementations, it is possible to... Figure 6 The inference result 649 or Figure 7 The inference result 749 includes, or otherwise indicates, a sum of such calculations of at least one prediction label 629 and prediction residual value 739.

[0257] The second machine learning model 730 may include or correspond to one or more of deep neural networks (DNNs), gradient boosting decision trees, and random forest models. That is, the first machine learning model 620 and the second machine learning model 730 may be architecturally different from each other. In some implementations, the second machine learning model 730 can be trained using one or more gradient boosting algorithms, one or more gradient descent algorithms, or a combination thereof.

[0258] The second machine learning model 730 can be trained using the same set of user profiles used to train the first machine learning model 620 and data indicating the difference between the true labels for such a set of user profiles and the predicted labels for such a set of user profiles as determined by the first machine learning model 620. Therefore, the process of training the second machine learning model 730 is performed after at least a portion of the process of training the first machine learning model 620 has been executed. Data for training the second machine learning model 730, such as data indicating the difference between the predicted labels and the true labels determined using the first machine learning model 620, can be generated or otherwise obtained by evaluating the performance of the first machine learning model 620. See below for reference. Figures 10 to 11 An example of such a process is described in further detail.

[0259] As mentioned above, at least in part, random projection logic 610, included in systems 600 and 700, can be used to obfuscate feature vectors with random noise, such as those included or indicated in user profiles 609 and other user profiles, to protect user privacy. To enable machine learning training and prediction, the random projection transformation applied through random projection logic 610 needs to preserve some concept of the distance between feature vectors. An example of a random projection technique that can be employed in random projection logic 610 includes the SimHash technique. This technique and the other techniques described above can be used to obfuscate feature vectors while preserving the cosine distance between such feature vectors.

[0260] While preserving the cosine distance between feature vectors may prove sufficient for k-NN models trained and using, such as the first machine learning model 620, it may be less ideal for other types of models trained and using, such as the second machine learning model 730. Therefore, in some implementations, it may be desirable to employ a random projection technique in the random projection logic 610 that can be used to obfuscate feature vectors while preserving the Euclidean distance between them. An example of such a random projection technique includes the Johnson-Lindenstrauss (JL) technique or transformation.

[0261] As mentioned above, one property of the JL transform is that it preserves the Euclidean distance between eigenvectors with probability. However, the JL transform is lossy, irreversible, and incorporates random noise. Therefore, even if two or more servers or computing systems in an MPC cluster collide, they will not be able to obtain a transformed version (P) of the user profile using the JL transform technique. i ′) Obtain the original user profile (P i The exact reconstruction of ) is described. In this way, the JL transformation technique can be used to provide user privacy protection for the purpose of transforming user profiles in one or more systems described herein. Furthermore, the JL transformation technique can be used as a dimensionality reduction technique. Therefore, an advantageous byproduct of using the JL transformation technique for transforming user profiles in one or more systems described herein is that it can be used to significantly improve the speed at which subsequent processing steps can be performed by such systems.

[0262] In general, given arbitrarily small ε>0, there exists a way to apply P to any 1≤i,j≤n. i Transform into P i ′、P j Transform into P j The JL transformation of ′, where n is the number of training examples, and:

[0263] (1-ε)×|P i -P j | 2 ≤|P′ i -P′ j | 2 ≤(1+ε)×|P i -P j | 2

[0264] In other words, applying the JL transform can change the Euclidean distance between two arbitrarily chosen training examples by no more than a small fraction ε. For at least the reasons mentioned above, the JL transform technique can be employed in some implementations, such as the random projection logic 610 described herein.

[0265] In some implementations, such as Figure 7 The system 700 described in the text is capable of representing, for example, systems composed of, etc. Figure 1 The system is implemented using an MPC cluster of 130. Therefore, it should be understood that in at least some of these implementations, two or more computing systems within an MPC cluster can provide secure and distributed computing capabilities as described in this paper. Figure 7 The elements shown describe some or all of the functionality. For example, each of the two or more computing systems in an MPC cluster can provide a reference to this document. Figure 7The corresponding functional share is described. In this example, two or more computing systems can operate in parallel to implement the chosen secret-sharing algorithm in order to collaboratively execute the algorithm referenced herein. Figure 7 The similar or equivalent operations described. In at least some of the foregoing implementations, user profile 609 can represent a secret share of the user profile. In such implementations, this document refers to... Figure 7 One or more of the other data or quantities described may also represent their secret share. It should be understood that, in providing this document for reference... Figure 7 The described functionality allows for additional operations to be performed by two or more computing systems for the purpose of protecting user privacy. See the example below for reference. Figure 12 Furthermore, examples of one or more of the aforementioned implementations are described in further detail elsewhere in this document.

[0266] Figure 8 This is a flowchart illustrating an example process 800 for generating inference results for a user profile at an MPC cluster with improved performance. For example, a reference can be executed at inference time. Figure 8 One or more of the operations described. The operation of process 800 can be performed, for example, by, such as Figure 1 The MPC cluster implementation of MPC cluster 130, and can also correspond to the above reference. Figure 7 One or more of the operations described. For example, a reference operation can be performed at inferred time. Figure 8 The described operation.

[0267] In some implementations, it can be achieved through, for example Figure 1 The MPC cluster 130 provides two or more computing systems in a secure and distributed manner as described in this article. Figure 8 The elements shown describe some or all of the functionality. For example, each of the two or more computing systems in an MPC cluster can provide a reference to this document. Figure 8 The corresponding share of the described functionality. In this example, two or more computing systems can operate in parallel to implement the selected secret-sharing algorithm in order to collaboratively perform the functions described herein. Figure 8 The description refers to similar or equivalent operations. It should be understood that this document is provided for reference only. Figure 8 The described functionality allows for additional operations to be performed by two or more computing systems for the purpose of protecting user privacy. See the example below for reference. Figure 12Examples of one or more of the aforementioned implementations are further described in detail elsewhere in this document. The operation of process 800 can also be implemented as instructions stored on one or more computer-readable media that may be non-transitory, and execution of the instructions by one or more data processing devices can cause the one or more data processing devices to perform the operation of process 800.

[0268] The MPC cluster receives an inference request associated with a specific user profile (802). For example, this could correspond to one or more operations similar to or equivalent to those performed by the MPC cluster 130 upon receiving an inference request from application 112, as referenced above. Figure 1 As described.

[0269] The MPC cluster determines a predicted label for a specific user profile (804) based on a specific user profile, a first machine learning model trained using multiple user profiles, and one or more of multiple real labels for the multiple user profiles. The label, as described herein, can be or include demographic user group identifiers or demographic features. For example, this could correspond to obtaining at least one predicted label 629 using the first machine learning model 620. One or more operations performed are similar to or equivalent to one or more operations, as referenced above. Figure 6-7 As described.

[0270] In this example, the multiple ground truth labels for multiple user profiles may correspond to ground truth labels included as a portion of encrypted label data 626, which are ground truth labels for the multiple user profiles used to train the first machine learning model 620. For example, one or more ground truth labels from the multiple ground truth labels on which the predicted label for a particular user profile is based may include at least one ground truth label for each of the k nearest neighbor user profiles identified by the k-NN model 622 of the first machine learning model 620. In some examples, each of the multiple ground truth labels is encrypted, as in Figures 6 to 7 The situation is similar to that in the example above. Several of the various ways to determine predicted labels using the true labels of the k nearest neighbor user profiles have been described in detail above. As will become apparent above, the way or method of using such true labels to determine predicted labels can depend at least in part on the type of inference technique employed (e.g., regression techniques, binary classification techniques, multi-class classification techniques, etc.).

[0271] The MPC cluster determines prediction residual values ​​(806) indicative of prediction errors in the predicted labels based on specific user profiles, a second machine learning model trained using multiple user profiles, and data indicating the differences between multiple true labels for multiple users and multiple predicted labels determined using a first machine learning model for multiple user profiles. For example, this could correspond to obtaining prediction residual values ​​739 (Residue) using the second machine learning model 730. i The second machine learning model performs one or more operations similar to or equivalent to those described above with reference to Reference 7. Therefore, in some implementations, the second machine learning model includes at least one of a deep neural network, a gradient boosting decision tree, and a random forest model.

[0272] The MPC cluster generates data representing the inference results based on the predicted labels and predicted residual values ​​(808). For example, this could correspond to the generation of the inference result 749 (Result) using the final result calculation logic 740. i One or more operations performed are similar to or equivalent to one or more operations, as referred to above. Figure 7 As described. Therefore, in some examples, the inference result includes or corresponds to the sum of the predicted label and the predicted residual value.

[0273] The MPC cluster provides data representing the inference result to the client device (810). For example, this could correspond to one or more operations similar to or equivalent to those performed by the MPC cluster 130 in providing the inference result to the client device 110 running on the application 112, as referenced above. Figures 1 to 2 As described.

[0274] In some implementations, process 800 further includes one or more operations in which the MPC cluster applies a transformation to a specific user profile to obtain a transformed version of the specific user profile. In these implementations, to determine the predicted label, the MPC cluster determines the predicted label at least in part based on the transformed version of the specific user profile. For example, this could correspond to the use of random projection logic 610 to transform user profile 609(P i Apply random projection transformation to obtain the transformed user profile 619(P) i One or more operations performed are similar to or equivalent to one or more operations, as referred to above. Figures 6 to 7As described. Therefore, in some examples, the aforementioned transformation can be a random projection. Furthermore, in at least some of these examples, the aforementioned random projection can be a Johnson-Lindenstrauss (JL) transformation. In at least some of the aforementioned implementations, to determine the predicted label, the MPC cluster provides a transformed version of a specific user profile as input to a first machine learning model to obtain a predicted label for the specific user profile as output. For example, this could correspond to receiving a transformed user profile 619 (P) with respect to the first machine learning model 620. i ′) as input and in response to it generate at least one predicted label 629 One or more operations performed are similar to or equivalent to one or more operations, as referenced above. Figures 6 to 7 As described.

[0275] As mentioned above, in some implementations, the first machine learning model includes a k-nearest neighbor model. In at least some of these implementations, to determine the predicted label, the MPC cluster identifies, at least in part, a number of k nearest neighbor user profiles considered most similar to the specific user profile among multiple user profiles based on the specific user profile and the k-nearest neighbor model, and determines the predicted label at least in part based on the true labels for each of the k nearest neighbor user profiles. In some such implementations, to determine the predicted label at least in part based on the true labels for each of the k nearest neighbor user profiles, the MPC cluster determines the sum of the true labels for the k nearest neighbor user profiles. For example, this could correspond to using a first machine learning model 620 to obtain at least one predicted label 629 in one or more implementations employing one or more regression and / or binary classification techniques. One or more operations performed are similar to or equivalent to one or more operations, as referenced above. Figures 6 to 7 As described. In some examples, the predicted labels include or correspond to the sum of the true labels for the k nearest neighbor user profiles.

[0276] In some of the aforementioned implementations, in order to determine the predicted label at least partially based on the true labels for each of the k nearest neighbor user profiles, the MPC cluster determines the set of predicted labels at least partially based on the set of true labels for each of the k nearest neighbor user profiles, respectively corresponding to a set of categories, and, in order to determine the set of predicted labels, the MPC cluster performs an operation on each category in the set. Such operations can include one or more operations in which the MPC cluster determines the majority vote or frequency of true labels in the set of true labels for user profiles in the k nearest neighbor user profiles whose true labels corresponding to a category are a first value. For example, this could correspond to obtaining at least one predicted label 629 using a first machine learning model 620 in one or more implementations employing one or more multi-class classification techniques. One or more operations performed are similar to or equivalent to one or more operations, as referenced above. Figures 6 to 7 As described.

[0277] Figure 9 This is a flowchart illustrating an example process 900 for preparing and executing the training of a second machine learning model at an MPC cluster to improve inference performance. The operation of process 900 can be performed, for example, by... Figure 1 The MPC cluster implementation of MPC cluster 130, and can also correspond to the above reference. Figure 2 , Figure 4 , Figure 6 and Figure 7 One or more of the operations described. In some implementations, this can be achieved through, for example... Figure 1 The MPC cluster 130 provides two or more computing systems in a secure and distributed manner as described in this article. Figure 9 The elements shown describe some or all of the functionality. For example, each of the two or more computing systems in an MPC cluster can provide a reference to this document. Figure 9 The description describes the corresponding secret share of functionality. In this example, two or more computing systems can operate in parallel to implement the chosen secret-sharing algorithm in order to collaboratively perform the functions described herein. Figure 9 The operations described are similar to or equivalent to those described. It should be understood that this document is provided for reference only. Figure 9 The described functionality allows for additional operations to be performed by two or more computing systems for the purpose of protecting user privacy. See the example below for reference. Figure 12Examples of one or more of the aforementioned implementations are further described in detail elsewhere in this document. The operation of process 900 can also be implemented as instructions stored on one or more computer-readable media that may be non-transitory, and execution of the instructions by one or more data processing devices can cause the one or more data processing devices to perform the operation of process 900.

[0278] The MPC cluster uses multiple user profiles to train a first machine learning model (910). For example, as described above, the first machine learning model can correspond to a first machine learning model 620. Similarly, the multiple user profiles used in training the first machine learning model can correspond to the number n user profiles used to train the first machine learning model 620, and the real labels for them can be included in an encrypted label dataset 626 as described above. As described herein, the labels can be or include user group identifiers or demographic features.

[0279] The MPC cluster evaluation is exemplified by the performance of the first machine learning model trained using multiple user profiles (920). See below for reference. Figures 10 to 11 Provide additional details regarding what such an evaluation might require.

[0280] In some implementations, the data generated in such evaluations can be used by the MPC cluster or another system communicating with the MPC cluster to determine whether the performance of a first machine learning model, such as a first machine learning model 620, is guaranteed to improve, for example, through a second machine learning model, such as a second machine learning model 730. See below for reference. Figure 10 The brief and residual dataset 1070 and Figure 11 Step 1112 further describes in detail examples of data generated in such an evaluation that can be utilized in this way.

[0281] For example, in some cases, an MPC cluster or another system communicating with an MPC cluster may determine, based on data generated in such an evaluation, that the performance (e.g., prediction accuracy) of a first machine learning model meets one or more thresholds, and therefore no improvement is guaranteed. In such cases, the MPC cluster may avoid training and implementing a second machine learning model based on this determination. However, in other cases, an MPC cluster or another system communicating with an MPC cluster may determine, based on data generated in such an evaluation, that the performance (e.g., prediction accuracy) of a first machine learning model meets one or more thresholds, and therefore an improvement is guaranteed. In these cases, based on this determination, the MPC cluster may receive a functional upgrade comparable to the functional upgrade that would be obtained during the transition from system 600 to system 700, as referenced above. Figures 6 to 7As described. To receive such a functional upgrade, the MPC cluster can continue training and implementing a second machine learning model, such as second machine learning model 730, to improve the performance of the first machine learning model. In some examples, the data generated in such an evaluation can be additionally or alternatively provided to one or more entities associated with the MPC cluster. In some such examples, one or more entities can make their own determination regarding whether the performance of the first machine learning model is guaranteed to improve, and proceed accordingly. Other configurations are possible.

[0282] The MPC cluster uses a collection of data, including data generated in evaluating the performance of the first machine learning model, to train the second machine learning model (930). Examples of such data can be included in the references below. Figure 10 The brief and residual dataset 1070 and Figure 11 The data described in step 1112.

[0283] In some implementations, process 900 further includes additional steps 912-916, which are described in further detail below. In such implementations, steps 912-916 are performed before steps 920 and 930, but can be performed after step 910.

[0284] Figure 10 This is a conceptual diagram of an exemplary framework for evaluating the performance of a first machine learning model in System 1000. In some implementations, such as... Figure 10 One or more of the elements 609-629 described above can be referenced separately. Figures 6 to 7 The described one or more elements 609-629 are similar or equivalent. In some examples, this document references... Figure 10 One or more of the operations described can correspond to the above references. Figure 9 Step 920 describes one or more of those operations. Similar to systems 600 and 700, system 1000 includes random projection logic 610 and a first machine learning model 620.

[0285] However, unlike systems 600 and 700, system 1000 further includes residual calculation logic 1060. Additionally, in Figure 10 In the example, user profile 609 (P) i This corresponds to one of the multiple user profiles used to train the first machine learning model 620; however, in Figure 6 and Figure 7 In the example, user profile 609 (P) iThe user profile used to train the first machine learning model 620 may not necessarily correspond to one of the multiple user profiles used to train the first machine learning model 620, but rather simply to the user profile associated with the inference request received at inference time. In some examples, the aforementioned multiple user profiles used to train the first machine learning model 620 can correspond to the above-mentioned references. Figure 9 Step 910 describes multiple user profiles. Residual calculation logic 1060 can be used to calculate based on at least one predicted label 629 and at least one true label 1059 (L...). i To generate residual values ​​1069 (Residue) that indicate the amount of error in at least one predicted label 629. i The labels described herein can be or include demographic user group identifiers or demographic features. At least one predictive label 629 and at least one real label 1059 (L i Both can be encrypted. For example, the residual calculation logic 1060 can use a secret share to calculate the difference between the values ​​of at least one predicted tag 629 and at least one true tag 1059. In some implementations, the residual value 1069 can correspond to the difference between the aforementioned values.

[0286] The residual value 1069 can be stored in association with the transformed user profile 619, for example, stored in memory as part of the profile and residual dataset 1070. In some examples, the data included in the profile and residual dataset 1070 may correspond to the data referenced above. Figure 9 The data described in step 930 and as referenced below Figure 11 One or both of the data described in step 1112. In some implementations, the residual value 1069 is in the form of a secret share to protect user privacy and data security.

[0287] In some implementations, such as Figure 10 The system 1000 described in the text is capable of representing components such as... Figure 1 The system is implemented using an MPC cluster of 130. Therefore, it should be understood that in at least some of these implementations, two or more computing systems within an MPC cluster can provide secure and distributed computing capabilities as described in this paper. Figure 10 The elements shown describe some or all of the functionality. For example, each of the two or more computing systems in an MPC cluster can provide a reference to this document. Figure 10 The corresponding functional share is described. In this example, two or more computing systems can operate in parallel to implement the chosen secret-sharing algorithm in order to collaboratively execute the algorithm referenced herein. Figure 10The operations described are similar to or equivalent to those described above. In at least some of the aforementioned implementations, user profile 609 can represent a secret share of the user profile. In such implementations, this document refers to... Figure 10 One or more of the other data or quantities described may also represent their secret share. It should be understood that, in providing this document for reference... Figure 10 When describing functionality, additional operations can be performed by two or more computing systems for the purpose of protecting user privacy. See the example below for reference. Figure 12 Furthermore, examples of one or more of the aforementioned implementations are described in further detail elsewhere in this document.

[0288] Figure 11 This is a flowchart illustrating an example process 1100 for evaluating the performance of a first machine learning model at an MPC cluster. The operation of process 1100 can be performed, for example, by means of... Figure 1 The MPC cluster implementation of MPC cluster 130, and can also correspond to the above reference. Figures 9 to 10 One or more of the operations described. In some examples, this article references... Figure 11 One or more of the operations described can correspond to the above references. Figure 9 Step 920 describes one or more of those operations. In some implementations, this can be achieved through, for example... Figure 1 The MPC cluster 130 provides two or more computing systems in a secure and distributed manner as described in this article. Figure 11 The elements shown describe some or all of the functionality. For example, each of the two or more computing systems in an MPC cluster can provide a reference to this document. Figure 11 The corresponding functional share is described. In this example, two or more computing systems can operate in parallel to implement the chosen secret-sharing algorithm in order to collaboratively execute the algorithm referenced herein. Figure 11 The operations described are similar to or equivalent to those described. It should be understood that this document is provided for reference only. Figure 11 The described functionality allows for additional operations to be performed by two or more computing systems for the purpose of protecting user privacy. See the example below for reference. Figure 12 Examples of one or more of the aforementioned implementations are further described in detail elsewhere in this document. The operation of process 1100 can also be implemented as instructions stored on one or more computer-readable media that may be non-transitory, and execution of the instructions by one or more data processing devices can cause one or more data processing devices to perform the operation of process 1100.

[0289] The MPC cluster selects the i-th user profile and at least one corresponding real label ([P i L iThe process 1100 involves recursively incrementing i (1102-1104) until i equals n (1114-1116), where n is the total number of user profiles used to train the first machine learning model. Labels can be or include demographic-based user group identifiers or demographic features. In other words, process 1100 includes performing steps 1106-1112 for each of the n user profiles used to train the first machine learning model, as described below.

[0290] In some implementations, the i-th user profile can represent a secret share of the user profile. In such implementations, this paper references... Figure 11 One or more of the other data or quantities described may also represent their share.

[0291] The MPC cluster provides the i-th user profile (P) i Apply random projection to obtain the transformed version (P) of the i-th user profile. i (1106). For example, this could correspond to the use of random projection logic 610 to process user profile 609 (P). i Apply random projection transformation to obtain the transformed user profile 619(P) i One or more operations performed are similar to or equivalent to one or more operations, as referred to above. Figure 10 As described.

[0292] The MPC cluster will transform the version (P) of the i-th user profile. i ′) is fed as input to the first machine learning model to obtain the transformed version (P) for the i-th user profile. i At least one predicted label of ′) As output (1108). For example, this could correspond to the transformed user profile 619 (P) received with respect to the first machine learning model 620. i ′) as input and in response to it generate at least one predicted label 629 One or more operations performed are similar to or equivalent to one or more operations, as referenced above. Figure 10 As described.

[0293] MPC clusters are at least partially based on the profile of the i-th user (P i At least one true label (L) i ) and at least one prediction label To calculate the residual value i (1110). For example, this could correspond to the use of residual calculation logic 1060 to at least partially base on at least one true label 1059 (L). i) and at least one prediction label 629 To calculate the residual value 1069 (Residue) i One or more operations performed are similar to or equivalent to one or more operations, as referred to above. Figure 10 As described.

[0294] The MPC cluster will calculate the residual value (Residue) i The transformed version of the i-th user profile (P) i ') is stored in association with (1112). For example, this could correspond to the residual value 1069 (Residue). i ) and the transformed user profile 619 (P i One or more operations that are similar to or equivalent to the operations performed in association (e.g., stored in memory as part of the profile and residual dataset 1070), as referenced above. Figure 10 As described. In some examples, this data may correspond to the reference above. Figure 9 The data described in step 930. Therefore, in these examples, some or all of the data stored in this step can be used as data for training a second machine learning model, such as the second machine learning model 730.

[0295] Referring again to steps 1108-1110, for at least some implementations where the first machine learning model is configured to employ regression techniques, the MPC cluster obtains at least one predicted label at step 1108. This can correspond to a single prediction label representing an integer. In these implementations, the residual value calculated by the MPC cluster at step 1110... i ) can correspond to indicating at least one real label (L) i ) with at least one predicted label The integer difference between the values. In at least some of the aforementioned implementations, at step 1108, the first machine learning model identifies the version considered to be the transformed version (P) of the i-th user profile. i Find the k most similar nearest neighbor user profiles, identify at least one true label for each of the k nearest neighbor user profiles, calculate the sum of the true labels for the k nearest neighbor user profiles, and use this sum as at least one predicted label. As mentioned above, the sum of the true labels for the k nearest neighbor user profiles, as determined in this step, is actually equivalent to the average of the true labels for the k nearest neighbor user profiles, scaled by a factor of k. In some examples, this sum can be used as at least one predicted label. Instead of averaging the true labels for the k nearest neighbor user profiles, this eliminates the need for division. It also considers at least one predicted label. Effectively equivalent to, for example, the average of the true labels for the k nearest neighbor user profiles scaled by a factor of k, the computation performed by the MPC cluster at step 1110, for at least some implementations where the first machine learning model is configured to employ regression techniques, is given by the following formula:

[0296]

[0297] Similarly, for at least some implementations where the first machine learning model is configured to employ binary classification techniques, the MPC cluster obtains at least one predicted label at step 1108. This can correspond to, for example, a single predicted label representing an integer determined at least in part based on the sum of the true labels for the k nearest neighbor user profiles. As mentioned above in the implementation where the first machine learning model is configured to employ regression techniques, such a sum of true labels for the k nearest neighbor user profiles is actually equivalent to, for example, the average of the true labels for the k nearest neighbor user profiles scaled by a factor of k.

[0298] However, unlike the implementation where the first machine learning model is configured to use regression techniques, in the implementation where the first machine learning model is configured to use binary classification techniques, each of the true labels for the k nearest neighbor user profiles can be a binary value of zero or one, such that the aforementioned average can be an integer value between zero and one (e.g., 0.3, 0.8, etc.). Although in the implementation using binary classification techniques, the MPC cluster can calculate and use the sum of the true labels for the k nearest neighbor user profiles (sum_of_labels) as at least one predicted label at step 1108. And use the formula described above, which employs regression techniques. To obtain a mathematically feasible residual value at step 1110. i Such residual values i This could potentially raise privacy concerns, for example, later when used to determine whether improving the first machine learning model is guaranteed, or later when used to train a second machine learning model such as the second machine learning model 730. More specifically, because each of the true labels in the k nearest neighbor user profiles can be a binary value of zero or one, in implementations employing binary classification techniques, such residual values... i The symbol ) can potentially indicate at least one real label (L) iThe value of ) and therefore can potentially be indicated by the residual value (Residue) which can be processed to some extent at or after step 1112. i Inferring data from one or more systems and / or entities.

[0299] For example, consider the use of binary classification techniques and L... i =1, k=15 and The first example. In this first example, at least one predicted label. The sum of the true labels for the k nearest neighbor user profiles (sum_of_labels) is actually equivalent to the average of the true labels for the k nearest neighbor user profiles scaled by a factor of k, where the aforementioned average is a non-integer value of 0.8. This first example will utilize the same formula as described above. For example, the residual value is calculated at step 1110. i In this first example, the residual value (Residue) i The following formula will be used to give Residue: i = (15)(1)-12=3. Therefore, in this first example, the residual value (Residue) i The value will be equal to (positive) 3. Now, consider a case where a binary classification technique will be used and Li = 0, but k and A second example where the values ​​are again equal to 15 and 12. Again, the same formula as described above will be used in this second example. For example, the residual value is calculated at step 1110. i In this second example, the residual value (Residue) i The result will be given by the following formula: Residue i = (15)(0)-12=-12. Therefore, in this first example, the residual value (Residue) i The value will be equal to -12. In fact, in the first and second examples above, the positive residual value (Residue) i ) can be with L i =l correlation, while negative residual (Residue) i ) can be with L i =0 related.

[0300] To understand why from Residue i Inferring L i It is possible to consider the residuals of a user profile used to train a first machine learning model, where the residuals are assumed to have a representation where the true label is zero. The normal distribution of the prediction errors (e.g., residual values) of the true labels equal to 0 (zero) is given, where μ0 and σ0 are the mean and standard deviation of the normal distribution of the prediction errors (e.g., residual values) of the true labels equal to 0 (zero) and are associated with the user profile used to train the first machine learning model, and it is assumed that the residuals of the training examples whose labels are equal to 1 satisfy the following condition. Where μ1 and σ1 are the mean and standard deviation of the normal distribution of the prediction error of the true label equal to 1 (a), respectively, and are associated with the user profile used to train the first machine learning model. Under such assumptions, it is clear that μ0 < 0, μ1 > 0, and σ0 = σ1 is not guaranteed.

[0301] Given the foregoing, as described below, in some implementations, different methods can be employed to perform one or more operations associated with steps 1108-1110 of the implementation employing a binary classification technique. In some implementations, to force the residuals of the two classes of training examples to have the same normal distribution, the MPC cluster can apply a transformation f to the sum of the true labels (sum_of_labels) of the k nearest neighbor user profiles, such that the L... i and The calculated residual values ​​cannot be used to predict L. i The transformation f, when applied to the initial predicted labels (e.g., the sum of the true labels in binary classification, the majority vote of the true labels in multi-class classification, etc.), can be used to remove biases that may exist in the predictions of the first machine learning model. To achieve this, the transformation f needs to satisfy the following property:

[0302] (i)f(μ0)=0

[0303] (ii)f(μ1)=1

[0304] (iii)σ0×f′(μ0)=σ1×f′(μ1)

[0305] Where f′ is the derivative of f.

[0306] An example of a transformation with the above properties that can be used in this type of implementation is the shape f(x) = a²x. 2 The quadratic polynomial transformation of +a1x+a0, where f′(x)=2a2x+a1. In some examples, the MPC cluster can deterministically find the values ​​of the coefficients {a2, a1, a0} based on three linear equations from the following three constraints:

[0307] make

[0308] (i)a′2=σ0-σ1

[0309] (ii)a′1=2(σ1μ1-σ0μ0)

[0310] (iii)a′0=μ0(μ0σ0+μ0σ1-2μ1σ1)

[0311] In these examples, the MPC cluster can compute the coefficients {a2, a1, a0} as: {a2, a1, a0} = D × {a2′, a1′, a0′}. The MPC cluster can compute {a2′, a1′, a0′}, for example, using addition and multiplication operations on the secret share. The transformation f(x) = a2x 2 +a1x+a0 also revolves around: Mirror symmetry.

[0312] To calculate the aforementioned coefficients and other values ​​dependent on them, the MPC cluster can first estimate the mean and standard deviation of the probability distributions of prediction errors (e.g., residual values) for true labels equal to zero, μ0, and σ0, respectively, and the mean and standard deviation of the probability distributions of prediction errors (e.g., residual values) for true labels equal to one, μ1, and σ1, respectively. In some examples, the variance σ0 of the probability distribution of prediction errors for true labels equal to zero can be determined as a supplement to or alternative to the standard deviation σ0. 2 Furthermore, as a supplement or substitute for the standard deviation σ1, the variance σ1 of the probability distribution of the prediction error of the true label, which is equal to one, can be determined. 2 .

[0313] In some instances, the given probability distribution of the prediction error may correspond to a normal distribution, while in others, it may correspond to a probability distribution other than a normal distribution, such as a Bernoulli distribution, a uniform distribution, a binomial distribution, a hypergeometric distribution, a geometric distribution, an exponential distribution, etc. In such other instances, the estimated distribution parameters may, in some examples, include parameters other than the mean, standard deviation, and variance, such as one or more parameters specific to the characteristics of the given probability distribution of the prediction error. For example, the distribution parameters estimated for a given probability distribution of prediction errors corresponding to a uniform distribution may include a minimum parameter and a maximum parameter (a and b), while the distribution parameters estimated for a given probability distribution of prediction errors corresponding to an exponential distribution may include at least one rate parameter (λ). In some implementations, it is possible to perform operations related to... Figure 11Process 1110 performs one or more operations similar to those operations, enabling the acquisition of data indicating the prediction error of the first machine learning model and the use of such data to estimate such distribution parameters. In at least some of the foregoing implementations, data indicating the prediction error of the first machine learning model can be acquired and used to (i) identify, from several different types of probability distributions (e.g., normal distribution, Bernoulli distribution, uniform distribution, binomial distribution, hypergeometric distribution, geometric distribution, exponential distribution, etc.), a specific type of probability distribution that most closely corresponds to the shape of the probability distribution of a given subset of the prediction error indicated by the data, and (ii) estimate one or more parameters of the probability distribution of the given subset of the prediction error indicated by the data based on the identified specific type of probability distribution. Other configurations are possible.

[0314] Referring again to the examples where the estimated distribution parameters include the mean and standard deviation, in which the MPC cluster is able to calculate such distribution parameters for a true label equal to zero:

[0315]

[0316]

[0317] in:

[0318]

[0319] count0=∑ i (1-L i )

[0320]

[0321] In some examples, MPC clusters are based on variance σ0 2 To calculate the standard deviation σ0, for example by calculating the variance σ0 2 The square root of . Similarly, to estimate such distribution parameters for true labels equal to one, the MPC cluster can compute:

[0322]

[0323]

[0324] in:

[0325]

[0326] count1=∑L i

[0327]

[0328] In some examples, MPC clusters are based on variance σ1 2 To calculate the standard deviation σ1, for example by calculating the variance σ1 2 The square root of.

[0329] Once these distribution parameters are estimated, the coefficients {a2, a1, a0} can be computed, stored, and later used to apply the corresponding transformation f to the sum of the true labels (sum_of_labels) for the k nearest neighbor user profiles. In some examples, these coefficients are used to configure a first machine learning model such that, forward, the first machine learning model applies the corresponding transformation f to the sum of the true labels for the k nearest neighbor user profiles in response to the input.

[0330] Similar to binary classification, in the case of multi-class classification, for each vector or set of true labels in the user profiles of the k nearest neighbor user profiles, each true label can be a binary value of zero or one. For this reason, a method similar to the one described above for binary classification can also be adopted in the implementation of multi-class classification techniques, making L-based... i and The calculated residual values ​​cannot be used to predict L. i However, in the case of multi-class classification, a corresponding function or transformation f can be defined and utilized for each class. For example, if each vector or set of true labels for each user profile will contain w different true labels corresponding to w different classes respectively, then w different transformations f can be determined and utilized. Furthermore, instead of calculating the sum of the true labels, in the case of multi-class classification, a frequency value is calculated for each class. Additional details on how such frequency values ​​can be calculated are provided above and immediately below. Other configurations are possible.

[0331] For any chosen j-th label, the MPC cluster can be based on l j Is it the training labels used to split the training examples into two groups? For l j These are training example groups for training labels; the MPC cluster can assume frequency. j This is based on a normal distribution and calculates the mean μ1 and variance σ1. On the other hand, for l j For training example groups that are not labeled with training data, the MPC cluster can assume frequency. j It is based on a normal distribution and calculates the mean μ0 and variance σ0.

[0332] Similar to binary classification, in the case of multi-class classification, the predictions of the k-NN model are likely to be biased (e.g., μ0 > 0 where it should have been 0, and μ1 < k where it should have been k). Additionally, there is no guarantee that σ0 == σ1. Thus, similar to binary classification, in the case of multi-class classification, the MPC cluster applies the transformation f on the predicted frequency j such that after the transformation, the Residue of the two groups j has substantially the same normal distribution. To achieve such a goal, the transformation f needs to satisfy the following properties:

[0333] (i) f(μ0) = 0

[0334] (ii) f(μ1) = k <00015​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​One or more of the operations described by at least one function or transformation method. In particular, steps 912-916 may be performed for an implementation in which one or more binary and / or multi-class classification techniques will be employed. As mentioned above, steps 912-916 are performed before steps 920 and 930 and may be performed after step 910.

[0344] The MPC cluster estimates a set of distribution parameters based on multiple real labels for multiple user profiles (912). For example, this could correspond to calculating the parameters μ0, σ0 as described above for the MPC cluster based on real labels associated with the same user profiles used in step 910. 2 , σ0, μ1, σ1 2 One or more operations that are similar to or equivalent to one or more operations performed in σ1.

[0345] The MPC cluster derives the function (914) based on the estimated set of distribution parameters. For example, this could correspond to one or more operations similar to or equivalent to those performed with respect to the parameters or coefficients (such as {a2, a1, a0}) that effectively define the function for the MPC cluster. Therefore, in some implementations, to derive the function at step 914, the MPC cluster derives the set of parameters of the function, for example, {a2, a1, a0}.

[0346] The MPC cluster configures the first machine learning model to generate initial predicted labels given a user profile as input and applies the derived function to the initial predicted labels to generate predicted labels for the user profile as output (916). For example, this could correspond to one or more operations similar to or equivalent to those performed on the MPC cluster configuring the first machine learning model such that the first machine learning model applies a corresponding transformation f in response to the input to the sum of the true labels for the k nearest neighbor user profiles (in the case of binary classification). In the case of multi-class classification, the transformation f could represent one of w different functions applied by the MPC cluster to a corresponding one of w different values ​​in a vector or set corresponding to w different classes. As mentioned above, each of these w different values ​​could correspond to a frequency value.

[0347] Having already performed steps 912-916 and configured the first machine learning model in this manner, the data generated in step 920 and subsequently utilized, for example, in step 930, may not be used to predict the true label (L). i ).

[0348] Refer again Figure 8In some implementations, process 800 may include the above references. Figures 9 to 11 One or more corresponding steps in the described operation.

[0349] In some implementations, process 800 further includes one or more operations in which the MPC cluster evaluates the performance of the first machine learning model. For example, this could correspond to the operations performed by the MPC cluster as described above. Figure 9 The described step 920 performs one or more operations similar to or equivalent to one or more operations. In these implementations, to evaluate the performance of the first machine learning model, for each of the multiple user profiles, the MPC cluster determines a predicted label for the user profile based at least in part on (i) the user profile, (ii) the first machine learning model, and (iii) one or more of the multiple real labels for the multiple user profiles, and determines a residual value of the user profile indicating the prediction error in the predicted label based at least in part on the predicted label determined for the user profile and the real label for the user profile included in the multiple real labels. For example, this could correspond to the MPC cluster performing as described above. Figure 11 The operations performed in steps 1108-1106 described herein are similar to or equivalent to one or more operations. Additionally, in these implementations, process 800 further includes one or more operations in which the MPC cluster uses data indicating residual values ​​determined for multiple user profiles to train a second machine learning model while evaluating the performance of the first machine learning model. For example, this could correspond to the operations performed by the MPC cluster as described above. Figure 9 The one or more operations performed in step 930 described are similar to or equivalent to one or more operations.

[0350] In at least some of the aforementioned implementations, the residual value of the user profile indicates the difference between the predicted label determined for the user profile and the true label for the user profile. For example, this could be for an example in which regression techniques are employed.

[0351] In at least some of the foregoing implementations, before the MPC cluster evaluates the performance of the first machine learning model, process 800 further includes one or more operations in which the MPC cluster derives a function based at least in part on multiple real labels and configures the first machine learning model to use the function to generate predicted labels for the user profile as output, given a user profile as input. For example, this could correspond to the above-referenced implementation of the MPC cluster. Figure 9The operations performed in steps 914-916 described are similar to or equivalent to one or more operations. Therefore, in some implementations, in order to derive the function at this step, the set of parameters of the MPC cluster derivation function is used, for example, {a2, a1, a0}.

[0352] In at least some of the foregoing implementations, process 800 further includes one or more operations in which the MPC cluster estimates a set of distribution parameters at least in part based on a plurality of true labels. In such implementations, in order to derive a function at least in part based on a plurality of true labels, the MPC cluster derives the function at least in part based on the estimated set of distribution parameters. For example, this could correspond to the MPC cluster performing as referenced above. Figure 9 The operations performed in steps 912-914 described herein are similar to or equivalent to one or more operations. Therefore, the aforementioned set of distribution parameters can include one or more parameters of the probability distribution of the prediction error of the true label for a first value among multiple true labels, such as the mean (μ0) and variance (σ0) of the normal distribution of the prediction error of the true label for a first value among multiple true labels, and one or more parameters of the probability distribution of the prediction error of the true label for a second value among multiple true labels, such as the mean (μ1) and variance (σ1) of the normal distribution of the prediction error of the true label for a second distinct value among multiple true labels. As described above, in some examples, the aforementioned set of distribution parameters can include other types of parameters. Furthermore, in at least some of the aforementioned implementations, the function is a quadratic polynomial function, for example, f(x) = a²x 2 +a1x+a0, where f′(x)=2a2x+a1.

[0353] In at least some of the aforementioned implementations, in order to configure the first machine learning model to use a function to generate predicted labels for a given user profile as output, the MPC cluster configures the first machine learning model, given a user profile as input, to: (i) generate initial predicted labels for the user profile, and (ii) apply a function to the initial predicted labels for the user profile to generate predicted labels for the user profile as output. For example, for an example employing a binary classification technique, this could correspond to one or more of the following operations, wherein the MPC cluster configures the first machine learning model, given a user profile as input, to: (i) compute the sum of the true labels (sum_of_labels) for the k nearest neighbor user profiles, and (ii) apply a function (transformation f) to the initial predicted labels for the user profile to generate predicted labels for the user profile. As output. A similar operation can be performed for implementations employing multi-class classification techniques. In some implementations, to apply a function to the initial predicted labels for a user profile, the MPC cluster applies a function defined as based on a derived set of parameters (e.g., (a2, a1, a0)). In some examples, to determine the predicted labels at least in part based on the true labels for each of the k nearest neighbor user profiles, the MPC cluster determines the sum of the true labels for the k nearest neighbor user profiles. This could be, for example, for implementations employing regression or binary classification techniques. In some of the foregoing examples, the predicted label for a particular user profile may correspond to the sum of the true labels for the k nearest neighbor user profiles. This could be, for example, for implementations employing regression classification techniques. In other such examples, in order to determine the predicted label based at least in part on the true labels for each of the k nearest neighbor user profiles, the MPC cluster applies a function to the sum of the true labels for the k nearest neighbor user profiles to generate a predicted label for a particular user profile. For example, this could be for an implementation that employs a binary classification technique.

[0354] As mentioned above, in some of the aforementioned implementations, in order to determine the predicted label at least partially based on the true labels for each of the k nearest neighbor user profiles, the MPC cluster determines the set of predicted labels at least partially based on the set of true labels for each of the k nearest neighbor user profiles, respectively corresponding to a set of categories, and, in order to determine the set of predicted labels, the MPC cluster performs an operation for each category in the set. Such operations can include one or more operations in which the MPC cluster determines the frequency of true labels in the set of true labels for user profiles in the k nearest neighbor user profiles whose true label corresponding to a category is a first value. For example, this can correspond to obtaining at least one predicted label 629 using a first machine learning model 620 in one or more implementations employing one or more multi-class classification techniques. One or more operations performed are similar to or equivalent to one or more operations, as referenced above. Figures 6 to 7 As described above. In at least some of the aforementioned implementations, in order to determine the set of predicted labels, for each category in the set, the MPC cluster applies a function corresponding to that category to the determined frequency to generate predicted labels corresponding to that category for a specific user profile. For example, the corresponding function may correspond to the one described in the above reference. Figure 9 Step 914 describes one of the w different functions derived by the MPC cluster for w different categories.

[0355] Figure 12This is a flowchart illustrating an example process 1200 for generating inference results for a user profile at the computing system of an MPC cluster to improve performance. Reference operations can be performed, for example, at inference time. Figure 12 One or more of the operations described. At least some of the operations in process 1200 can be performed, for example, by, such as Figure 1 The first computing system implementation of the MPC cluster 130 of the MPC1 MPC cluster, and also capable of corresponding to the above reference. Figure 8 One or more of the operations described. However, in process 1200, one or more operations can be performed on the secret share in order to provide user data privacy protection. Generally, in at least some implementations, a "share," as described below and elsewhere in this document, can correspond to a secret share. Other configurations are possible. For example, a reference can be performed at inference time. Figure 12 One or more of the operations described.

[0356] The first computing system of the MPC cluster receives an inference request (1202) associated with a given user profile. For example, this could correspond to one or more operations similar to or equivalent to one or more operations performed by MPC1 of MPC cluster 130 upon receiving an inference request from application 112, as referenced above. Figure 1 As described. In some implementations, this can correspond to the information mentioned in the above reference. Figure 8 The one or more operations performed in step 802 described herein are similar to or equivalent to one or more operations.

[0357] The first computing system of the MPC cluster determines predicted labels (1204-1208) for a given user profile. Labels can be or include demographic-based user group identifiers or demographic characteristics associated with the user profile. In some implementations, this may correspond to information as referenced above. Figure 8The one or more operations performed in step 804 described are similar to or equivalent to one or more operations. However, in steps 1204-1208, it is possible to perform the determination of a predicted label for a given user profile on a secret share in order to provide user data privacy protection. In order to determine the predicted label for a given user profile, the first computing system of the MPC cluster (i) determines the first share of the predicted label (1204) at least in part based on the first share of the given user profile, a first machine learning model trained using multiple user profiles, and one or more real labels from multiple real labels for multiple user profiles, (ii) receives from the second computing system of the MPC cluster data indicating the second share of the predicted label determined by the second computing system of the MPC cluster at least in part based on the second share of the given user profile and a first set of one or more machine learning models, and (iii) determines the predicted label at least in part based on the first share and the second share of the predicted label (1208). For example, the second computing system of the MPC cluster may correspond to Figure 1 MPC cluster 130 MPC2.

[0358] In this example, the multiple ground truth labels for multiple user profiles may correspond to ground truth labels included as a portion of encrypted label data 626, which are ground truth labels for the multiple user profiles used to train and / or evaluate the first machine learning model 620. In some examples, the multiple ground truth labels may correspond to a share of another set of ground truth labels. One or more ground truth labels from the multiple ground truth labels on which the predicted label for a given user profile is based may, for example, include at least one ground truth label for each of the k nearest neighbor user profiles identified by the k-NN model 622 of the first machine learning model 620. In some examples, each of the multiple ground truth labels is encrypted, such as Figures 6 to 7 The situation is similar to the example above. Several of the various ways to determine predicted labels using the true labels of the k nearest neighbor user profiles have been described in detail above. As becomes apparent above, the way or method of using such true labels to determine predicted labels can depend at least in part on the type of inference technique employed (e.g., regression techniques, binary classification techniques, multi-class classification techniques, etc.). See the references above. Figures 1 to 5 Additional details are provided regarding secret share swaps that can be performed in association with k-NN computation.

[0359] The first computing system of the MPC cluster determines the prediction residual values ​​(1210-1214) that indicate the prediction error in the prediction labels. In some implementations, this may correspond to the values ​​mentioned above. Figure 8The one or more operations performed in step 806 described herein are similar to or equivalent to one or more operations. However, in steps 1210-1214, it is possible to perform the determination of the prediction residual value on the secret share in order to provide user data privacy protection. In order to determine the prediction residual value, the first computing system of the MPC cluster (i) determines the first share of the prediction residual value of the given user profile (1210) based at least in part on the first share of the given user profile and the second machine learning model trained using multiple user profiles and data indicating the differences between multiple real labels for multiple user profiles and multiple predicted labels as determined for multiple user profiles using the first machine model, (ii) receives data from the second computing system of the MPC cluster indicating the second share of the prediction residual value of the given user profile determined at least in part by the second computing system of the MPC cluster based on the second share of the given user profile and the second set of one or more machine learning models (1212), and (iii) determines the prediction residual value of the given user profile (1214) based at least in part on the first share and the second share of the prediction residual value.

[0360] The first computing system of the MPC cluster generates data representing the inference results based on the predicted labels and predicted residuals (1216). In some implementations, this may correspond to the information mentioned in the above reference. Figure 8 The steps described in step 808 are one or more operations similar to or equivalent to one or more operations. Therefore, in some examples, the inference result includes or corresponds to the sum of the predicted labels and the predicted residuals.

[0361] The first computing system of the MPC cluster provides data representing the inference results to the client device (1218). In some implementations, this may correspond to the information mentioned in the above reference. Figure 8 The described step 810 performs one or more operations similar to or equivalent to one or more operations. For example, this could correspond to one or more operations similar to or equivalent to those performed by MPC cluster 130 to provide inference results to client device 110 running on application 112, as referenced above. Figures 1 to 2 As described.

[0362] In some implementations, process 1200 further includes one or more operations in which a first computing system of the MPC cluster applies a transformation to a first share of a given user profile to obtain a first transformed share of the given user profile. In these implementations, in order to determine the predicted label, the first computing system of the MPC cluster determines the first share of the predicted label at least in part based on the first transformed share of the given user profile. For example, this could correspond to the use of random projection logic 610 to transform user profile 609(P iApply random projection transformation to obtain the transformed user profile 619(P) i One or more operations performed are similar to or equivalent to one or more operations, as referred to above. Figures 6 to 8 As described.

[0363] In at least some of the aforementioned implementations, in order to determine the first share of the predicted label, the first computing system of the MPC cluster provides the first transformed share of a given user profile as input to the first machine learning model to obtain the first share of the predicted label for the given user profile as output. For example, this could correspond to receiving the transformed user profile 619 (P) with respect to the first machine learning model 620. i ′) as input and in response to it generate at least one predicted label 629 One or more operations performed are similar to or equivalent to one or more operations, as referenced above. Figures 6 to 7 As described.

[0364] In some examples, the aforementioned transformation can be a random projection. Furthermore, in at least some of these examples, the aforementioned random projection can be a Johnson-Lindenstrauss (JL) transformation.

[0365] In some implementations, to apply the JL transform, the MPC cluster can generate the projection matrix R in ciphertext. To transform the n-dimensional P... i Projecting to k dimensions, the MPC cluster can generate an n×k random matrix R. For example, the first computing system (e.g., MPC1) can create an n×k random matrix A, where A i,j With a 50% probability = 1 and A i,j With a 50% probability of 0, the first computational system can split A into two shares [A1] and [A2], discard A, keep [A1] confidentially, and give [A2] to the second computational system (e.g., MPC2). Similarly, the second computational system can create an n×k random matrix B whose elements have the same distribution as the elements of A. The second computational system can split B into two shares [B1] and [B2], discard B, keep [B2] confidentially, and give [B1] to the first computational system.

[0366] The first computational system is then able to compute [R1] as 2 × ([A1] == [B1]) - 1. Similarly, the second computational system is then able to compute [R2] as 2 × ([A2] == [B2]) - 1. In this way, [R1] and [R2] are two secret shares of R whose elements have equal probability of being 1 or -1.

[0367] The actual random projection onto P of dimension 1×n iThe secret share is compared with the projection matrix R of dimension n×k to produce a 1×k result. Assuming n >> k, the JL transformation reduces the dimension of the training data from n to k. To perform the above projection on the encrypted data, the first computational system is able to compute [P] i,1 ]⊙[R i,1 This requires multiplication between two shares and addition between two shares.

[0368] As mentioned above, in some implementations, the first machine learning model includes a k-nearest neighbor model maintained by a first computing system of the MPC cluster, and the first set of one or more machine learning models includes a k-nearest neighbor model maintained by a second computing system of the MPC cluster. In some examples, the two aforementioned k-nearest neighbor models may be identical or nearly identical to each other. That is, in some examples, the first and second computing systems maintain copies of the same k-NN model and each stores its own share of the true labels. In some examples, a model rooted in one or more prototype methods can be implemented instead of one or both of the aforementioned k-nearest neighbor models.

[0369] In at least some of these implementations, to determine the predicted label, a first computing system of the MPC cluster (i) identifies a first set of nearest-neighbor user profiles based at least in part on a first share of a given user profile and a k-nearest neighbor model maintained by the first computing system of the MPC cluster; (ii) receives data from a second computing system of the MPC cluster indicating the second set of nearest-neighbor profiles identified by the second computing system of the MPC cluster based at least in part on a second share of a given user profile and a k-nearest neighbor model maintained by the second computing system of the MPC cluster; (iii) identifies k nearest-neighbor user profiles considered most similar to a given user profile among a plurality of user profiles based at least in part on the first and second sets of nearest-neighbor profiles, and determines a first share of the predicted label based at least in part on the true label for each of the k nearest-neighbor user profiles. For example, this could correspond to obtaining at least one predicted label 629 using a first machine learning model 620 in one or more implementations employing one or more regression and / or binary classification techniques. One or more operations performed are similar to or equivalent to one or more operations, as referenced above. Figures 6 to 8 As described. In some examples, the predicted labels include or correspond to the sum of the true labels for the k nearest neighbor user profiles.

[0370] In some of the aforementioned implementations, to determine a first share of the predicted label, a first computing system of the MPC cluster (i) determines a first share of the sum of the true labels for the k nearest neighbor user profiles, (ii) receives a second share of the sum of the true labels for the k nearest neighbor user profiles from a second computing system of the MPC cluster, and (iii) determines the sum of the true labels for the k nearest neighbor user profiles based at least in part on the first and second shares of the sum of the true labels for the k nearest neighbor user profiles. For example, this could correspond to obtaining at least one predicted label 629 using a first machine learning model 620 in one or more implementations employing one or more multi-class classification techniques. One or more operations performed are similar to or equivalent to one or more operations, as referenced above. Figures 6 to 8 As described.

[0371] In some implementations, the second machine learning model includes at least one of a deep neural network (DNN), a gradient boosting decision tree (GBDT), and a random forest model maintained by the first computing system of the MPC cluster, and a second set of one or more machine learning models includes at least one of a DNN, GBDT, and random forest model maintained by the second computing system of the MPC cluster. In some examples, the two models (e.g., DNN, GBDT, random forest, etc.) maintained by the first and second computing systems may be identical or nearly identical to each other.

[0372] In some implementations, process 1200 further includes one or more operations in which the MPC cluster evaluates the performance of a first machine learning model and trains a second machine learning model using data indicating prediction residual values ​​determined for multiple user profiles in evaluating the performance of the first machine learning model. For example, this could correspond to the operations performed by the MPC cluster as described above. Figures 8 to 9The one or more operations performed in step 920 described are similar to or equivalent to one or more operations. However, in such implementations, one or more operations can be performed on the secret share to provide user data privacy protection. In these implementations, to evaluate the performance of the first machine learning model, for each of the plurality of user profiles, the MPC cluster determines a predicted label for the user profile and determines a residual value of the user profile indicating the prediction error in the predicted label. To determine the predicted label for the user profile, the first computing system of the MPC cluster (i) determines the first share of the predicted label for the user profile based at least in part on the first share of the user profile, the first machine learning model, and one or more real labels among the real labels for the plurality of user profiles, (ii) receives from the second computing system of the MPC cluster data indicating the second share of the predicted label for the user profile determined by the second computing system of the MPC cluster based at least in part on the second share of the user profile and a first set of one or more machine learning models maintained by the second computing system of the MPC cluster, and (iii) determines the predicted label for the user profile based at least in part on the first share and the second share of the predicted label. To determine the residual value of a user profile indicating errors in the predicted labels, a first computing system of the MPC cluster (i) determines a first share of the residual value of the user profile based at least in part on the predicted labels determined for the user profile and a first share of the real labels for the user profile included in a plurality of real labels; (ii) receives from a second computing system of the MPC cluster data indicating a second share of the residual value of the user profile determined by the second computing system of the MPC cluster based at least in part on the predicted labels determined for the user profile and the second share of the real labels for the user profile; and (iii) determines the residual value of the user profile based at least in part on the first and second shares of the residual value. For example, this could correspond to the above reference regarding the performance of the MPC cluster. Figure 11 The operations performed in steps 1108-1106 described herein are similar to or equivalent to one or more operations. Additionally, in these implementations, process 1200 further includes one or more operations in which the MPC cluster uses data indicating residual values ​​determined for multiple user profiles to train a second machine learning model while evaluating the performance of the first machine learning model. For example, this could correspond to the operations performed by the MPC cluster as described above. Figure 9 The one or more operations performed in step 930 described are similar to or equivalent to one or more operations.

[0373] In at least some of the aforementioned implementations, a first share of the residual value of the user profile indicates the difference between the predicted label determined by the first machine learning model for the user profile and the first share of the true label for the user profile, and a second share of the residual value of the user profile indicates the difference between the predicted label determined by the first machine learning model for the user profile and the second share of the true label for the user profile. This could be, for example, for an example in which regression techniques are employed.

[0374] In at least some of the aforementioned implementations, before the MPC cluster evaluates the performance of the first machine learning model, process 1200 further includes one or more operations, wherein the MPC cluster (i) derives a function and (ii) configures the first machine learning model to generate an initial predicted label for a user profile given a user profile as input and applies the function to the initial predicted label for the user profile to generate a first share of the predicted label for the user profile as output. For example, this could correspond to the above-referenced implementation of the MPC cluster. Figures 8 to 9 The one or more operations performed in steps 914-916 described herein are similar to or equivalent to one or more operations. To derive the function, the first computing system of the MPC cluster (i) derives a first share of the function based at least in part on a first share of each of the plurality of true labels, (ii) receives data from the second computing system of the MPC cluster indicating a second share of the function derived by the second computing system of the MPC cluster based at least in part on a second share of each of the plurality of true labels, and (iii) derives the function based at least in part on the first and second shares of the function. For example, for an example employing a binary classification technique, this could correspond to one or more of the following operations, wherein the MPC cluster configures a first machine learning model, given a user profile as input, to: (i) compute the sum of true labels (sum_of_labels) for the k nearest neighbor user profiles, and (ii) apply a function (transformation f) to the initial predicted labels for the user profile to generate predicted labels for the user profile. As output. A similar operation can be performed for cases where multi-class classification techniques are used.

[0375] When implemented on the secret share, the first computing system (e.g., MPC1) is able to compute:

[0376]

[0377] [count 0,1 ]=∑ i (1-[L i,1 ])

[0378]

[0379] Similarly, when implemented on secret shares, the second computational system (e.g., MPC2) is able to compute:

[0380]

[0381] [count 0,2 ]=∑ i (1-[L i,2 ])

[0382]

[0383] The MPC cluster can then reconstruct sum0, count0, and sum_of_square0 in plaintext as described above, and compute the distribution.

[0384] Similarly, in order to calculate the distribution The first computing system (e.g., MPC1) is capable of computing:

[0385]

[0386] [count 1,1 ]=∑ i [L i,1 ]

[0387]

[0388] Furthermore, the second computing system (e.g., MPC2) is capable of calculating:

[0389]

[0390] [count 1,2 ]=∑ i [L i,2 ]

[0391]

[0392] The MPC cluster can then reconstruct sum1, count1, and sum_of_square1 in plaintext as described above, and compute the distribution.

[0393] In at least some of the aforementioned implementations, when evaluating the performance of the first machine learning model, the MPC cluster can employ one or more fixed-point computation techniques to determine the residual value for each user profile. More specifically, when evaluating the performance of the first machine learning model, to determine a first share of the residual value for each user profile, the first computation system of the MPC cluster scales the corresponding real label or its share by a specific scaling factor, scales the coefficients {a2, a1, a0} associated with the function by a specific scaling factor, and rounds the scaled coefficients to the nearest integer. In such implementations, a second computation system of the MPC cluster can perform similar operations to determine a second share of the residual value for each user profile. The MPC cluster is thus able to compute the residual value with the secret share, reconstruct the plaintext residual value from the two secret shares, and divide the plaintext residual value by the scaling factor.

[0394] In at least some of the foregoing implementations, process 1200 further includes one or more operations in which the first computing system of the MPC cluster estimates a first share of the set of distribution parameters based at least in part on a first share of each of the plurality of real labels. In some such implementations, in order to derive the first share of a function based at least in part on the first share of each of the plurality of real labels, the first computing system of the MPC cluster derives the first share of the function based at least on the first share of the set of distribution parameters. For example, this could correspond to the above reference regarding the execution of MPC clusters. Figures 8 to 9 The operations performed in steps 912-914 described herein are similar to or equivalent to one or more operations. Therefore, the aforementioned set of distribution parameters can include one or more parameters of the probability distribution of the prediction error of the true label for a first value among multiple true labels, such as the mean (μ0) and variance (σ0) of the normal distribution of the prediction error of the true label for a first value among multiple true labels, and one or more parameters of the probability distribution of the prediction error of the true label for a second value among multiple true labels, such as the mean (μ1) and variance (σ1) of the normal distribution of the prediction error of the true label for a second distinct value among multiple true labels. As described above, in some examples, the aforementioned set of distribution parameters can include other types of parameters. Furthermore, in at least some of the aforementioned implementations, the function is a quadratic polynomial function, for example, f(x) = a²x 2 +a1x+a0, where f′(x)=2a2x+a1, however, in some examples, other functions can be used.

[0395] In some examples, to determine a first share of the predicted label, a first computing system of the MPC cluster (i) determines a first share of the sum of the true labels for the k nearest neighbor user profiles, (ii) receives a second share of the sum of the true labels for the k nearest neighbor user profiles from a second computing system of the MPC cluster, and (iii) determines the sum of the true labels for the k nearest neighbor user profiles based at least in part on the first and second shares of the sum of the true labels for the k nearest neighbor user profiles. This could be an implementation employing regression or binary classification techniques. In some of the foregoing examples, the first share of the predicted label may correspond to the sum of the true labels for the k nearest neighbor user profiles. This could be an implementation employing regression classification techniques. In other such examples, to determine the first share of predicted labels, the MPC cluster applies a function to the sum of the true labels for the k nearest neighbor user profiles to generate a predicted label for a given user profile. For example, this could be for an implementation that employs a binary classification technique.

[0396] As mentioned above, in some of the aforementioned implementations, in order to determine a first share of the predicted label based at least in part on the true labels for each of the k nearest neighbor user profiles, the first computing system of the MPC cluster determines a first share of the predicted label set based at least in part on the set of true labels for each of the k nearest neighbor user profiles corresponding to the set of categories. To determine the first share of the predicted label set, for each category in the set, the first computing system of the MPC cluster (i) determines a first share of the frequency of true labels whose true labels corresponding to the category are a first value in the set of true labels for user profiles in the k nearest neighbor user profiles, (ii) receives a second share of the frequency of true labels whose true labels corresponding to the category are a first value in the set of true labels for user profiles in the k nearest neighbor user profiles, and (iii) determines the frequency of true labels whose true labels corresponding to the category are a first value in the set of true labels for user profiles in the k nearest neighbor user profiles based at least in part on the first and second shares of the frequency of true labels whose true labels corresponding to the category are a first value in the set of true labels for user profiles in the k nearest neighbor user profiles. Such operations can include one or more operations in which a first computing system of an MPC cluster determines the frequency of true labels whose class-corresponding true labels are a first value in a set of true labels for user profiles in k nearest neighbor user profiles. For example, this could correspond to obtaining at least one predicted label 629 using a first machine learning model 620 in one or more implementations employing one or more multi-class classification techniques. One or more operations performed are similar to or equivalent to one or more operations, as referenced above. Figures 6 to 8 As described.

[0397] In at least some of the aforementioned implementations, in order to determine a first share of the set of predicted labels, for each category in the set, the first computing system of the MPC cluster applies a category-corresponding function to the frequency of true labels whose true labels corresponding to the category are first values ​​in the set of true labels for user profiles in the k nearest neighbor user profiles, to generate a first share of predicted labels corresponding to the category for a given user profile. For example, the corresponding function may correspond to the above reference. Figures 8 to 9 Step 914 describes one of the w different functions derived by the MPC cluster for w different categories.

[0398] For multi-class classification problems, when evaluating the performance (e.g., quality) of the first machine learning model, for each training example / query, the MPC cluster is able to find k nearest neighbors and compute the frequency of their labels on the secret share.

[0399] For example, consider the assumption that for a multi-class classification problem there exist w valid labels (e.g., classes) {l1, l2, ... ln}. w Example of}. In the case of {id1, id2, ... id} k Of the k identified neighbors, the first computing system (e.g., MPC1) can use [l] as the [l] j The frequency of the j-th tag in [1] is calculated as follows:

[0400]

[0401] The first calculation system can calculate the frequency based on the real label [label1] as follows:

[0402] [expected_frequency j,1 ]=k×([label1]==j)

[0403] Therefore, the first computing system is able to calculate:

[0404] [Residue j,1 ] = [expected_frequency] j,1 ]-[frequency j,1 ]

[0405] Furthermore, [Residue] j,1 Equivalent to:

[0406]

[0407] Similarly, the second computing system (e.g., MPC2) is capable of computing:

[0408]

[0409] In the case of binary classification and regression, for each inference, the residual value can be a secret message of integer type. Conversely, in the case of multi-class classification, for each inference, the residual value can be a secret message of integer vector, as shown above.

[0410] Population Statistics Report

[0411] Digital component provider 160 may have several activities (e.g., digital component distribution activities) that may involve different digital components. For each activity and digital component displayed to various users on client device 110, digital component provider 160 may expect feedback indicating the performance of that digital component or the activity including that digital component. To provide such feedback, content platform 150 is able to implement demographic reporting to generate and provide digital component provider 160 with reports indicating the effectiveness of each activity and / or each digital component for various demographic-based user groups. In one example, the report may include tables—such as those shown in Table 7 below—and analyses associated with the data shown in the tables. Each content provider 160 is displayed one or more reports specific to the activities and / or digital components for that particular content provider 160.

[0412]

[0413]

[0414] Table 7

[0415] While the digital component provider 160 shown in the report is described as including multiple activities involving one or more digital components, in other implementations, any digital component provider 160 can have any number of activities, any demographic categories (e.g., age range; gender; parental status; household income; lifestyle interests such as tech enthusiasts, sports fans, cooking enthusiasts, etc.; market segments such as product purchase interests; and / or any other categories), and any type of event or any combination of events (e.g., impressions, clicks and / or conversions, and / or their absence). The analysis generated for each report can vary accordingly. Demographic categories can correspond to various user groups that have been assigned to them through extended or self-reporting methods.

[0416] Because the data in the tables (e.g., Table 7) is calculated for all users as a whole rather than for individual users, the privacy of individual users is protected.

[0417] In some implementations, additional or alternative privacy protections can be implemented for demographic reporting, such as differential noise addition, record deidentification, k-anonymization, granularity-based techniques, etc., as described below. For differential noise addition, application 112, secure MPC cluster 130, content platform 150, and / or aggregation system 180 can add a controllable amount of differential noise from a preset distribution (e.g., a Laplace or Gaussian distribution) to one or more functions involving user-private data (e.g., identification information such as IP addresses and / or timestamps). For record deidentification, application 112 can simply send a set of records without any identification information such as IP addresses and / or timestamps. To deidentify records, application 112 and / or content platform 150 can remove identification information. For k-anonymity, application 112, secure MPC cluster 130, content platform 150, and / or aggregation system 180 can implement k-anonymization techniques, wherein at least "k" values ​​of user attributes within user data can be anonymized to enhance privacy, and operations such as aggregation can be performed on the anonymized data. K-anonymity requires that reports be aggregated on a given key and are only revealed if the key is shared with at least k records. For granularity-based techniques, application 112 can be programmed to allow digital component provider 160 to specify the granularity required for the report (e.g., time-based granularity such as one day or one hour, or geographic granularity such as a specific state, province, city, or country), and application 112, secure MPC cluster 130, content platform 150, and / or aggregation system 180 can perform computations on the specified granularity.

[0418] Systems for Demographic Reporting

[0419] The data presented in the report (e.g., the number of impressions, clicks, and / or conversions in each of the various demographic categories and / or the absence / absence of impressions, clicks, and / or conversions) could be generated using third-party cookies. However, to avoid cookies in order to protect user privacy, the report is executed using Environment 100's system (which can also be referred to as a framework).

[0420] Content platform 150 (e.g., DSP or SSP, and in some implementations, a separate reporting platform) can receive data from digital component provider 160 identifying activities involving digital components and one or more demographic categories for which digital component provider 160 expects to conduct demographic reporting. Digital component provider 160 can input this data into an application (e.g., a browser or native application). The first set of categories may include, for example, women and income greater than $100,000. In some implementations, the data identifying activities and the first set of one or more demographic categories can be included in an aggregation key, which can be a composite or cascaded key with multiple values ​​or multiple columns. These values ​​can be data identifying activities, and each column can represent a different demographic category.

[0421] Content platform 150 (e.g., DSP or SSP, and in some implementations a separate reporting platform coupled to content platform 150) enables a user of client device 110, on which digital components are being (or will be) displayed, to associate with a second set of one or more demographic categories. In one example, the second set of one or more demographic categories includes female, parent, and income greater than $100,000. The association can be performed in at least one of two ways. In the first way, content platform 150 (or, in some implementations, a separate reporting platform) is able to receive self-identification of the user to the second set of one or more demographic categories provided by application 112 implemented on client device 110, and then content platform 150 (or, in some implementations, a separate reporting platform) is able to map the user to the second set of one or more demographic categories to perform the association. In the second way, content platform 150 is able to transmit the user's browsing history (e.g., the browsing history may be or include a user profile, on which content platform 150 can...) to MPC cluster 130. Figure 2 The process 210 and 212 transmits the user profile to the MPC cluster 130; then the content platform 150 (or, in some implementations, a separate reporting platform) is able to receive inferences from the machine learning model within the MPC cluster 130, which outputs a second set of demographic categories; and subsequently, the content platform 150 (or, in some implementations, a separate reporting platform) is able to map the user to one or more of the second set of demographic categories to perform association.

[0422] This machine learning model can be a k-nearest neighbor model and can utilize the modeling techniques described above relative to demographic-based digital component distribution. However, this machine learning model used for reporting can be trained using different data from one or more machine learning models used for demographic-based digital component distribution because machine learning is used for different purposes in both distribution and reporting. For example, in the case of demographic-based digital component distribution, the purpose of machine learning is to suggest user groups to a user or the user's application (e.g., a browser) so that relevant digital components of interest to the user can be displayed; however, in the case of demographic reporting, the purpose of machine learning is to determine, in the report to digital component provider 160, the category (which can also be referred to as a bucket) into which a user who has been shown a digital component can be placed. Given these different purposes, the machine learning models used for demographic-based digital component distribution and reporting are trained differently and thus generate different outputs (i.e., the same probability output is classified into different categories—i.e., users are classified differently). For example, a machine learning model for demographic-based digital component distribution might always classify outputs that are 95% male as male, but a reporting model might report 100 users who are all 95% male as 95 males and 5 females; in this example, the MPC cluster 130 might be softer (i.e., easier or less strict) in committing specific labels for reporting purposes than for demographic-based digital component distribution purposes. The content platform 150 can control or modify this strictness of classification for the machine learning model implemented in the MPC cluster 130 for demographic-based digital component distribution and / or reporting. In some implementations, a separate reporting platform can control or modify this strictness of classification for the machine learning model implemented in the MPC cluster 130 for reporting.

[0423] Although the machine learning models are described as being trained differently for demographic-based digital component distribution and demographic reporting, in some implementations those models can be trained similarly or even in the same manner. Furthermore, while the machine learning models for demographic-based digital component distribution and demographic reporting are shown residing in MPC cluster 130, in some other implementations, the machine learning models for demographic-based digital component distribution and / or reporting can be implemented on client device 110, so that user classification into demographic groups occurs on client device 110 instead of MPC cluster 130. These implementations are typically implemented when client device 110 has sufficient storage capacity and computing power. Such alternative implementations can advantageously save bandwidth by preventing excessive communication with MPC cluster 130.

[0424] If a first set of one or more demographic categories (representing the demographic categories for which the digital component provider 160 expects to report demographic information; e.g., female and income greater than $100,000, as data input into application 112 by the digital component provider 160) and a second set of one or more demographic categories (representing inferences about user groups of users, such as those generated by machine learning or user groups self-identified by users on application 112; e.g., female, parents, and income greater than $100,000) have at least one common demographic category (e.g., female and income greater than $100,000), then the content platform 150 (or a separate reporting platform in some implementations) is able to transmit browsing events (e.g., impressions, clicks, and / or conversions, and / or their absence / non-existence) input on client device 110 and at least one common demographic category to the aggregation API. In the example given above for the first set of one or more demographic categories and the second set of one or more demographic categories, note that the category of female and income greater than $X is common. Therefore, in this example, the content platform 150 (or a separate reporting platform in some implementations) transmits browsing events (e.g., impressions, clicks and / or conversions, and / or their absence / non-existence) and data identifying common demographic categories entered on the client device 110 to the aggregation API.

[0425] In some implementations, the report may be in response to a request from the digital content provider 160. In some implementations, the report may be in response to a request from the content platform 150. In some implementations, the report may be in response to a specific type of user interaction (e.g., the display of specific digital content of a content item, one or more clicks on a digital content item, conversions associated with a digital content item, such as navigation to a product purchase page for purchasing products promoted using digital content items).

[0426] The aggregation API combines browsing events (e.g., impressions, clicks, and / or conversions, and / or their absence / non-existence) and at least one common demographic category with at least one browsing event from other users and at least one related demographic category as part of a first set of one or more demographic categories to generate aggregated data. In the example used above, the aggregation API combines data on browsing events for the category of Women and Income Greater Than X with browsing events from other users within the category of Women and Income Greater Than X—relative to this digital component. In this example, the aggregation does not take into account non-common categories for Parents (i.e., not common between the first set of one or more demographic categories and the second set of one or more demographic categories) because the digital component provider expects reports only for the specified category of Women and Income Greater Than X, not for the category of Parents. In some examples, the aggregated data can be the same as or similar to Table 7 discussed above. Aggregation can aggregate data for a preset time period (e.g., 1 hour, 12 hours, 1 day, 2 days, 5 days, 1 month, or any other time period). For example, in Table 7, the counts of impressions, clicks, and conversions are daily.

[0427] Aggregation is performed securely to prevent fraud and protect user privacy. The aggregation API communicates with aggregation system 180. Aggregation system 180 can be one or more computers communicatively coupled to content platform 150, a separate reporting platform, client device 110, website 142, publisher 140, and / or digital component provider 160. Aggregation system 180 can generate aggregated network measurements based on data received from client device 110. In some implementations, the data to be aggregated is sent to aggregation system by application 112, which can be a web browser or a native application. In several implementations, the data to be aggregated can be sent to aggregation system by the operating system of client device 110; in such implementations, web browsers and / or native applications on client device 110 can be configured to report impressions, clicks, and / or conversions to the operating system. The operating system can perform each of the operations described below for reporting impressions and conversions performed by application 112.

[0428] Application 112 on client device 110 can provide the aggregation system 180 with measurement data elements including encrypted data representing network data. The network data can include data regarding impressions, clicks, and / or conversions. For example, application 112 can generate and send measurement data elements for each conversion to the aggregation system 180, with conversion data stored at client device 110 for each conversion. For each of one or more digital components, the aggregated network measurement result can include the total number of impressions, clicks, and / or conversions across multiple client devices 110.

[0429] Application 112, secure MPC cluster 130, content platform 150 and / or aggregation system 180 can protect privacy by implementing various techniques such as thresholding schemes or two-party or other MPC computing systems as described below.

[0430] In some implementations, application 112 can use an (t,n) threshold scheme to generate data in the measurement data elements. In some implementations, when application 112 detects a conversion or receives conversion data for a conversion, application 112 generates a group key (e.g., a polynomial function) based on data about impressions, clicks, and / or conversions. Application 112 can then generate a group membership key that represents a portion of the group key and can be used to regenerate the group key only when a sufficient number of group membership keys for the same set of impressions, clicks, and conversions are received. In this example, the measurement data element for a conversion can include the group membership key generated by application 112 and a label corresponding to the set of impressions, clicks, and conversions. Each unique set of impressions, clicks, and conversions can have a corresponding unique label, allowing aggregation system 180 to use its labels to aggregate measurement data elements for each set of impressions, clicks, and / or conversions.

[0431] In the (t,n) threshold encryption scheme, the aggregation server needs to receive at least t group member keys for the same set of impressions, clicks, and / or conversions in order to decrypt the impression and conversion data. If fewer than t group member keys are received, the aggregation server cannot decrypt the data regarding impressions, clicks, and / or conversions. Once at least t measurement data elements for the same pair of impressions and conversions are received from the client device 110, the aggregation system 180 is able to determine the group key from the at least t group member keys and obtain the impression and conversion data from that group key.

[0432] Threshold encryption techniques, such as (t,n) threshold encryption schemes, can use web data (e.g., impressions, clicks, and / or conversions) or a portion thereof or a derivative thereof as a seed for generating a group key, which is then split among multiple applications (e.g., web browsers or native applications) on multiple client devices reporting the measurement of web data. This allows each application running on different client devices to generate the same group key, which uses the same web data to encrypt the web data, without collaboration between applications (or client devices) and without requiring a central system to distribute keys to each application. Alternatively, each application where a web event (e.g., impressions and associated conversions) occurs can use web data received, for example, from digital components and / or remote servers, to generate a group key to encrypt the web data.

[0433] Each application can use different information to generate a group membership key, which, when combined with a sufficient number of other group membership keys, can be used to regenerate a group key or another representation of the group key. For example, each application can use its unique identifier to generate its group membership key, allowing each application to generate a group membership key different from each other without inter-application collaboration. This generation of different group membership keys by each application makes it possible to regenerate the group key when any combination of group membership keys that total at least a threshold number "t" is received. Therefore, network data can be decrypted when at least t group membership keys are received, but cannot be decrypted when fewer than t group membership keys are received. By enabling this secret sharing between applications without inter-application collaboration, user privacy is protected by eliminating communication between user devices, reducing bandwidth consumption through such communication, and preventing measurement fraud that could occur if a single private key were simply sent to each application.

[0434] The aggregation system 180 can determine the number of clicks and / or conversions for an impression based on the number of received measurement data elements, including data about impressions, clicks, and / or conversions for a set of impressions, clicks, and / or conversions. For example, after obtaining impression, click, and conversion data using at least t group membership keys, the aggregation system 180 can determine the number of group membership keys received for the set of impressions, clicks, and / or conversions as the number of conversions. The aggregation system 180 can report the data about impressions, clicks, and / or conversions to the content platform 150 (or a separate reporting platform in some implementations) via the aggregation API.

[0435] Content platform 150 (or a separate reporting platform in some implementations) can receive aggregated data from the aggregation API. Content platform 150 (or a separate reporting platform in some implementations) can use the aggregated data to generate reports. To generate reports, content platform 150 (or a separate reporting platform in some implementations) can arrange the aggregated data in tables (e.g., a table like Table 7 or similar), generate analyses based on the aggregated data in the tables, and combine and present the tables and analyses in the report. Because the tables (e.g., Table 7) and the data in the report are calculated for all users collectively rather than for individual users, and because such tables or reports are not controlled by application 112, vendors, privacy experts, or any other such entity, such reports protect user privacy and prevent the leakage of user data, while avoiding the use of cookies. Content platform 150 (or a separate reporting platform in some implementations) can transmit reports on the aggregated data to the application (e.g., a browser or native application) of digital component provider 160. The reports can be presented on a user interface displayed by an application implemented on the computing device of digital component provider 160.

[0436] Content platform 150 (or a separate reporting platform in some implementations) is capable of generating reports in response to a report generation request. In some implementations, content platform 150 (or a separate reporting platform in some implementations) is capable of receiving report requests from digital content provider 160. In one example consistent with these implementations, digital component provider 160 may request reports for a specific digital component of digital component provider 160. In another example, digital component provider 160 may request reports for several digital components of digital component provider 160. In yet another example, digital component provider 160 may request reports for one or more activities of digital component provider 160. In some examples, digital component provider 160 may request reports by specifying one or more activity IDs, one or more digital component IDs, one or more demographic user group IDs, and / or one or more events for which a report is expected (e.g., impressions, clicks and / or conversions, and / or their absence / non-existence). In other implementations, requests for report generation can be generated automatically. The automatic generation of requests via scripts can occur at (a) preset time intervals and / or (b) when the count of one or more events associated with a digital component or activity of the content provider 160—such as impressions, clicks and / or conversions, and / or their absence / non-existence—exceeds a specific threshold (e.g., when the number of clicks by women exceeds 1000). In some implementations, the script can generate requests automatically in response to a request from the content provider. Typically, the reports are not time-sensitive, and taking a minute or more to generate a report may not be disadvantageous.

[0437] Although a thresholding scheme has been described above, in some implementations, the aggregation system 180 can be a secure multi-party (e.g., two-party) computation system. The aggregation system 180 can allow information across multiple sites to be folded into a single privacy-preserving report, which is made possible by a write-only per-source data store that flushes the data to the reporting endpoint after an aggregation threshold is reached across many client devices 110. That is, data is reported only when server-side aggregation services are used to fully aggregate data across browsers (or other application users).

[0438] Technology for population reporting

[0439] Figure 13 This is a diagram illustrating an example process 1300 for demographic reporting performed by content platform 150. While the reporting is described as being performed by content platform 150, in some implementations, the reporting can be performed by a separate reporting platform, as indicated above. Content platform 150 can receive at 1302 data from an application (e.g., a browser or native application) of digital component provider 160 that identifies activities involving the digital component and one or more demographic categories that digital component provider 160 expects to demographically report. Digital component provider 160 can input this data into such an application. The first set of categories may include, for example, women and income greater than X dollars. In some implementations, the data identifying the activities and the first set of one or more demographic categories can be included in an aggregation key, which can be a composite or cascaded key with multiple values ​​or multiple columns of values.

[0440] Content platform 150 is able at 1304 to associate a user of client device 110, on which (or to be) displaying digital components, with a second set of one or more demographic categories. In one example, the second set of one or more demographic categories includes female, parent, and income greater than X dollars. The association can be performed in at least one of two ways. In the first way, content platform 150 is able to receive a self-identification of the user to the second set of one or more demographic categories provided by application 112 implemented on client device 110, and then content platform 150 is able to map the user to the second set of one or more demographic categories to perform the association. In the second way, content platform 150 is able to transmit the user's browsing history to MPC cluster 130; then content platform 150 is able to receive an inference from a machine learning model within MPC cluster 130, which outputs an inference including the second set of demographic categories; and subsequently, content platform 150 is able to map the user to the second set of one or more demographic categories to perform the association.

[0441] If a first set of one or more demographic categories (representing the demographic categories for which the digital component provider 160 expects to report demographic data; e.g., female and income greater than $100,000, as data input into application 112 by the digital component provider 160) and a second set of one or more demographic categories (representing inferences about user groups of users, such as those generated by machine learning or user groups self-identified by users on application 112; e.g., female, parent, and income greater than $100,000) have at least one common demographic category (e.g., female and income greater than $100,000), then the content platform 150 is able to transmit browsing events (e.g., impressions, clicks, and / or conversions, and / or their absence / non-existence) input on client device 110 and at least one common demographic category to the aggregation API at 1306. In the example given above for the first set of one or more demographic categories and the second set of one or more demographic categories, note that the category of female and income greater than $X is common. Therefore, in this example, the content platform 150 transmits browsing events (e.g., impressions, clicks and / or conversions, and / or their absence / non-existence) and data identifying common demographic categories to the aggregation API input on the client device 110.

[0442] The aggregation API combines browsing events (e.g., impressions, clicks, and / or conversions, and / or their absence / non-existence) and at least one common demographic category with at least one browsing event from other users and at least one related demographic category as part of a first set of one or more demographic categories to generate aggregated data. In the example used above, the aggregation API combines data on browsing events for the category of Women and Income Greater Than X dollars with browsing events from other users within the category of Women and Income Greater Than X dollars—relative to this digital component. In this example, the aggregation does not take into account non-common categories for Parents (i.e., not common between the first set of one or more demographic categories and the second set of one or more demographic categories) because the digital component provider expects reports only for the specified category of Women and Income Greater Than X dollars, not for the Parents category. In some examples, the aggregated data can be the same as or similar to Table 7 discussed above. Aggregation can aggregate data for a preset time period (e.g., 1 hour, 12 hours, 1 day, 2 days, 5 days, 1 month, or any other time period). For example, in Table 7, the counts of impressions, clicks, and conversions are daily.

[0443] The aggregation API is described in detail above.

[0444] Content platform 150 can receive aggregated data from the aggregation API at point 1308. Content platform 150 can use the aggregated data to generate reports. To generate reports, content platform 150 can arrange the aggregated data in tables (e.g., a table like Table 7), generate analyses based on the aggregated data in the tables, and combine and present the tables and analyses in the report. Because the data in the tables (e.g., Table 7) and the reports is calculated for all users collectively rather than for individual users, and because such tables or reports are not controlled by application 112, vendors, privacy experts, or any other such entity, such reports protect user privacy and prevent the leakage of user data while avoiding the use of cookies.

[0445] Content platform 150 is capable of generating reports in response to a report generation request. In some implementations, content platform 150 is capable of receiving report requests from digital content provider 160. In one example consistent with these implementations, digital component provider 160 may request reports for a specific digital component of digital component provider 160. In another example, digital component provider 160 may request reports for several digital components of digital component provider 160. In yet another example, digital component provider 160 may request reports for one or more activities of digital component provider 160. In some examples, digital component provider 160 may request reports by specifying one or more activity IDs, one or more digital component IDs, one or more demographic user group IDs, and / or one or more events for which a report is expected (e.g., impressions, clicks and / or conversions, and / or their absence / non-existence). In other implementations, requests for report generation can be generated automatically. The automatic generation of requests via scripts can occur at (a) preset time intervals and / or (b) when the count of one or more events associated with a digital component or activity of content provider 160—such as impressions, clicks and / or conversions, and / or their absence / non-existence—exceeds a specific threshold (e.g., when the number of clicks by women exceeds 1000). In some implementations, the script can generate requests automatically in response to a request from the content provider.

[0446] Content platform 150 can transmit reports on aggregated data at point 1310 to the application (e.g., a browser or native application) of digital component provider 160. The reports can be presented on a user interface displayed by an application implemented on the computing device of digital component provider 160. Typically, the reports are not latency-sensitive. For example, taking a minute or more to generate a report may not be disadvantageous.

[0447] Figure 14This is a diagram illustrating an example process 1400 executed by client device 110 to facilitate the distribution and demographic reporting of demographic-based digital components. Application 112 of client device 110 is able to receive, at 1402, data identifying inferred demographic characteristics of the application's users from one or more computers (e.g., an MPC cluster 130 including two computing systems, MPC1 and MPC2). Application 112 is able to display digital content at 1404, which includes computer-readable code for reporting events related to the digital component, data specifying a set of permitted demographic-based user group identifiers, and an activity identifier for the digital component. The digital content for which reporting is being performed can be content from an electronic resource, such as a web page or native application. In another example, the digital content for which reporting is being performed can be a digital component displayed in a digital component slot of an electronic resource.

[0448] Application 112 can determine at 1406 that a given inferred feature matches a given allowed demographic user group identifier from a allowed list of demographic user group identifiers. The allowed list of demographic user group identifiers can include demographic user group identifiers that are permitted to report on digital content. For example, the list can include one or more user groups whose owners have enabled reporting for digital content. The list can be included in the script of the digital content. To determine if a match exists, application 112 can compare the allowed demographic user group identifiers with a list of user group identifiers that include user groups whose members are users.

[0449] In response to determining that a given inferred demographic characteristic matches a given permitted demographic-based user group identifier, application 112 can use computer-readable code at 1408 to generate and transmit a request to update one or more event counts for the digital component and the given permitted demographic-based user group identifier. For example, the request could be to increment (e.g., expand) the number of users in a demographic-based user group identified by the given permitted demographic-based user group identifier who have been presented with and interacted with the digital content, such as selecting the digital content or performing some other action related to the digital content.

[0450] Receiving data at 1402 that identifies the inferred demographic characteristics of the user of application 112 includes receiving at client device 110 an inferred demographic user group identifier for adding the user to a demographic user group. In some implementations, an inference request including the user's profile can be transmitted to one or more computers. In some implementations, two or more computers are required to perform secure multi-party computation. This inference request can include the user's profile. The inferred demographic user group identifier can be received from one or more computers in response to the inference request.

[0451] The transmission of an inferred request to one or more computers can include sending a corresponding secret share of a user profile to each MPC computing system (e.g., an MPC server) that forms the one or more computers. MPC cluster 130 uses one or more machine learning models to perform a secure MPC process to generate an inferred secret share of a demographically based user group identifier and transmits the inferred secret share of the demographically based user group identifier to application 112.

[0452] In response to displaying a digital component, application 112 may transmit an inference request to one or more computers for inferred demographic characteristics of a user of application 112. The inference request may include a user profile and contextual signals associated with at least one of: (i) a digital component slot in which the digital component is displayed, or (ii) the digital component itself, wherein an inferred demographic user group identifier is received from one or more computers in response to the inference request. In some examples, contextual signals may include context-level signals such as the Uniform Resource Locator (URL) of a resource, the location of client device 110, the spoken language setting of application 112, the number of digital component slots, first screen or below, etc. In some instances, contextual signals for a digital component may include authoring signals such as information about the digital component, the format of the digital component (e.g., image, video, audio, etc.), the size of the digital component, etc.

[0453] A request to generate and transmit an update of one or more event counts for a digital component and a given permitted demographic user group identifier using computer-readable code can include (i) generating an aggregate key that includes an activity identifier and a given permitted demographic user group identifier, and (ii) transmitting the aggregate key along with the request.

[0454] Using computer-readable code to generate and transmit requests for updating one or more event counts for digital components and given permitted demographic-based user group identifiers can include calling the application programming interface (API) of application 112 to send the request.

[0455] In some implementations, process 1400 can be used to update potentially incorrect event counts. For example, when a user is presented with digital content, the user may have been inferred to be a member of a first demographic-based user group. However, some time later, the user may be inferred to be in a second demographic-based user group, different from the first. In this example, the request could be to decrement the event count for the first demographic-based user group and increment the event count for the second demographic-based user group.

[0456] Figure 15 This is a block diagram of an example computer system 1500 capable of performing the operations described above. System 1500 includes a processor 1510, memory 1520, storage device 1530, and input / output device 1540. Each of components 1510, 1520, 1530, and 1540 can be interconnected, for example, using a system bus 1550. Processor 1510 is capable of processing instructions for execution within system 1500. In some implementations, processor 1510 is a single-threaded processor. In another implementation, processor 1510 is a multi-threaded processor. Processor 1510 is capable of processing instructions stored in memory 1520 or on storage device 1530.

[0457] Memory 1520 stores information within system 1500. In one implementation, memory 1520 is a computer-readable medium. In some implementations, memory 1520 is a volatile memory cell. In another implementation, memory 1520 is a non-volatile memory cell.

[0458] Storage device 1530 provides mass storage for system 1500. In some implementations, storage device 1530 is a computer-readable medium. In various implementations, storage device 1530 may include, for example, a hard disk drive, an optical disk drive, a storage device shared by multiple computing devices over a network (e.g., a cloud storage device), or some other mass storage device.

[0459] Input / output device 1540 provides input / output operations for system 1500. In some implementations, input / output device 1540 may include one or more network interface devices, such as an Ethernet card, a serial communication device (e.g., an RS-232 port), and / or a wireless interface device (e.g., an 802.11 card). In another implementation, the input / output device may include a driver device configured to receive input data and send output data to external device 1560, such as a keyboard, printer, and display device. However, other implementations, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc., are also possible.

[0460] Although already Figure 15 An example processing system is described herein, but the subject matter and functional operations described herein can be implemented using other types of digital electronic circuit systems or computer software, firmware, or hardware (including the structures disclosed herein and their equivalents) or a combination thereof.

[0461] The embodiments of the subject matter and operation described in this specification can be implemented using digital electronic circuit systems or computer software, firmware, or hardware (including the structures disclosed in this specification and their equivalents), or a combination thereof. The embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium (or media) for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on artificially generated propagating signals—e.g., machine-generated electrical, optical, or electromagnetic signals—that are generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus. The computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof, or be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof. Furthermore, although the computer storage medium is not a propagating signal, it can be a source or destination of computer program instructions encoded in artificially generated propagating signals. Computer storage media can also be one or more separate physical components or media (e.g., multiple CDs, disks or other storage devices) or included in one or more separate physical components or media (e.g., multiple CDs, disks or other storage devices).

[0462] The operations described in this specification can be implemented as operations performed by a data processing device on data stored on one or more computer-readable storage devices or received from other sources.

[0463] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines for processing data, including, for example, programmable processors, computers, systems-on-a-chip (SoCs), or combinations thereof. Apparatus can include special-purpose logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, apparatus can include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. Apparatus and execution environments can implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0464] Computer programs (also called programs, software, software applications, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, objects, or other units suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program can be stored as a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or code sections). A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a communication network.

[0465] The processes and logic flows described in this specification can be executed by one or more programmable processors executing one or more computer programs to perform actions by manipulating input data and generating output. The processes and logic flows can also be executed by a dedicated logic circuit system, and the device can also be implemented as a dedicated logic circuit system, such as an FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit).

[0466] As an example, processors suitable for executing computer programs include both general-purpose microprocessors and special-purpose microprocessors. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. Essential components of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to such mass storage devices, or both. However, a computer does not necessarily need to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory can be supplemented by dedicated logic circuitry systems or incorporated into dedicated logic circuitry systems.

[0467] To provide interaction with the user, embodiments of the subject matter described herein can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse or trackball, that the user can use to provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending web pages to a web browser on the user's client device in response to a request received from a web browser.

[0468] Embodiments of the subject matter described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or middleware components (e.g., an application server), or front-end components (e.g., a client computer having a graphical user interface or web browser that a user can interact with through an implementation of the subject matter described herein), or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., self-organizing peer-to-peer networks).

[0469] A computing system can include clients and servers. Clients and servers are typically geographically separated and usually interact via a communication network. The client-server relationship occurs through computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., HTML pages) to the client device (e.g., for the purpose of displaying data to a user interacting with the client device and receiving user input from the user interacting with the client device). It is possible to receive data generated at the client device (e.g., the result of user interaction) at the server.

[0470] While this specification contains numerous details of specific implementations, these should not be construed as limiting the scope of any invention or potentially claimed content, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in this specification within the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as acting in certain combinations and even initially claimed in this way, it is possible in some cases to remove one or more features from a claimed combination, and the claimed combination may involve sub-combinations or variations of sub-combinations.

[0471] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in sequential order, or requiring all illustrated operations to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous. Furthermore, the separation of various system components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0472] Therefore, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing can be advantageous.

Claims

1. A method comprising: receiving, by a content platform from an application of a digital content provider, data identifying digital content and a first set of one or more demographic categories to perform a demographic report for; associating, by the content platform, a user of a client device on which the digital content is being displayed with a second set of one or more demographic categories; if the first set of one or more demographic categories and the second set of one or more demographic categories have at least one demographic category in common, transmitting, by the content platform to an API, a browsing event entered on the client device and the at least one demographic category in common, wherein the API combines the browsing event and the at least one demographic category in common with at least one browsing event of other users and at least one demographic category related as one of the first set of one or more demographic categories to generate aggregated data; receiving, by the content platform from the API, the aggregated data; generating, by the content platform, a report comprising the aggregated data; and transmitting, by the content platform, the report to the application of the digital content provider.

2. The method of claim 1, wherein, associating the user with the second set of one or more demographic categories comprises: receiving self-identification by the user on the application of the second set of one or more demographic categories; and mapping the user to the second set of one or more demographic categories.

3. The method of claim 1, wherein, associating the user with the second set of one or more demographic categories comprises: transmitting a browsing history of the user to a multi-party computation cluster; and receiving, from a machine learning model within the multi-party computation cluster, an inference output by the machine learning model, the inference comprising the second set of demographic categories; and mapping the user to the second set of one or more demographic categories.

4. The method of claim 1, wherein, generating the report comprises: arranging the aggregated data in a table; generating an analysis based on the aggregated data; and providing the table and the analysis in the report.

5. The method of claim 1, wherein, the report is generated in response to a request for the report.

6. The method of claim 1, wherein, the report is automatically generated at a preset time interval.

7. The method of claim 5, wherein, the request for the report is generated by one or more of the digital content provider that developed and provides the application, the content platform, a secure multi-party computation cluster, or a publisher.

8. The method of claim 5, wherein, the request for the report is automatically generated at a preset time interval or after a count of events exceeds a preset threshold.

9. The method of claim 1, wherein, the data identifying the digital content and the first set of one or more demographic categories are included in an aggregation key, wherein the data identifying the digital content comprises a uniform resource locator (URL) of a resource displaying the digital content.

10. The method of claim 1, further comprising: determining a third set of demographic categories related to the user; and limiting the application of the client device to display digital content indicated by a corresponding digital content provider as being associated with a category of the third set of demographic categories.

11. The method of claim 10, wherein, determining the third set of one or more demographic categories includes: receiving self-identification by the user on the application of the third set of one or more demographic categories that are relevant to the user.

12. The method of claim 10, wherein, determining the third set of one or more demographic categories includes: transmitting, to a multi-party computation cluster, a browsing history of the user; and receiving, from the multi-party computation cluster, data identifying the third set of demographic categories.

13. A system comprising: at least one programmable processor; and a machine-readable medium storing instructions that, when executed by the at least one programmable processor, cause the at least one programmable processor to perform operations comprising: receiving, by a content platform from an application of a digital content provider, data identifying digital content and a first set of one or more demographic categories for which to perform demographic reporting; associating, by the content platform, a user of a client device on which the digital content is being displayed with a second set of one or more demographic categories; if the first set of one or more demographic categories and the second set of one or more demographic categories have at least one demographic category in common, transmitting, by the content platform to an API, a browsing event entered on the client device and the at least one demographic category in common, wherein the API combines the browsing event and the at least one demographic category in common with at least one browsing event of another user and at least one demographic category that is relevant as one of the first set of one or more demographic categories to generate aggregated data; receiving, by the content platform from the API, the aggregated data; generating, by the content platform, a report comprising the aggregated data; and transmitting, by the content platform, the report to the application of the digital content provider.

14. The system of claim 13, wherein, associating the user with the second set of one or more demographic categories includes: receiving self-identification by the user on the application of the second set of one or more demographic categories; and mapping the user to the second set of one or more demographic categories.

15. The system of claim 13, wherein, associating the user with the second set of one or more demographic categories includes: transmitting, to a multi-party computation cluster, a browsing history of the user; and receiving, from a machine learning model within the multi-party computation cluster, an inference output by the machine learning model, the inference comprising the second set of demographic categories; and mapping the user to the second set of one or more demographic categories.

16. The system of claim 13, wherein, generating the report includes: arranging the aggregated data in a table; generating analysis based on the aggregated data; and providing the table and the analysis in the report.

17. The system of claim 13, wherein, the report is responsive to a request for the report.

18. The system of claim 13, wherein, the report is generated automatically at a preset time interval. the report is generated automatically at a preset time interval.

19. The system of claim 17, wherein, The request for the report is generated by one or more of the digital content provider that developed and provided the application, the content platform, a secure multi-party computation cluster, or a publisher.

20. The system of claim 17, wherein, The request for the report is automatically generated at a preset time interval or after a count of events exceeds a preset threshold.

21. The system of claim 13, wherein, The data identifying the digital content and the first set of one or more demographic categories are included in an aggregate key, or wherein the data identifying the digital content includes a uniform resource locator (URL) of a resource that displays the digital content.

22. A non-transitory computer program product storing instructions that, when executed by at least one programmable processor, cause the at least one programmable processor to perform operations comprising: receiving, by a content platform from an application of a digital content provider, data identifying digital content and a first set of one or more demographic categories for which to perform a demographic report; associating, by the content platform, a user of a client device on which the digital content is being displayed with a second set of one or more demographic categories; if the first set of one or more demographic categories and the second set of one or more demographic categories have at least one demographic category in common, transmitting, by the content platform to an API, a browsing event entered on the client device and the at least one demographic category in common, wherein the API combines the browsing event and the at least one demographic category in common with at least one browsing event of other users and at least one demographic category related as one of the first set of one or more demographic categories to generate aggregate data; receiving, by the content platform from the API, the aggregate data; generating, by the content platform, a report including the aggregate data; and transmitting, by the content platform, the report to the application of the digital content provider.

Citation Information

Patent Citations

  • Probabilistic inference of demographic information from user selection of content

    US9003441B1