Method and system for privacy-preserving activity aggregation mechanism
By allocating randomized groups and digital signature certificates on the browser, the problem of personalized digital component selection when the browser does not support third-party cookies is solved, user anonymity monitoring and fraud detection are realized, and privacy protection and system security are improved.
Patent Information
- Application Number
- CN202180019433.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-03
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-03-03
AI Technical Summary
In the case where the browser does not support third-party cookies, the selection and provision of personalized digital components becomes difficult, resulting in waste of computing resources and bandwidth, while user privacy protection is difficult to achieve.
The randomized grouping technology is adopted to ensure the anonymity and statistical utility of the randomized group by assigning a randomized group constructed based on randomly selected identifiers and timestamps on the user device, and providing corresponding digital signature certificates and unique public and private keys.
It realizes monitoring and analysis of network activities without sacrificing user anonymity, preventing user information leakage, and detecting fraud and abuse, improving user privacy protection and system security.
Smart Images

Figure CN115380506B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the aggregation of web activity, data processing, and protecting user privacy in an online environment. The increase in user privacy online has led many browser developers to change the way user data is processed. For example, some browsers no longer support certain types of cookies, but the deprecation of third-party (3P) cookies can lead to fraud and abuse. Background Art
[0002] Aggregating web activity allows a user's browsing experience to be personalized and can deliver more relevant content to the user more quickly than would be possible without monitoring. However, existing mechanisms, such as cookies, can be linked to an individual user and information about that user. Such precision can make users feel that they are too easily identified and their information too easily compromised. Summary of the invention
[0003] Generally, an innovative aspect of the subject matter described in this specification can be an embodiment of a method for privacy-preserving network activity monitoring, the method comprising: receiving a request for digital content from a domain from an application on a user device of a user; assigning to the application at a first time a randomized group constructed based on a randomly selected identifier and a timestamp indicating the first time the randomized group was assigned to the application; and providing to the application at the first time (i) a digital signature certificate corresponding to the randomly selected identifier and the timestamp, and (ii) a unique public key associated with the certificate and a corresponding unique private key, wherein, within a predetermined time period in which the randomized group is assigned to the application, the randomly selected identifier is also assigned to at least a threshold number of other applications executed on other user devices.
[0004] In some embodiments, the method includes: receiving a second request for digital content from the domain from the application; and providing, by the application at a second time to the domain, an obfuscated identifier corresponding to the randomly selected identifier and an age bucket of a randomized group indicating an age range of a cookie containing an age of the randomized group, wherein the age of the randomized group is calculated based on a difference between the second time and the first time.
[0005] In some embodiments, the method further includes: detecting, by the domain, anomalous activity associated with the randomly selected identifier based on the received randomized cohort age bucket and at least one of: a number of interactions associated with the randomly selected identifier, a randomized cohort age distribution, and a probability distribution associated with specific interactions and a specific time period.
[0006] In some embodiments, wherein assigning the randomized group to the application comprises: assigning the randomized group to the application by the domain, and wherein the randomly selected identifier is assigned to at least a threshold number of other applications, wherein the randomly selected identifier is a randomly generated identifier selected from two or more randomly generated identifiers, and wherein the unique public key is generated by the domain.
[0007] In some embodiments, assigning the randomized group to the browser includes: assigning the randomized group to the application by a central server; wherein the randomly selected identifier is assigned to at least a threshold number of other applications, wherein the randomly selected identifier is a randomly generated identifier selected from two or more randomly generated identifiers, and wherein the unique public key is generated by the central server.
[0008] In some embodiments, the method includes: providing, by an application, a request for digital content and a randomized group from a second domain different from the first domain; receiving, by the application, a certification request including a challenge from the second domain in response to providing the request for digital content from the second domain; and providing, by the application, a digitally signed certificate to a verification system, the digitally signed certificate triggering the verification system to (i) create a fuzzy certificate including a randomly selected identifier, a randomized group age bucket, and a challenge, (ii) sign the fuzzy certificate, and (iii) provide the fuzzy certificate to the second domain, wherein the challenge is shielded from the verification system using a blinding scheme.
[0009] In some embodiments, the method includes: providing, by the application, a digitally signed certificate to a verification system; and verifying, by the verification system, that the randomized group is assigned to at least a threshold number of persons.
[0010] Other embodiments of this aspect include corresponding systems, apparatus, and computer programs encoded on computer storage devices, the programs configured to perform the actions of the methods.
[0011] The subject matter described in this specification can be implemented in specific embodiments to realize one or more of the following advantages.
[0012] The manner in which a digital component distribution system selects and distributes personalized digital components (e.g., generating selection parameters and / or the selection parameters themselves) has historically included the use of user information (e.g., browsing information, interest group information, etc.) obtained from third-party cookies, which are cookies dropped on a client device by a domain that is different from the domain of the web page presented on the client device (e.g., eTLD+1). However, some browsers are blocking the use of third-party cookies, making it more difficult to select and provide personalized digital components, which means that computing resources and bandwidth may be wasted by selecting and distributing content that is not of interest to the user to the user. In addition, functions that a computer system could previously perform using third-party cookies can no longer be performed, resulting in a computer system that is less efficient and less effective. To overcome this problem, privacy-preserving technologies that can monitor, aggregate, and analyze network activity while hindering the tracking of users and simultaneously preventing the leakage of user information across computing systems can be used. In other words, the techniques discussed herein are changing the way computing systems operate to overcome the problems that arise when browsers do not support the use of third-party cookies.
[0013] The privacy-preserving monitoring mechanism described herein provides network activity monitoring functionality with randomized groups. The randomized group includes an identifier and a timestamp. The combination of the identifier and timestamp does not uniquely identify a particular browser or user device, but rather. However, the timestamp can be obfuscated while still providing useful information by generating an age bucket to which the randomized group belongs and providing a combination of the identifier and the age bucket. The identifier and age bucket combination is assigned to at least a threshold number of unique browsers running on different user devices, thereby ensuring anonymity without sacrificing the statistical utility of the randomized group. The randomized group can be used to generate statistics about group activity and other information while ensuring user anonymity. Users are randomly grouped into groups of size k so that the domain to which the randomized group applies can track the activity of the group rather than the activity of any one user.
[0014] In addition, randomized groups can be used in security contexts with third parties, such as content providers and hosts, to detect fraudulent activity or coordinated abuse. For example, randomized groups allow existing anti-abuse techniques to combat participatory abuse. Participatory abuse can include behaviors such as click fraud, view count inflation, rating manipulation, ranking manipulation, etc. Randomized groups can be used to detect suspicious network activity that indicates fraudulent use while providing users with a specific level of privacy that the user did not previously have. For example, randomized groups can be used to provide users with k-anonymity guarantees. The k-anonymity guarantee ensures that at least k random users are associated with a single randomized group, which can be identified by a randomized group identifier and a timestamp. For example, the k-anonymity guarantee of a randomized group identifier with k=100 ensures that at least 100 random users are associated with the randomized group identifier, so that information associated with a particular randomized group is anonymized to a certain extent while still facilitating applications such as statistical analysis and abuse detection.
[0015] The described monitoring mechanism, randomization groups, improves user experience and trust by providing privacy guarantees that can be externally verified by an independent third party. The described system can include one or more verification servers that are independent of the source of the randomization group, so that the specific identities of the randomization group and the user remain hidden, while allowing the verification server to determine the statistical properties of a specific randomization group identifier. This allows users to confirm through an independent source that their anonymity is maintained and that the privacy-preserving system is working as promised. Users who can independently verify the privacy guarantees may feel more comfortable adopting a system that uses the described monitoring mechanism.
[0016] Additionally, randomized groups can be used as a replacement in systems that use traditional web activity monitoring methods. For example, randomized groups can be used in existing systems with little to no adjustments under certain conditions, allowing system designers to reuse existing infrastructure to provide relevant statistics about web activity and perform security functions to protect third parties while improving user privacy.
[0017] The techniques discussed throughout this document can also be used to detect irregular activity (e.g., network attacks) and shut down irregular activity. For example, these techniques can detect higher than usual levels of network requests or traffic and use this information to throttle further network requests or traffic, or block further requests from the group of computing devices responsible for the high level of network requests or traffic. The techniques can also be used to detect overuse of specific computing resources and perform load balancing to improve the efficiency of computer systems.
[0018] Various features and advantages of the foregoing subject matter are described below with reference to the accompanying drawings. Additional features and advantages will be apparent from the subject matter described herein and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a block diagram of a system in which a privacy-preserving monitoring mechanism is implemented.
[0020] Figure 2 is a data flow diagram of an example process for publishing and implementing a privacy-preserving monitoring mechanism.
[0021] Figure 3A is a swim lane diagram illustrating an example process for publishing and implementing a privacy-preserving monitoring mechanism that is limited in scope to a single publishing domain.
[0022] Figure 3B is a swim lane diagram illustrating an example process for publishing and implementing a privacy-preserving monitoring mechanism for use in network activities performed across different domains.
[0023] Figure 4 is a flow chart illustrating an example process for publishing and implementing a privacy-preserving monitoring mechanism.
[0024] Figure 5 is a block diagram of an example computer system.
[0025] The same reference numbers and names in different drawings represent the same elements. DETAILED DESCRIPTION
[0026] In general, this document describes systems and techniques for ensuring specified privacy levels for users associated with monitoring mechanisms, randomized groups, which also provide statistical tracking and abuse detection capabilities.
[0027] Figure 1 1 is a block diagram of an environment 100 for privacy-preserving data collection and analysis. The example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. The network 102 connects an electronic document server 104 ("electronic document server"), a user device 106, a digital component distribution system 110 (also referred to as DCDS 110), and one or more authentication servers 130. The example environment 100 may include many different electronic document servers 104, user devices 106, authentication servers 130, and trusted domain servers 140. For ease of explanation, one trusted domain server 140 is shown.
[0028] The user device 106 is an electronic device capable of requesting and receiving resources (e.g., electronic documents) through the network 102. Example user devices 106 include personal computers, wearable devices, smart speakers, tablet devices, mobile communication devices (e.g., smart phones), smart appliances, and other devices capable of sending and receiving data through the network 102. In some embodiments, the user device can include a speaker that outputs audible information to the user and a microphone (e.g., spoken input) that accepts audible input from the user. The user device can also include a digital assistant that provides an interactive voice interface for submitting input and / or receiving output provided in response to input. The user device can also include a display for presenting visual information (e.g., text, images, and / or video). The user device 106 typically includes a user application, such as a web browser, to facilitate sending and receiving data through the network 102, but the local application executed by the user device 106 can also facilitate sending and receiving data 102 through the network.
[0029] The user device 106 includes software 107. The software 107 can be, for example, a browser or an operating system. In some embodiments, the software 107 allows a user to access information through a network such as the network 102, retrieve information from a server, and display the information on a display of the user device 106. In some embodiments, the software 107 manages the hardware and software resources of the user device 106 and provides common services to other programs on the user device 106. The software 107 can act as an intermediary between programs and the hardware of the user device 106.
[0030] The software 107 is specific to each user device 106. As described in detail below, the privacy-preserving data analysis and collection innovations provide device-specific solutions that are resource efficient and secure.
[0031] An electronic document is data that presents a collection of content at a user device 106. Examples of electronic documents include web pages, word processing documents, portable document format (PDF) documents, images, videos, search result pages, and feeds. Local applications (e.g., "apps"), such as applications installed on mobile, tablet, or desktop computing devices, are also examples of electronic documents. Electronic document servers 104 can provide electronic documents 105 ("electronic documents") to user devices 106. For example, electronic document servers 104 can include servers that host publisher websites, such as a network domain (e.g., eTLD+1). Each of electronic document servers 104 can be a server in a separate domain (e.g., a different eTLD+1) or associated with a separate domain.
[0032] In this example, user device 106 can initiate a request for a given publisher web page, and electronic document server 104 hosting the given publisher web page can respond to the request by sending machine Hypertext Markup Language (HTML) code that initiates rendering of the given web page at user device 106 .
[0033] An electronic document can include a variety of content. For example, an electronic document 105 can include static content (e.g., text or other specified content) that is within the electronic document itself and / or does not change over time. An electronic document can also include dynamic content that can change over time or on a per-request basis. For example, a publisher of a given electronic document can maintain a data source for populating portions of an electronic document. In this example, a given electronic document can include a tag or script that causes the user device 106 to request content from a data source when the given electronic document is processed (e.g., rendered or executed) by the user device 106. The user device 106 integrates the content obtained from the data source into the rendering of the given electronic document to create a composite electronic document that includes the content obtained from the data source.
[0034] In some cases, a given electronic document can include a digital content tag or digital content script that references the DCDS 110. In these cases, when the given electronic document is processed by the user device 106, the digital content tag or digital content script is executed by the user device 106. The execution of the digital content tag or digital content script configures the user device 106 to generate a request 108 for digital content, which is transmitted to the DCDS 110 via the network 102. For example, the digital content tag or digital content script can enable the user device 106 to generate a packetized data request including header and payload data. The request 108 can include data such as the name (or network location) of the server from which the digital content is requested, the name (or network location) of the requesting device (e.g., the user device 106), and / or information that the DCDS 110 can use to select the digital content to be provided in response to the request. The request 108 is transmitted by the user device 106 to the server of the DCDS 110 via the network 102 (e.g., a telecommunications network).
[0035] Request 108 can include data specifying characteristics of the electronic document and the location where the digital content can be presented. For example, data specifying a reference (e.g., a URL) to an electronic document (e.g., a web page) where the digital content is to be presented, available locations of the electronic document that can be used to present the digital content (e.g., digital content slots), the size of the available locations, the location of the available locations in the presentation of the electronic document, and / or the types of media eligible for presentation in these locations can be provided to DCDS 110. Similarly, data specifying selected keywords designated for the electronic document ("document keywords") or entities (e.g., people, places, or things) referenced by the electronic document can also be included in request 108 (e.g., as payload data) and provided to DCDS 110 to facilitate identification of digital content items, such as electronic documents or digital components, that are eligible for presentation with the electronic document.
[0036] The request 108 can also include data related to other information, such as information that the user has provided, geographic information indicating the state or region from which the request is being submitted, or other information that provides context for the environment in which the digital content will be displayed (e.g., the type of device on which the digital content will be displayed, such as a mobile device or tablet device). The user-provided information can include demographic data of the user of the user device 106. For example, demographic information can include age, gender, geographic location, education level, marital status, household income, occupation, hobbies, social media data, and whether the user owns a particular item, among other characteristics.
[0037] Data specifying characteristics of the user device 106 can also be provided in the request 108, such as information identifying the model of the user device 106, the configuration of the user device 106, or the size (e.g., physical size or resolution) of an electronic display (e.g., a touch screen or desktop monitor) on which the electronic document is presented. The request 108 can be transmitted, for example, over a packetized network, and the request 108 itself can be formatted as packetized data having a header and payload data. The header can specify the destination of the packet, and the payload data can include any of the information discussed above.
[0038] In addition to the privacy protection techniques discussed throughout this document, controls may be provided to users to allow them to choose whether and when the systems, programs, or features described herein enable the collection of user information (e.g., information about the user's social network, social behavior or activities, occupation, user preferences, or the user's current location) and whether to send content or communications from a server to the user. In addition, before certain data is stored or used, it may be processed in one or more ways so that personally identifiable information is removed. For example, the user's identity may be processed so that the user's personally identifiable information cannot be determined, or the user's geographic location may be generalized (e.g., to the city, zip code, or state level) when location information is obtained so that the user's specific location cannot be determined. Thus, users can control what information is collected about the user, how that information is used, and what information is provided to the user.
[0039] The DCDS 110 selects digital content to be presented with a given electronic document in response to receiving the request 108 and / or using information included in the request 108. In some embodiments, the DCDS 110 is implemented in a distributed computing system that includes, for example, a server and a collection of multiple computing devices that are interconnected and identify and distribute digital content in response to the request 108. The collection of multiple computing devices operates together to identify a collection of digital content that is eligible to be presented in an electronic document from a corpus of millions or more available digital content. The millions or more available digital content can be indexed, for example, in a digital component database 112. Each digital content index entry can reference the corresponding digital content and / or include distribution parameters (e.g., selection criteria) that regulate the distribution of the corresponding digital content.
[0040] The identification of eligible digital content can be split into multiple tasks, which are then distributed among the computing devices within the set of multiple computing devices. For example, different computing devices can each analyze different portions of the digital component database 112 to identify various digital content having distribution parameters that match the information included in the request 108.
[0041] DCDS 110 aggregates the results received from the set of multiple computing devices and uses information associated with the aggregated results to select one or more instances of digital content to be provided in response to request 108. DCDS 110 can then generate and transmit reply data 114 (e.g., digital data representing the reply) over network 102 that enables user device 106 to integrate the selected set of digital content into a given electronic document so that the selected set of digital content and the content of the electronic document are presented together at a display of user device 106.
[0042] The DCDS 110 can forward requests 108 from the software 107 of the user device 106 to a data source, such as the electronic document server 104, and can forward replies 114 from the electronic document server 104 to the software 107 of the user device 106. For example, the DCDS 110 acts as an intermediary between the electronic document server 104 and the user device 106 and / or the software 107 running on the user device 106.
[0043] The randomization group generator 121 (RCX generator) allows the electronic document server 104 to generate randomization groups, i.e., the privacy-preserving monitoring / aggregation mechanism described herein. In this document, a randomization group refers to a specific format of a privacy-preserving monitoring mechanism having a randomization group identifier and a randomization group timestamp. A randomization group is generated in response to an initial third-party request from an application such as software 107 or a device such as user device 106, and the randomization group includes an identifier (i.e., a randomization group identifier) and a timestamp (i.e., a randomization group timestamp). For example, a randomization group can be data represented by rcx (rcx.id, rcx.timestamp), where rcx represents a randomization group, rcx.id represents a randomization group identifier, and rcx.timestamp represents a timestamp. The randomization group is assigned to the software 107 or user device 106 from which the initial request was received. For simplicity of explanation, in this example, a randomization group is generated in response to an initial request from a browser 107. In other examples, the randomized group can be generated in response to an initial request from a particular user device 106 .
[0044] Because the randomization group generators 121 are associated with specific domains, the randomization groups generated by each generator 121 can be domain-wide, meaning that the randomization group data is used within the domain to which the randomization group generator 121 and / or the electronic document server 104 is associated, and the randomization groups are not provided to or shared with other domains or servers.
[0045] The initial request can be a request for an electronic document 105 from an electronic document server 104. The initial request can be a request for a content item from a third party server 150 that provides content, such as a digital component that can be provided for display with the content requested from the electronic document server 104.
[0046] The randomized group identifier is a randomly selected or constructed identifier that is also assigned to multiple other browsers 107 executed on other user devices 106 that have also provided initial requests to the electronic document server 104. For example, the randomized group identifier can be a randomly generated 64-bit identifier selected from a set of existing identifiers or created in response to an initial request from a browser 107. The number of other applications 107 to which the randomized group identifier is assigned is based on a predetermined threshold privacy level guaranteed to the user. For example, the electronic document server 104 can achieve a guarantee of k-anonymity, which means that each randomized group identifier is assigned to at least k browsers 107 running on different user devices 106. Each electronic document server 104 can independently select a k to guarantee. In some examples, each electronic document server 104 guarantees the same level of k-anonymity.
[0047] The randomization group timestamp indicates the time when the randomization group identifier was requested and / or assigned to the software 107 in response to the request 108. For example, the randomization group timestamp can indicate the time when the request 108 was received by the randomization group generator 121. In another example, the randomization group timestamp can indicate the time when the randomization group identifier was selected in response to the request 108. In another example, the randomization group timestamp can indicate the time when the randomization group identifier was assigned to the software 107 in response to the request 108. One or more of these actions can be performed simultaneously, and thus the randomization group timestamp can represent the time when one or more of these actions are performed. Because the randomization group identifier is assigned to at least k browsers 107 (i.e., for k=3000, assigned to 3000 different browsers 107) for the purpose of maintaining k-anonymity, the combination of the randomization group timestamp and the randomization group identifier can serve as a unique identifier. In order to protect privacy when providing a randomization group, the randomization group generator 121 can also anonymize the randomization group timestamp, creating a parameter representing the age bucket to which the randomization group belongs. The age bucket represents a broad range of age values within which the age of the randomized group falls, but cannot be used to uniquely identify the browser 107 to which the randomized group is assigned. For example, the randomized group generator 121 can determine the difference between the current time and the randomized group timestamp to determine the age of the randomized group. The randomized group generator 121 can then generate values for the age bucket based on information such as, for example, the k value and the age range or predetermined age range required to maintain k-anonymity, as well as other parameters.
[0048] In addition to the randomization group including the randomization group identifier and the randomization group timestamp, the randomization group generator 121 also generates a certificate that can be used to prove the validity of the randomization group. For example, the randomization group generator 121 can generate a certificate containing a public verification key signed by the electronic document server 104. The electronic document server 104 can also generate a public key / private key pair. The certificate generation process and the verification process are described in more detail below.
[0049] The randomization group including both the randomization group identifier and the randomization group timestamp is a flexible privacy-preserving monitoring mechanism that allows anonymity as well as unique identification. As described in further detail below, when the browser 107 proves the validity of its randomization group, the certificate can be used to uniquely identify the browser 107. For example, for verification purposes, the browser 107 can transmit the certificate provided by the randomization group generator 121 to the verification system.
[0050] In some embodiments, the publishing domain or electronic document server 104 does not provide metadata other than the randomization group to be stored on the user device 106 on which the browser 107 is stored. This additional restriction further improves user privacy by reducing the amount of data collected and stored, eliminating the possibility of leaking certain types of user data that are not collected and therefore cannot be linked to a specific user or randomization group identifier.
[0051] The analyzer 123 analyzes the randomized group data to monitor user network activity. The analyzer 123 can receive the randomized group identifier and the randomized group age data and the request for data from the electronic document server 104 associated with the analyzer 123, and use the received randomized group identifier and the randomized group age data to perform security functions. For example, the analyzer 123 can detect certain types of fraudulent activities or coordinated abuse of the system or content from the electronic document server 104 based on the randomized group identifier and the randomized group age data. Figure 1 As illustrated, each electronic document server 104 can have a separate analyzer 123 customized for its own needs. In some examples, the electronic document servers 104 can share a centralized analyzer 123, which can be implemented as a remote or separate analysis server or service.
[0052] The system 100 includes one or more third-party independent verification services that a user of the user device 106 or browser 107 can elect to use. The third-party independent verification service independently verifies the privacy attributes of the randomized groups assigned by the publishing service, such as the electronic document server 104 as described above and the trusted domain server 140 as described below. The independent verification of the privacy attributes of the randomized groups is optional for the user and is described in more detail below.
[0053] The verification server 130 is a server independent of the electronic document server 104 that performs verification of the statistical properties of the randomization group identifier and the randomization group identifier and the randomization group age parameter pair. The verification server 130 acts as an independent server that does not publish the randomization group and does not participate in monitoring the randomization group data or otherwise interact with a server such as the electronic document server 104 that monitors and / or analyzes the randomization group data. The verification server 130 allows users to verify that a publishing server such as the electronic document server 104 maintains a guaranteed level of privacy for a particular randomization group. By giving users the opportunity to verify through an independent service that their privacy is being maintained by participating publishing domains, the system 100 encourages user trust and improves the user experience. Additionally, this allows users to discern whether a particular publishing domain is compliant and hold the publishing domain accountable, thereby improving the experience for all users.
[0054] The trusted domain server 140 is a server independent of the electronic document server 104 that is capable of issuing randomization groups to the software 107 in response to the request 108. The trusted domain server 140 issues randomization groups of global scope, meaning that the randomization groups can be provided to requesting servers from different domains and are not limited to use within a specific domain, as is the case with domain-wide randomization groups generated by the electronic document server 104. The trusted domain server 140 acts as a central source of randomization groups, guaranteeing one or more privacy levels for each randomization group generated and assigned. In some embodiments, the trusted domain server 140 is separate from the electronic document server 104, does not share information with any electronic document server 104, and does not provide or host content. For example, the trusted domain server 140 is not involved in the content distribution process and is involved in generating and assigning randomization groups in the system 100 to maintain the privacy of users of the system 100.
[0055] The randomization group generator 142 (RCX generator) is a generator that operates similarly to the randomization group generator 121 described above, but is associated with the trusted domain server 140 rather than an electronic document server.
[0056] Figure 22 is a data flow diagram of an example process 200 for publishing and implementing a privacy-preserving monitoring mechanism. As described below, a domain supporting randomized groups needs to assign randomized groups to users in a manner that protects certain verifiable k-anonymity properties. The operations of process 200 can be implemented, for example, by electronic document server 104, user device 106, verification server 130, trusted domain server 140, and / or third-party server 150. The operations of process 200 can also be implemented as instructions stored on one or more computer-readable media that can be non-transitory, and execution of the instructions by one or more data processing devices can cause one or more data processing devices to perform the operations of process 200.
[0057] As mentioned above Figure 1 As discussed, DCDS 110 may act as an intermediary, forwarding requests 108 from user device 106 to electronic document server 104, and forwarding replies 114 from electronic document server 104 to user device 106. In this particular example, DCDS 110 may transmit data between user device 106 and electronic document server 104, but is not illustrated in the data flow.
[0058] Process 200 begins at stage A-1, where user device 106 provides a request for content to electronic document server 104. For example, software 107 provides a request for content, such as a particular web page, to electronic document server 104 associated with a particular network domain. In addition to web pages, software 107 can also provide requests for third-party content, such as digital components. For example, software 107 can provide a request to third-party server 140. The request in stage A-1 can be a request from user device 106 that is a request for content or a request for a digital component. The request can be provided to electronic document server 104 or third-party server 140. In this particular example, a request labeled Request 1 is provided to electronic document server 104.
[0059] Request 1 is an initial request from software 107. The initial request indicates that software 107 has not previously issued a request to electronic document server 104, that software 107 has not previously published a randomization group, and / or that a previously published randomization group has expired, been reset by a user associated with software 107, or is otherwise invalid. For example, request 1 can indicate that software 107 has not previously issued a request to electronic document server 104, and therefore electronic document server 104 should publish and assign a randomization group to software 107. In some implementations, if system 100 uses a single central publishing entity such as trusted domain server 140, stage A-2 occurs, wherein data indicating request 1 is forwarded from electronic document server 104 to trusted domain server 140. If system 100 uses a separate publishing entity, such as electronic document server 104, stage A-2 does not occur. In this particular example, stage A-2 does not occur because system 100 uses a separate publishing entity. In other examples and implementations, stage A-2 can occur even if system 100 uses a separate publishing entity.
[0060] The process 200 continues at stage B, where the randomization group generator generates / constructs a randomization group 201 in response to receiving data indicative of request 1 and assigns the randomization group to the requesting user device 106 or software 107 .
[0061] If the system 100 uses a single central publishing entity such as the trusted domain server 140, the randomization group generator 142 performs phase B by generating a randomization group in response to receiving data indicating the request 1. If the system 100 uses a separate publishing entity, such as the electronic document server 104, the randomization group generator 121 performs phase B by generating a randomization group in response to receiving data indicating the request 1. In this particular example, phase B is described as being performed by the randomization group generator 121. In other examples and implementations, phase B can be performed by the randomization group generator 142 if, for example, the system 100 uses a single central publishing entity such as the trusted domain server 140.
[0062] The randomization group generator 121 generates each randomization group based on a randomly selected identifier and a timestamp. The identifier can be randomly generated, or the identifier can be selected from existing identifiers already in use when the system 100 has a sufficient number of different browsers 107 running on different user devices 106 or a sufficient number of different user devices 106. In this particular example, the randomization group generator 121 selects a 64-bit identifier as the randomization group identifier.
[0063] The concept of k-anonymity applied to the randomization group identifier guarantees that each randomization group identifier is assigned to at least k browsers 107. For example, k different browsers 107 can be used as a proxy metric to guarantee k different devices 106 corresponding to k different users. Each publishing domain guarantees a specific k for the randomization group it publishes, so that each randomization group identifier is assigned to at least k browsers 107 running on different user devices 106. In order to ensure this level of privacy, the publishing domain must track the number of each randomization group identifier that is published and associate the browser 107 with the specific randomization group identifier that has been assigned. During the assignment process, the browser 107 can be assigned to the cluster of browsers 107 assigned to the specific randomization group identifier. For example, the randomization group generator 121 can randomly assign the browser 107 from which the request 1 is received to the cluster of browsers 107 associated with the specific randomization group identifier.
[0064] Each publishing domain performs an allocation process to provide an even distribution (or within a threshold of an even distribution) of identifiers. For example, the randomized group generator 121 can randomly assign a browser 107 to one of a set of randomized group identifiers that have been generated and associated with the browser 107. The randomized group generator 121 balances the number of different randomized group identifiers with the number of browsers in order to maintain its k-anonymity guarantee for users. For example, the randomized group generator 121 can balance the number of different randomized group identifiers assigned based on the expected and current number of browsers 107 providing requests for content to the corresponding electronic document server 104. In one example, the randomized group generator 121 can generate or receive a threshold number of different randomized group identifiers based on the expected browsers 107 providing requests for content to the electronic document server 104, and can adjust the threshold number of different randomized group identifiers based on the actual current traffic from different browsers 107 providing requests for content to the electronic document server 104. For example, the randomization group generator 121 can increase the number of different randomization group identifiers associated with its domain based on an increase in the number of different browsers 107 or a rate of change in the number. In another example, at the initialization of the system 100, the randomization group generator 121 can receive multiple requests from different browsers 107 running on separate user devices 106, and randomly assign a cluster of k browsers 107 to a specific randomization group identifier to ensure k-anonymity.
[0065] To guarantee k-anonymity, the domain must have enough traffic to make the statistical parameters meaningful and provide anonymization so that a specific randomized group identifier is assigned to k different browsers running on different user devices. The system 100 is able to ensure that the system has enough traffic by setting and adjusting thresholds on network activity levels and ensure that the domain maintains enough traffic to make the randomized group privacy properties verifiable.
[0066] Once a randomization group is generated and assigned to a browser 107, the randomization group can have an expiration time. For example, the randomization group can expire after a predetermined time period, or until a user resets or clears the randomization group. In some examples, the randomization group can expire after a time period specified by a user of the browser 107, a default setting of the browser 107, a publishing domain, a time period determined by the browser 107 based on the network activity of the user of the browser 107 and the habits of the user or users with similar browsing habits to the user. For example, if a user often clears the randomization group associated with the publishing domain on a regular basis, the expiration period of the randomization group can be adjusted so that the randomization group is cleared on a regular basis similar to the schedule that the users of the domain usually use.
[0067] When a randomized group identifier is assigned and expires or is cleared, the randomized group generator 121 can reassign a previously used randomized group identifier. For example, the randomized group generator 121 can use a list of available randomized group identifiers (including previously expired randomized group identifiers) to randomly assign a randomized group identifier to the browser 107, or randomly generate a randomized group identifier based on the needs of the system.
[0068] In some embodiments, the randomization group identifier can become "stale"; in other words, the number of users assigned to a particular randomization group identifier can decrease over time due to users resetting their randomization groups or their randomization groups expiring. Although a publishing domain such as the electronic document server 104 can guarantee k users when publishing a randomization group, the number of users behind the FC can decrease due to the randomization group expiring or being reset. The system 100 can offset this reduction in users per bucket by, for example, storing one or more additional states on the server that indicate the recycling status of randomization group identifier / age buckets with fewer than k users. In addition, the verification server 130 can also detect a lack of activity in a particular user bucket (i.e., a user bucket assigned to a particular randomization group identifier over a period of time). The verification server 130 can also provide feedback to the browser 107 about the distribution of the randomization group identifier and can provide feedback to the randomization group generator 121.
[0069] In addition, the randomized group generator 121 generates a timestamp indicating the time when the randomized group identifier is assigned or generated, as described above with respect to Figure 1 For example, upon receiving a request from the software 107, the randomized group generator generates a randomized group timestamp. The timestamp indicates when the randomly selected identifier is assigned to the publishing domain that randomly assigns the randomized group to the user.
[0070] The user of the browser 107 can reset the randomization group associated with their browser 107 at any time, just as other monitoring mechanisms can be cleared and / or reset. When the user resets the randomization group published by the publishing domain associated with their browser 107, the next request sent from the browser 107 to the publishing domain, such as the electronic document server 104, is received as the initial request for which the randomization group will be assigned to the browser 104.
[0071] In addition to the randomization group assigned to the browser 107, the randomization group generator 121 also generates a certificate containing the randomization group and a public / private key pair. The certificate provided by the randomization group generator 121 contains identifiable information of the browser 107, including a randomization group identifier and a randomization group timestamp. The randomization group generator 121 signs the certificate, which indicates the authenticity of the certificate and its assignment to the browser 107. The signed certificate is only used for verification purposes, as described in further detail below, and will never be transmitted to a domain other than the verification server or can be accessed by a domain other than the verification server. The randomization group generator 121 uses a cryptographic algorithm to generate a public key that may be known to others and a private key that may never be known to anyone other than the browser 107. The public / private key pair is used to prove or confirm the identity of the browser 107 that provides a signed certificate including a randomization group. In particular, the public key is provided to the requesting entity, and the browser 107 can use the private key to cryptographically confirm that it is assigned the certificate. In addition, the certificate can be encrypted using the public / private key pair.
[0072] Verification server 130 verifies that the signature on the certificate is valid. For example, if examplecoolvideoplatform.com issued a certificate, verification server 130 will retrieve the public key of examplecoolvideoplatform.com and verify that the signature on the certificate is valid. For example, the key can be distributed through a technology such as a public key infrastructure (PKI), and the key does not need to be distributed by the electronic document server 104 associated with the randomization group generator 121 that generates the certificate. Verification server 130 then generates a challenge and transmits the challenge to browser 107. For example, the challenge can be a random number or some variable known to verification server 130. Browser 107 can then sign the challenge using its private key and return the signature to verification server 130. Verification server 130 can then use the public verification key published on the signed certificate to verify that the signature is valid.
[0073] The process 200 continues to stage C, where the electronic document server 104 provides the randomization group, the signed certificate, the public / private key pair, and the response to the request to the browser 107. In some embodiments, the electronic document server 104 provides each of the randomization group, the signed certificate, the public / private key pair, and the response to the request to the browser 107 at the same time. In other embodiments, the electronic document server 104 provides one or more of the randomization group, the signed certificate, the public / private key pair, and the response to the request to the browser 107 separately. In some embodiments, the private / public key pair is distributed through a standard technology such as PKI. As described above, the specific stages of the process 200 performed by the electronic document server 104. For example, stage C can be performed by the trusted domain server 140.
[0074] Process 200 continues to stage D, where browser 107 provides a subsequent request and randomization group data to electronic document server 104. The subsequent request can be identical in format to the initial request, and is not required to provide any additional information indicating that the request follows the initial request. Browser 107 detects that the request is provided to a domain associated with the randomization group to which browser 107 has been assigned, and provides the randomization group data, thereby indicating to electronic document server 104 that the request is a subsequent request. The randomization group data allows the domain associated with electronic document server 104 to monitor the activities of different browsers within the domain, while protecting the privacy of the browser user to a greater extent than previously possible.
[0075] The browser 107 provides randomized group data, including a randomized group identifier assigned by the browser 107 to a domain associated with the electronic document server 104 and data representing an obfuscated age of the randomized group of the domain. Specifically, the browser 107 generates a randomized group age bucket to be provided, rather than a randomized group timestamp, to obfuscate the specific identity of the browser 107. For example, the browser 107 uses the function RCX.AgeBucketFn() to calculate the randomized group age bucket. For example, the function can be a bucketing function, such as log2(currentTime-randomized cohort.Timestamp), where currentTime represents the current time and randomized cohort.Timestamp represents the randomized group timestamp.
[0076] Because the browser 107 provides subsequent requests to both the randomized group identifier and the randomized group timestamp, the publishing domain (in this example, the electronic document server 104 of the publishing domain) is responsible for ensuring that at least k users or user agents of the browser 107 are assigned to each pair of randomized group identifier and randomized group age bucket to satisfy its k-anonymity guarantee. In some embodiments, the publishing domain can guarantee different levels of k-anonymity with respect to the randomized group identifier and the randomized group identifier and randomized group timestamp pairs.
[0077] Optionally, the user of the browser 107 can choose to verify the statistics and / or privacy attributes of the randomization group assigned to the browser 107. Any third party (such as the electronic document server 104 or the trusted domain server 140) that is not the publishing domain of the randomization group to be verified can maintain a verification server 130, and the user can choose to direct their browser 107 to any verification server 130. In some embodiments, the electronic document server 104 can maintain a verification server 130 to verify the statistics and / or privacy attributes of the randomization group published by other electronic document servers 104 associated with other domains. In some embodiments, the browser 107 can automatically request verification from the verification server 130 without being instructed by the user. For example, the system 100 can require participating browsers 107 to periodically request verification from a randomly selected qualified verification server 130. In some embodiments, the browser 107 does not request verification unless instructed by the user.
[0078] The process 200 continues to stage E, where the software 107 provides the certificate to the verification server 130. For example, the browser 107 provides the certificate generated by the randomization group generator 121 in stage B to the verification server 130. The browser 107 can also provide the public key to the verification server for certification purposes. In some embodiments, the browser 107 provides randomization group information, such as a randomization group identifier and a randomization group age bucket, instead of a certificate, which does not obfuscate the randomization group age.
[0079] Process 200 continues to stage F, where verification server 130 performs a verification process to verify the statistical and / or privacy properties of the randomized group data provided by browser 107 .
[0080] The verification server 130 verifies that the set of randomized group identifiers to which the verification server 130 has access is uniformly distributed. For example, the verification server 130 can use the certificate to determine the randomized group identifier and / or the randomized group age bucket. The verification server 130 can then compare the randomized group identifier and / or the randomized group age bucket with the list of randomized group identifiers and / or randomized group age buckets to which the verification server has access to determine whether the number of browsers running on different user devices assigned to each different randomized group identifier within an appropriate range is uniformly distributed or uniformly distributed within a threshold distance. As described above, the range can be within a specific domain or globally within a network of participating domains. For example, the verification server 130 can maintain a list of randomized group identifiers and randomized group age bucket pairs it receives from the browsers 107 participating in its verification service. The verification server 130 can then determine whether the number of different browsers running on different user devices assigned to each different randomized group identifier within a specific age bucket and / or domain is normally distributed. In some implementations, the verification server 130 determines whether the number of different browsers running on different user devices assigned to each different randomized group identifier within a particular domain is normally distributed without the additional requirement of being within the same age bucket.
[0081] The browser 107 can automatically request authentication at certain intervals from an authentication service such as the authentication server 130. For example, the browser 107 can request authentication from the authentication server 130 every week to ensure that the electronic document server 104 complies with its k-anonymity guarantee obligations and delivers the level of privacy it promised to users.
[0082] Alternatively, during a content delivery process where the randomization group is globally scoped and not assigned by each domain, the electronic document server 104 from which the browser 107 requests content can include a certification request in its response to the initial request from the browser 107 .
[0083] The browser 107 provides such a randomization group to each electronic document server 104 on each request, regardless of which domain the electronic document server 104 is associated with. Using a global-scope randomization group prevents malicious domains from colluding to combine randomization groups across domains in a way that creates more traceable entropy, thereby reducing the risk of compromising user privacy due to collusion between domains. However, the global-scope randomization group is globally readable, which does not eliminate the potential for abuse due to strategies such as replay attacks. For example, a malicious actor could purchase traffic to a domain, obtain a randomization group (including a randomization group identifier and a randomization group timestamp), and then use the observed distribution to attack a different domain. Because the entity that generates the randomization group is separate and independently different from the server from which the content is requested, such as the electronic document server 104 or the third-party server 150, the server from which the content is requested will need to be able to perform a proof step to verify that the browser 107 that provides the randomization group identifier and the randomization group timestamp has actually been assigned a randomization group from a trusted server such as the trusted domain server 140.
[0084] The attestation request can specify that the browser 107 should request the verification server 130 to determine whether the randomization group identifier and the randomization group age bucket provided to the electronic document server 104 with the browser 107's request for content were actually assigned to the browser 107 by a trusted server, such as the trusted domain server 140. For example, the electronic document server 104 can request the verification server 130 to require the browser 107 to attest to its identity as the browser 107 to which the certificate and associated randomization group information were assigned by a trusted server, such as the trusted domain server 140. The verification server 130 can facilitate this attestation process by generating an anonymous certificate that is different from the certificate provided and signed by the randomization group generator 121, which attests that a particular certificate is assigned to a particular browser 107. This process is not relevant where the randomization group is locally scoped for each domain such that each domain issues a randomization group to the browser 104, and thus will be able to determine whether the browser 107 is the browser to which it assigned the randomization group based on the signed certificate issued by the randomization group generator 121. For example, if the randomization group generator 121 associated with the electronic document server 104 generated the randomization group, it will be able to verify its own signature on the certificate it received associated with the randomization group or by using public key encryption using the public / private key pair generated as described above with respect to phase B.
[0085] In one illustrative example, a user visits ExampleNewsWebsite.com using browser 107. The user reads an article that links to a particular video on ExampleVideoHostingPlatform.com. In this example, browser 107 has previously been assigned a randomization group and certificate from trusted domain server 140, and therefore browser 107 sends a request for the particular video to ExampleVideoHostingPlatform.com associated with electronic document server 104. In the request, browser 107 includes its assigned randomization group identifier and randomization group age bucket.
[0086] Because the global scope randomization group assigned to browser 107 by the trusted server is not generated by randomization group generator 121 associated with electronic document server 104, ExampleVideoHostingPlatform.com can request browser 107 to prove its identity and forward the request to verification server 130 for completion. Electronic document server 104 provides secret X to browser 107. The secret can have any value and can be, for example, a randomly generated 16-bit value. Browser 107 provides secret X and certificate provided by randomization group generator 142 to verification server 130 to request verification server 130 to perform the certification process. Browser 107 can mask the value of X from verification server 130 to prevent secret X from being linked to browser 107. For example, browser 107 can mask X from verification server 130 using a partially blinded signature scheme. Verification server 130 then receives browser 107's certificate and blinded secret X and generates an anonymous certificate indicating a randomization group identifier and a randomization group age bucket based on the certificate. Verification server 130 then signs this anonymous certificate and returns the signed, anonymous certificate and X to electronic document server 104 that issued the attestation request. Electronic document server 104 can compare the randomized group information in the signed anonymous certificate from verification server 130 with the randomized group identifier and randomized group age bucket that electronic document server 104 received from browser 107. If the information matches, browser 107 has successfully attested that its identity is that of the browser to which trusted domain server 140 assigned the randomized group information provided to electronic document server 104.
[0087] In some embodiments, if multiple verification servers 130 collaborate and / or share resources, a more robust verification process can be performed. For example, the verification servers 130 can determine whether multiple randomized group identifiers within a particular age bucket comply with the k-anonymity guarantee provided by the electronic document server 104 and designated by the system 100 to be considered compliant, etc., in the context of a larger group that better represents the total population of participating browsers and domains. The verification servers 130 can use, for example, the total number of unique public keys provided with the certificates received across each verification server can be calculated, thereby allowing the user to verify that there are approximately at least k users with the same randomized group identifier.
[0088] In another example implementation, electronic document server 104 can return a unique tracking URL to browser 107 instead of secret X. Browser 107 can then follow the unique tracking URL and publish the signed anonymous certificate from verification server 130 to the destination at the unique tracking URL.
[0089] Independent verification servers 130 can collaborate to represent that certain statistical properties of the entire population of randomized groups are satisfied and privacy guarantees are maintained at scale. Due to the nature and scale of the number of browsers running on different user devices representing different users, the verification servers 130 provide a level of protection in the form of checkpoints that are difficult to spoof, rather than formal proofs of correctness or legitimacy.
[0090] In some examples, a malicious publisher can use a subset of the bits of the randomized group identifier to encode sensitive data about the user or otherwise embed information in a hidden manner. The system 100 can perform a consistency check at the bit level to check the randomized group identifier and determine whether the randomized group identifier is uniformly selected and distributed. For example, the verification server 130 can randomly generate a mask, perform a bitwise AND operation on all randomized group identifiers and re-aggregate. If the randomized group identifier has been uniformly selected and assigned, the result should also be uniform. In some embodiments, the verification server 130 remembers the randomized group identifier previously assigned to each user and tests the bit-level correlation over time. In some embodiments, the verification server 130 performs a test to ensure that two or more browsers 107 running on different user devices 106 are not consistently assigned to the same randomized group identifier.
[0091] The process 200 continues to stage G, where the verification server 130 provides the results of the verification process to the browser 107. The verification server 130 can provide the statistics and / or privacy properties of the randomized group identifier and the randomized group age bucket to the browser 107 in raw quantities. For example, the verification server 130 can provide the distribution of the randomized group identifier assignments or the number of different randomized group browsers running on different user devices to the browser 107 associated with the randomized group identifier. The browser 107 can then compare the distribution of the raw number of randomized group identifier assignments to a threshold deviation from a uniform distribution, or compare the number of different randomized group browsers running on different user devices of the browser 107 associated with the randomized group identifier to a threshold number k of different browsers running on different user devices that should be associated with the randomized group identifier to ensure k-anonymity. The browser 107 can provide an indication to the user of the browser 107 when one or more thresholds are not met. For example, if the number of different randomized group browsers running on different user devices from the browser 107 associated with the randomized group identifier does not meet a threshold number k of different browsers running on different user devices, the browser 107 can display a visual message, play audio, create a vibration, etc. to indicate to the user that the threshold has not been met.
[0092] In addition to the advantage of providing a monitoring mechanism that provides users with an independently verifiable level of privacy, randomized groups also provide protection for content providers and hosts of the system. One form of abuse in content distribution systems is through the manipulation of engagement statistics. For example, in a scheme where content providers pay per interaction (e.g., a pay-per-click system), users can be incentivized to click on specific content items to increase the payment required by the content provider. Additionally, on video content platforms, users can collude to coordinate the manipulation of video content recommendation systems by artificially increasing the popularity of certain videos by providing fraudulent views. Randomized groups provide a monitoring mechanism that allows detection of such abuse.
[0093] From the user's perspective, the randomized groups act similarly to traditional monitoring mechanisms while providing a higher level of privacy. Thus, the described system for a privacy-preserving monitoring mechanism requires little change to the user's experience while improving their privacy and reducing the likelihood that user information will be compromised.
[0094] Randomization groups rely on the use of statistical parameters to guarantee user privacy and leverage a greater amount of entropy for the purpose of performing abuse detection without violating user privacy. For k=1 in a k-anonymity context, randomization groups provide the same level of abuse detection as traditional tracking mechanisms (such as cookies). As k increases, randomization groups provide users with a higher level of privacy while still being useful for abuse detection. Additionally, randomization groups do not require asserting trust in a first-party context, where randomization groups are locally scoped and limited to use by the domain that published the randomization group.
[0095] Once the browser 107 has provided the randomization group information to the electronic document server 104, the electronic document server 104 can perform security functions such as statistical analysis to detect abuse of the system 100. For example, when the randomization group is local in scope to a particular domain and when the randomization group is global in scope, each analyzer 123 of a particular electronic document server 104 can perform a statistical analysis specified by the electronic document server 104 associated with the analyzer 123.
[0096] The electronic document server 104 is able to protect its domain from participant abuse in a manner that preserves user privacy by using the described monitoring mechanism, namely the randomized group. Because the system guarantees anonymity of k, where k is adjustable, the use of the randomized group provides a system that provides continuous protection from participant abuse that maintains a higher level of privacy protection than is currently possible.
[0097] The analyzer 123 of the electronic document server 104 monitors anonymous user network activities by using the randomized group information, and detects signs of engaging in abuse or collusion based on the statistical properties of the distribution of the randomized group identifiers and the activities associated with the randomized group identifiers. For example, the analyzer 123 can detect statistically abnormal behavior, such as irregular resource requests or abnormal activities from user groups associated with specific randomized group identifiers and / or age buckets.
[0098] For a given participating domain, the number of clicks for each randomized group identifier should be evenly distributed. Inconsistencies in the number of clicks can indicate anomalous activity. For example, deviations greater than a threshold deviation of the number of clicks associated with the randomized group identifiers, or spikes in activity (the number or speed of deviations) over a period of time associated with one or more randomized groups can statistically indicate anomalous activity that can indicate abuse. For example, the analyzer 123 can determine a statistical test to verify statistically anomalous activity through metrics such as the risk ratio of the activity and its impact on domain resources compared to the risk and potential damage caused by the malicious activity.
[0099] The analyzer 123 can calculate the age distribution of the randomized group age buckets received in combination with the randomized group identifier and test for anomalies against, for example, a known global background distribution. For example, the analyzer 123 can detect that the amount of activity from browsers having randomized group age bucket values below a threshold age value is equal to or below a threshold amount and determine that there is an abnormal amount of activity.
[0100] Analyzer 123 can detect coordinated attacks involving groups of malicious users by performing fraud detection using an anti-masquerading algorithm, using provable limits on undetectable fraud to provide an upper limit on the effectiveness of fraud. For example, analyzer 123 can limit the probability that two or more randomized group identifier / randomized group age bucket pairs visit the same set of long-tail websites within a certain time window. In another example, analyzer 123 can construct a bipartite graph between users and websites, and then try to find a bipartite graph (a specific type of bipartite graph in which each vertex of the first set is connected to each vertex of the second set) corresponding to clusters of users acting in a coordinated manner.
[0101] The analyzer 123 is capable of detecting entropy created fraudulently for the purpose of tracking one or more users. For example, a malicious actor could attempt to assign a single browser 107 to a randomized group identifier, where the other k-1 browsers 107 are bots. However, the electronic document server 104 assigns randomized group identifiers to users in a random manner, and therefore attacks of this nature would be difficult to conduct without the attacker gaining access to the server. A malicious or compromised domain electronic document server 104 may be able to conduct such fraud against multiple users associated with the browser 107, but it would be difficult for the malicious domain to perform any meaningful scale of attack while evading detection by an independent verification server such as the verification server 130. In addition to providing notification of such randomized group identifier / age bucket pairs to the participating domain's electronic document servers 104, enabling the domain to prevent abuse from the identified randomized group identifier, such scaled attacks can be effectively contained by allowing the verification server 130 to perform abuse detection and prevention, allowing the verification server 130 to detect and block or delete randomized groups assigned to botnets over a period of time.
[0102] When the analyzer 123 detects statistically abnormal behavior, the analyzer 123 can set or adjust limits to prevent a particular user from requesting resources from a particular electronic document server 104 or requesting resources generally over a period of time. For example, the analyzer 123 can detect that a particular randomized group identifier / randomized group age bucket pair is being provided to the electronic document server 104 and is associated with an abnormal amount of activity over the past five minutes, and stop the particular randomized group identifier / randomized group age bucket from requesting resources from the electronic document server 104. Thus, the system 100 can reduce the amount of resources used to facilitate fraudulent activity. However, when the randomized group is an independent, domain-wide randomized group, each verification server 130 can only access the randomized group identifier and randomized group age bucket of users who have chosen to use that particular verification server 130 as their independent verification service. In some embodiments, whether the randomized group is domain-wide or global-wide, the verification server 130 is federated and grouped with other verification servers 130 to access a threshold number of randomized groups in order to provide statistically meaningful results to the browser 107.
[0103] When the randomization group is global in scope, the analyzer 123 can communicate with the trusted domain server 140 to identify suspected fraud or anomalous activity associated with the global-scope randomization group identifier / randomization group age bucket pairs, which can be communicated to other electronic document servers 104 to monitor and prevent abuse across participating domains.
[0104] Figure 3A and 3B 300 and 350 are swim lane diagrams illustrating example processes 300 and 350 for publishing and implementing a monitoring mechanism for privacy protection of network activities performed across different domains. The operations of processes 300 and 350 can be implemented, for example, by user device 106 and / or software 107, electronic document server 104, DCDS 110, verification server 130, trusted domain server 140, and / or third-party server 150. All of the processes in processes 300 and 350 can also be implemented as instructions stored on one or more computer-readable media, which can be non-transitory, and execution of the instructions by one or more data processing devices can cause one or more data processing devices to perform the operations of processes 300 and 350.
[0105] Reference now Figure 3A , process 300 begins at step 1, where an initial request from software 107 is provided to electronic document server 104. In some examples, the request can be forwarded to electronic document server 104 via DCDS 110, which Figure 3A104. For example, ExampleImageHostingPlatform.com associated with the ExampleImageHostingPlatform network domain can receive a request for an image from browser 107. In this example, browser 107 has not yet been assigned a randomization group from electronic document server 104 of ExampleImageHostingPlatform, and thus the request is an initial request. In another example, if browser 107 was previously assigned a randomization group from electronic document server 104 of ExampleImageHostingPlatform, but the randomization group has expired or has been cleared by the user of browser 107, then the request is also considered an initial request.
[0106] The process 300 continues to step 2, where the electronic document server 104 generates a randomization group, which includes a randomization group identifier, a randomization group timestamp, a signed certificate, and a public / private key pair, and assigns the randomization group to the software 107. For example, the randomization group generator 121 of the electronic document server 104 can generate a randomization group, a signed certificate, and a public / private key pair, and assign the randomization group to the browser 107, as described above with respect to Figure 2 As described above, the electronic document server 104 is associated with a particular network domain.
[0107] Process 300 continues to step 3, where the electronic document server 104 provides a reply to the software 107 with the requested content and the randomized group, which includes the randomized group identifier, the randomized group timestamp, the signed certificate and the public / private key pair, as described above with respect to Figure 2 For example, the electronic document server 104 can transmit to the browser 107 a response including the requested image, the randomization group, the signed certificate, and the public / private key pair.
[0108] Process 300 continues to step 4, where software 107 provides a subsequent request to electronic document server 104, which previously published the randomization group to software 107. Software 107 includes the randomization group identifier and age bucket with its request. For example, electronic document server 104 for domain ExampleImageHostingPlatform.com can receive a request for a second image from browser 107 along with the assigned randomization group identifier and randomization group age bucket.
[0109] Optionally, a user of the software 107 can select a verification server to independently verify the privacy and / or statistical properties of the randomized group identifier associated with the software 107 , as described with respect to steps 5 , 6 , and 7 .
[0110] Process 300 continues to step 5, where the software 107 provides its certificate to the verification server 130. For example, the browser 107 can provide its signed certificate generated by the randomized group generator 121 to the verification server 130.
[0111] Process 300 continues to step 6, where verification server 130 performs a verification process based on the certificate. For example, verification server 130 can perform a verification process based on the certificate to determine statistical and / or privacy protection properties of the randomized group identifier provided in the certificate.
[0112] Optionally, the verification server 130 can also authenticate the provided certificate before performing the verification process to determine that the provided randomized group identifier is assigned to the software 107 that is providing the certificate. For example, the verification server 130 can use a public / private key pair to determine that the provided randomized group identifier is assigned to the browser 107 that is providing the certificate.
[0113] Process 300 continues to step 7, where verification server 130 provides the results of the verification process to software 107. For example, verification server 130 can output to browser 107 raw counts or results of statistical analysis and comparisons to thresholds to determine whether statistically abnormal or fraudulent activity has been detected, or whether a guaranteed level of privacy for browser 107 is being maintained.
[0114] Reference now Figure 3B , the process 350 for publishing and implementing a monitoring mechanism for privacy protection of network activities performed across different domains is performed by the central trusted domain server 140 rather than the separate electronic document server 104. The randomization group published by the trusted domain server 140 is global in scope. The process 350 begins at step 1, where an initial request from software 107 is provided to the electronic document server 104. In some examples, the request can be made by Figure 3B The DCDS 110 (not shown) is forwarded to the electronic document server 104.
[0115] Process 350 continues to step 2, where electronic document server 104 forwards the request information to trusted domain server 140. For example, electronic document server 104 can receive a request for several lines of text from its domain ExampleTextandImageHostingPlatform.com. In this example, browser 107 has not yet been assigned a randomization group and does not provide a randomization group to electronic document server 104, and therefore the request is an initial request.
[0116] The process 350 continues with step 3, where the trusted domain server 140 generates a randomization group, including a randomization group identifier, a randomization group timestamp, a certificate, and a public / private key pair, and assigns the randomization group to the software 107. For example, the randomization group generator 142 of the trusted domain server 140 can generate a randomization group, a signed certificate, and a public / private key pair, and assign the randomization group to the browser 107, as described above with respect to Figure 2 As described above, the trusted domain server 140 is not associated with a specific network domain and generates a global-scope randomization group that can be used across different domains.
[0117] Process 350 continues to step 4A, where electronic document server 104 provides a reply to software 107. For example, electronic document server 104 can transmit a response to browser 107 that includes the requested text.
[0118] Process 350 continues to step 4B, where the trusted domain server 140 provides the randomization group, including the randomization group identifier, the randomization group timestamp, the certificate, and the public / private key pair to the software 107. For example, the trusted domain server 140 can send the randomization group, the signed certificate, and the public / private key pair to the browser 107.
[0119] In some embodiments, steps 4A and 4B occur simultaneously. In some embodiments, steps 4A and 4B occur asynchronously.
[0120] Process 350 continues to step 5, where software 107 provides subsequent requests to electronic document server 104. Because trusted domain server 140 has assigned a globally scoped randomization group to software 107, software 107 includes the randomization group identifier and the age bucket with its request to electronic document server 104. For example, electronic document server 104 for domain ExampleTextandImageHostingPlatform.com can receive a request for an image from browser 107 along with the assigned randomization group identifier and the randomization group age bucket.
[0121] Optionally, a user of the software 107 can select a verification server to independently verify the privacy and / or statistical properties of the randomized group identifier associated with the software 107 , as described with respect to steps 6 , 7 , and 8 .
[0122] Process 350 continues to step 6, where the software 107 provides its certificate to the verification server 130. For example, the browser 107 may provide its signed certificate generated by the randomized group generator 142 to the verification server 130.
[0123] Process 350 continues to step 7, where verification server 130 performs a verification process based on the certificate. For example, verification server 130 can perform a verification process based on the certificate to determine statistical and / or privacy-preserving properties of the randomized group identifier provided in the certificate.
[0124] Optionally, the verification server 130 can also authenticate the provided certificate before performing the verification process to determine that the provided randomized group identifier is assigned to the software 107 that provided the certificate. For example, the verification server 130 can use a public / private key pair to determine that the provided randomized group identifier is assigned to the browser 107 that provided the certificate. This protocol prevents the verification server 130 from considering statistical data from browsers 107 with forged or stolen certificates, improving the stability of the statistical activity data collected and reducing resources selected for fraudulent use.
[0125] Process 350 continues to step 8, where verification server 130 provides the results of the verification process to software 107. For example, verification server 130 can output raw numbers or results of statistical analysis and comparisons to thresholds to browser 107 to determine whether statistically abnormal or fraudulent activity has been detected, or whether a guaranteed level of privacy for browser 107 is being protected.
[0126] Figure 4 4 is a flow chart illustrating an example process 400 for publishing and implementing a monitoring mechanism for privacy protection. The operations of process 400 can be implemented, for example, by user device 106 and / or software 107, electronic document server 104, DCDS 110, verification server 130, trusted domain server 140 and / or third-party server 150. Process 400 can also be implemented as instructions stored on one or more computer-readable media, which can be non-transitory, and the execution of the instructions by one or more data processing devices can cause one or more data processing devices to perform the operations of process 400.
[0127] Process 400 begins by receiving a request for digital content from a domain from an application on a user device of a user (402). For example, electronic document server 104 can receive a request for digital content from a domain associated with electronic document server 104 from browser 107 of user device 106, as described above with respect to Figure 1 , 2 and as described in 3A-3B.
[0128] Process 400 continues by assigning to the application at a first time a randomized group generated based on the randomly selected identifier and a timestamp indicating the first time the randomized group was assigned to the application (404). For example, electronic document server 104 can assign a local scope randomized group associated with a domain to browser 107, as described above with respect to Figure 1 , 2 In another example, the trusted domain server 140 can assign a global randomization group to the browser 107 that can be used, as described above with respect to Figure 1 , 2 and as described in 3B.
[0129] Process 400 continues by providing, at a first time, to the application (i) a certificate corresponding to the randomly selected identifier and a timestamp and signed with a unique public key and (ii) a unique private key corresponding to the unique public key, wherein the randomly selected identifier is also assigned to at least a threshold number of other applications executed on other user devices within a predetermined time period (406). For example, the electronic document server 104 or the trusted domain server 140 provides the randomization group and / or the certificate to the browser 107. If the electronic document server 104 generates a randomization group of local scope, the electronic document server 104 also provides the requested digital content to the browser 107. If the trusted domain server 140 generates a randomization group of global scope, the electronic document server 104 provides the requested digital content to the browser 107. Step 406 can be as described above with respect to Figure 1 , 2 Perform as described in 3A-3B.
[0130] In some implementations, process 400 can include receiving, from the application, a second request for digital content from the domain, and providing, by the application, to the domain at a second time, an obfuscated identifier generated based on the randomly selected identifier and a randomized group age bucket indicating an age range of a cookie containing an age of the randomized group, wherein the age of the randomized group is calculated based on a difference between the second time and the first time. For example, browser 107 can calculate the randomized group age bucket based on the randomized group and / or certificate provided by electronic document server 104 or trusted domain server 140, as described above with respect to Figure 1 , 2and 3A-3B. In some embodiments, the fuzzy identifier includes a randomly selected identifier and a randomized group age bucket. In some embodiments, the fuzzy identifier includes data and parameters derived from the randomly selected identifier and / or the randomized group age bucket. For example, the fuzzy identifier can include parameters derived by applying a multiplier, adding a constant, or applying some other set of operations to the randomly selected identifier and / or the randomized group age bucket. The fuzzy identifier can also include data and parameters in addition to the randomly selected identifier and the randomized group age bucket.
[0131] In some embodiments, assigning the randomized group to the application includes assigning the randomized group to the application by the domain, wherein the randomly selected identifier is assigned by the domain to at least a threshold number of other applications, wherein the randomly selected identifier is a randomly generated identifier selected from two or more randomly generated identifiers, and wherein the unique public key is generated by the domain. For example, the electronic document server 104 can generate and assign the randomized group to the browser 107 and ensure that the randomized group identifier is assigned to at least k-1 other browsers 107 by the electronic document server 104, wherein the randomized group identifier is selected between the two or more randomly generated identifiers, and wherein the unique public key is generated by the electronic document server 104, as described above with respect to Figure 1 , 2 and as described in 3A.
[0132] In some embodiments, process 400 includes providing, by the application, a certificate to a verification system, and verifying, by the verification system, that the randomized group is assigned to at least a threshold number of people. For example, verification server 130 can determine the statistical and / or privacy properties of the randomized group indicated in the certificate, as described above with respect to Figure 1 , 2 and as described in 3A.
[0133] In some embodiments, assigning the randomized group to the application includes assigning, by the central server, the randomized group to the application, wherein the randomly selected identifier is assigned to at least a threshold number of other applications, wherein the randomly selected identifier is a randomly generated identifier selected from two or more randomly generated identifiers, and wherein a unique public key is generated by the central server. For example, the trusted domain server 140 can generate and assign the randomized group to the browser 107 and ensure that the randomized group identifier is assigned to at least k-1 other browsers 107, wherein the randomized group identifier is selected from two or more randomly generated identifiers, and wherein the unique public key is generated by the trusted domain server 140, as described above with respect to Figure 1 , 2 and as described in 3B.
[0134] In some embodiments, process 400 includes: providing, by an application, a request for digital content from a second domain different from the first domain and a randomized group; in response to providing the request for digital content from the second domain, receiving, by the application, an attestation request including a challenge from the second domain; and providing, by the application, a certificate to a verification system, the certificate triggering the verification system to (i) create a fuzzy certificate including a randomly selected identifier, the randomized group age bucket, and the challenge, (ii) sign the fuzzy certificate, and (iii) provide the fuzzy certificate to the second domain, wherein the challenge is shielded from the verification server using a blinding scheme. For example, verification server 130 can generate an anonymous certificate based on the attestation request, as described above with respect to Figure 1 , 2 and as described in 3B.
[0135] In some embodiments, process 400 includes detecting, by the domain, unusual activity associated with the randomly selected identifier based on the received randomized group age bucket and at least one of: a number of interactions associated with the randomly selected identifier, a randomized group age distribution, and a probability distribution associated with specific interactions and specific time periods. For example, analyzer 123 of electronic server 103 can detect unusual activity associated with the randomized group identifier and / or the randomized group identifier / age bucket pair, as described above with respect to Figure 1 , 2 and as described in 3A-3B.
[0136] The present technology therefore allows Internet activity to be tracked (e.g., to help prevent abuse or fraud) while helping to improve individual user privacy. The use of randomized group identifiers that are assigned to multiple users and age buckets rather than specific timestamps means that it is not possible to identify individual users associated with a given randomized group. The privacy of that user is therefore improved. At the same time, the use of randomized group identifier / age bucket pairs provides sufficient statistical information over time to identify anomalies in Internet activity that potentially indicate fraud. In addition, by using verification systems and challenges issued by third-party domains, the third-party domains are able to verify the authenticity of the randomized group without obtaining identifiable information about individual users. The present technology therefore provides a stable way to combat abusive or fraudulent Internet activity while helping to improve individual user privacy.
[0137] Figure 55 is a block diagram of an example computer system 500 that can be used to perform the above operations. System 500 includes a processor 510, a memory 520, a storage device 530, and an input / output device 540. Each of components 510, 520, 530, and 540 can be interconnected, for example, using a system bus 550. Processor 510 can process instructions for execution within system 500. In some embodiments, processor 510 is a single-threaded processor. In another embodiment, processor 510 is a multi-threaded processor. Processor 510 can process instructions stored in memory 520 or on storage device 530.
[0138] Memory 520 stores information within system 500. In one implementation, memory 520 is a computer readable medium. In some implementations, memory 520 is a volatile memory unit. In another implementation, memory 520 is a non-volatile memory unit.
[0139] The storage device 530 can provide mass storage for the system 500. In some embodiments, the storage device 530 is a computer-readable medium. In various embodiments, the storage device 530 can include, for example, a hard disk device, an optical disk device, a storage device shared by multiple computing devices (e.g., a cloud storage device) over a network, or some other mass storage device.
[0140] Input / output device 540 provides input / output operation for system 500.In some embodiments, input / output device 540 can include network interface device, such as Ethernet card, serial communication device, such as, RS-232 port, and / or wireless interface device, such as, one or more of 802.11 card.In another embodiment, input / output device can include driver device configured to receive input data and send output data to external device 560, such as, keyboard, printer and display device.However, other embodiments can also be used, such as mobile computing device, mobile communication device, set-top box TV client device etc.
[0141] Although already Figure 4 An example processing system is described in the specification, but the subject matter and implementation of the functional operations described in this specification can be implemented in other types of digital electronic circuits, or in computer software, firmware or hardware including the structures disclosed in this specification and their structural equivalents, or in a combination of one or more of them.
[0142] The embodiments of the subject matter and operations described in this specification can be implemented in digital electronic circuits or in computer software, firmware or hardware, including the structures disclosed in this specification and their structural equivalents, or in one or more combinations thereof. The embodiments of the subject matter described in this specification can be implemented as one or more computer programs, that is, one or more circuits of computer program instructions, which are encoded on one or more computer storage media for execution by a data processing device or for controlling the operation of a data processing device. Alternatively or in addition, program instructions can be encoded on an artificially generated propagation signal, for example, a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device for execution by a data processing device. A computer storage medium can be or be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. In addition, although a computer storage medium is not a propagation signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagation signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (eg, multiple CDs, disks, or other storage devices).
[0143] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0144] The term "data processing apparatus" encompasses all kinds of apparatus, devices and machines for processing data, including, for example, a programmable processor, a computer, a system on a chip, or a plurality or combination of the foregoing. The apparatus can include dedicated logic circuits, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the apparatus can also include code that creates an execution environment for the computer program involved, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.
[0145] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or code portions). A computer program can be deployed to be executed on one computer or on multiple computers located in one location or distributed across multiple locations and interconnected by a communication network.
[0146] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuits, and the apparatus can also be implemented as special purpose logic circuits, for example, FPGAs (field programmable gate arrays) or ASICs (application specific integrated circuits).
[0147] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as a magnetic disk, a magneto-optical disk, or an optical disk, or be operably coupled to receive data from or transmit data to or both of the one or more mass storage devices. However, a computer does not need to have such a device. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few examples. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0148] To provide interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having: a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user; and a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, voice, or tactile input. In addition, the computer can interact with the user by sending documents to or receiving documents from a device used by the user; for example, sending a web page to a web browser on a user's client device in response to a request received from the web browser.
[0149] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, such as a data server, or includes an intermediate component, such as an application server, or includes a front-end component, such as a client computer with a graphical user interface or a web browser, or includes any combination of one or more such back-end, intermediate or front-end components, through which a user can interact with implementations of the subject matter described in this specification. The components of the system can be interconnected by digital data communication in any form or medium, such as a communication network. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), inter-networks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0150] The computing system can include a client and a server. The client and the server are usually remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by means of computer programs running on the respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to the client device (e.g., for the purpose of displaying data to a user interacting with the client device and receiving user input from the user). Data generated at the client device (e.g., the result of user interaction) can be received from the client device at the server.
[0151] Although this specification contains many specific implementation details, these should not be interpreted as limitations on the scope of any invention or the claimed content, but as descriptions of features specific to a particular embodiment of a particular invention. Certain features described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. In addition, although features may be described above as working in certain combinations and even initially claimed as such, one or more features from the claimed combination can be removed from the combination in some cases, and the claimed combination can be directed to a sub-combination or a variation of the sub-combination.
[0152] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed, to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0153] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions described in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing may be advantageous.
Claims
1. A method for privacy-preserving network activity monitoring, comprising: receiving, from an application on a user device of a user, a request for digital content from a first domain; assigning a randomized group to the application at a first time, the randomized group being constructed based on a randomly selected identifier and a timestamp indicating the first time the randomized group was assigned to the application; providing, at the first time, to the application (i) a digitally signed certificate corresponding to the randomly selected identifier and the timestamp, and (ii) a unique public key and a corresponding unique private key associated with the certificate; receiving, from the application, a second request for digital content from the first domain; as well as receiving, from the application at a second time, an obfuscated identifier corresponding to the randomly selected identifier and a randomized group age bucket indicating an age range of a cookie containing an age of the randomized group, wherein the age of the randomized group is calculated based on the difference between the second time and the first time, and Wherein, within a predetermined time period of assigning the randomized group to the application, the randomly selected identifier is also assigned to at least a threshold number of other applications executed on other user devices.
2. The method according to claim 1, wherein: Assigning a randomized group to the application comprises: assigning, by the first domain, the randomization group to the application, and wherein the randomly selected identifier is assigned to the at least threshold number of other applications, wherein the randomly selected identifier is a randomly generated identifier selected from two or more randomly generated identifiers, and The unique public key is generated by the first domain.
3. The method according to claim 1, wherein: Assigning the randomized group to the application comprises: assigning the randomized group to the application by a central server; wherein the randomly selected identifier is assigned to the at least threshold number of other applications, wherein the randomly selected identifier is a randomly generated identifier selected from two or more randomly generated identifiers, and Wherein, the unique public key is generated by the central server.
4. The method according to claim 1, further comprising: providing, by the application, a request for digital content from a second domain different from the first domain and the randomized group; receiving, by the application, a certification request including a challenge from the second domain in response to providing the request for digital content from the second domain; as well as providing the digitally signed certificate to a verification system by the application, the digitally signed certificate triggering the verification system to (i) create an obfuscated certificate containing the randomly selected identifier, the randomized group age bucket, and the challenge, (ii) sign the obfuscated certificate, and (iii) provide the obfuscated certificate to the second domain, Therein, the challenge is shielded from the verification system using a blinding scheme.
5. The method according to claim 1, further comprising: The application provides the digitally signed certificate to the verification system; as well as The randomized group is verified, by the verification system, to be assigned to at least a threshold number of applications.
6. The method according to claim 1, further comprising: The first domain detects anomalous activity associated with the randomly selected identifier based on the received randomized cohort age bucket and at least one of: a number of interactions associated with the randomly selected identifier, a randomized cohort age distribution, and a probability distribution associated with interactions and time periods.
7. A system for privacy-preserving network activity monitoring, comprising: one or more processors; as well as One or more memory elements including instructions that, when executed, cause the one or more processors to perform operations comprising: receiving, from an application on a user device of a user, a request for digital content from a first domain; assigning a randomized group to the application at a first time, the randomized group being constructed based on a randomly selected identifier and a timestamp indicating the first time the randomized group was assigned to the application; providing, at the first time, to the application (i) a digitally signed certificate corresponding to the randomly selected identifier and the timestamp, and (ii) a unique public key and a corresponding unique private key associated with the certificate; receiving, from the application, a second request for digital content from the first domain; and receiving, from the application at a second time, an obfuscated identifier corresponding to the randomly selected identifier and a randomized group age bucket indicating an age range of a cookie containing an age of the randomized group, wherein the age of the randomized group is calculated based on the difference between the second time and the first time, and Wherein, within a predetermined time period of assigning the randomized group to the application, the randomly selected identifier is also assigned to at least a threshold number of other applications executed on other user devices.
8. The system according to claim 7, wherein: Assigning the randomized group to the application comprises: assigning, by the first domain, the randomization group to the application, and wherein the randomly selected identifier is assigned to the at least threshold number of other applications, wherein the randomly selected identifier is a randomly generated identifier selected from two or more randomly generated identifiers, and The unique public key is generated by the first domain.
9. The system according to claim 7, wherein: Assigning the randomized group to the application comprises: assigning the randomized group to the application by a central server; wherein the randomly selected identifier is assigned to the at least threshold number of other applications, wherein the randomly selected identifier is a randomly generated identifier selected from two or more randomly generated identifiers, and Wherein, the unique public key is generated by the central server.
10. The system according to claim 7, wherein: The operations also include: providing, by the application, a request for digital content from a second domain different from the first domain and the randomized group; receiving, by the application, a certification request including a challenge from the second domain in response to providing the request for digital content from the second domain; and providing the digitally signed certificate to a verification system by the application, the digitally signed certificate triggering the verification system to (i) create an obfuscated certificate containing the randomly selected identifier, the randomized group age bucket, and the challenge, (ii) sign the obfuscated certificate, and (iii) provide the obfuscated certificate to the second domain, Therein, the challenge is shielded from the verification system using a blinding scheme.
11. The system according to claim 7, wherein: The operations also include: The application provides the digitally signed certificate to a verification system; and The randomized group is verified, by the verification system, to be assigned to at least a threshold number of applications.
12. The system according to claim 7, wherein: The operations also include: The first domain detects anomalous activity associated with the randomly selected identifier based on the received randomized cohort age bucket and at least one of: a number of interactions associated with the randomly selected identifier, a randomized cohort age distribution, and a probability distribution associated with interactions and time periods.
13. A non-transitory computer storage medium encoded with instructions that, when executed by a distributed computing system, cause the distributed computing system to perform operations comprising: receiving, from an application on a user device of a user, a request for digital content from a first domain; assigning a randomized group to the application at a first time, the randomized group being constructed based on a randomly selected identifier and a timestamp indicating the first time the randomized group was assigned to the application; providing, at the first time, to the application (i) a digitally signed certificate corresponding to the randomly selected identifier and the timestamp, and (ii) a unique public key and a corresponding unique private key associated with the certificate; receiving, from the application, a second request for digital content from the first domain; as well as receiving, from the application at a second time, an obfuscated identifier corresponding to the randomly selected identifier and a randomized group age bucket indicating an age range of a cookie containing an age of the randomized group, wherein the age of the randomized group is calculated based on the difference between the second time and the first time, and Wherein, within a predetermined time period of assigning the randomized group to the application, the randomly selected identifier is also assigned to at least a threshold number of other applications executed on other user devices.
14. The non-transitory computer storage medium of claim 13, wherein: Assigning the randomized group to the application comprises: assigning, by the first domain, the randomization group to the application, and wherein the randomly selected identifier is assigned to the at least threshold number of other applications, wherein the randomly selected identifier is a randomly generated identifier selected from two or more randomly generated identifiers, and The unique public key is generated by the first domain.
15. The non-transitory computer storage medium of claim 13, wherein: Assigning the randomized group to the application comprises: assigning the randomized group to the application by a central server; wherein the randomly selected identifier is assigned to the at least threshold number of other applications, wherein the randomly selected identifier is a randomly generated identifier selected from two or more randomly generated identifiers, and Wherein, the unique public key is generated by the central server.
16. The non-transitory computer storage medium of claim 13, wherein: The operations also include: providing, by the application, a request for digital content from a second domain different from the first domain and the randomized group; receiving, by the application, a certification request including a challenge from the second domain in response to providing the request for digital content from the second domain; and providing the digitally signed certificate to a verification system by the application, the digitally signed certificate triggering the verification system to (i) create an obfuscated certificate containing the randomly selected identifier, the randomized group age bucket, and the challenge, (ii) sign the obfuscated certificate, and (iii) provide the obfuscated certificate to the second domain, Therein, the challenge is shielded from the verification system using a blinding scheme.
17. The non-transitory computer storage medium of claim 13, wherein: The operations also include: The application provides the digitally signed certificate to a verification system; and The randomized group is verified, by the verification system, to be assigned to at least a threshold number of applications.
18. The non-transitory computer storage medium of claim 13, wherein: The operations also include: The first domain detects anomalous activity associated with the randomly selected identifier based on the received randomized cohort age bucket and at least one of: a number of interactions associated with the randomly selected identifier, a randomized cohort age distribution, and a probability distribution associated with interactions and time periods.
Citation Information
Patent Citations
Cryptographic methods and systems for managing digital certificates with linkage values
US20190089547A1