Automated detection of unauthorized re-identification

By analyzing browser content through machine learning models, the system automatically identifies and blocks third-party tracking behavior, solving the problem of privacy leaks after users disable cookies and achieving effective privacy protection.

CN115865454BActive Publication Date: 2026-03-17GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-12
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing technologies cannot effectively identify and prevent third-party content providers from still tracking user behavior after users have disabled the cookie option, leading to user privacy leaks.

Method used

By using machine learning models (such as neural networks) to analyze user browser content, and through content characteristics and user history data, it can automatically identify whether third parties are tracking users, and block content providers from transmitting content when high-probability tracking behavior is identified.

Benefits of technology

Accurately identify and prevent third-party tracking behavior, protect user privacy, and prevent content providers from infringing on users' browser activity privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115865454B_ABST
    Figure CN115865454B_ABST
Patent Text Reader

Abstract

The present disclosure relates to automatically detecting unauthorized re-identification, including: retrieving a log of content items provided to anonymous computing devices; identifying a first content item provided to a plurality of anonymous computing devices within a first predetermined time period; for each anonymous computing device of the plurality of anonymous computing devices, generating a set of identifications of second content items retrieved by the anonymous computing device prior to receiving the first content item within a second predetermined time period; determining that a signal or combination of signals having a highest predictive power between a first set of identifications and a second set of identifications exceeds a threshold; identifying a provider of the first content item; and preventing transmission of a request for the content item by an anonymous computing device to the identified provider if the signal or combination of signals having the highest predictive power exceeds the threshold.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Case Analysis

[0002] This application is a divisional application of Chinese Invention Patent Application No. 202080008803.5, filed on May 12, 2020.

[0003] Cross-references to related applications

[0004] This application claims priority to U.S. Patent Application Serial No. 16 / 412,027, filed May 14, 2019, the contents of which are incorporated herein by reference. Background Technology

[0005] Nowadays, it's common for people to buy goods online rather than in physical stores. When people visit different web pages and domains to shop, information about them is often tracked by third parties using cookies. Cookies allow third parties to track which web pages a person visits, how often they visit each page, how long they spend on each page before viewing a new page, any choices they make on each page, the content of each page, and any data entries they execute on each page. Third parties are often content providers who use the data collected from cookies to select user-targeted content for each user based on the data collected through cookies. User-targeted content can be displayed to users when they view web pages that may not be related to the content shown to them.

[0006] To stop unwanted tracking and thus user-directed content selected by third-party content providers, websites and web browsers typically present users with options to enable or disable cookies. Unfortunately, even if a user chooses to disable cookies, third-party content providers can still ignore the user's choice and use cookies anyway, or bypass the user's choice to disable cookies using re-identification techniques (such as deterministic or probabilistic methods, email-based identity synchronization, phone-based identity synchronization, server-to-server synchronization, remote procedure calls, etc.). These re-identification techniques allow third-party content providers to track user activity after the user has opted to disable cookies and therefore believes they are not being tracked or receiving user-directed content.

[0007] Even after a user selects the option to disable cookies, a computer system or web browser may struggle to determine whether the user is being tracked by a third-party content provider. Regardless of whether the user is being tracked, the same content can be requested by the computer device and presented to the user. For example, a user can open a browser and select the option to disable cookies. The user can visit a webpage discussing cars and then visit another webpage where content describing cars appears in the sidebar of the website. The same content could be presented as a result of tracking the user from a website about cars, or it could be randomly selected and presented. Even though the user has selected the option to disable cookies, previous systems and methods could not determine how the content was selected (e.g., targeted or random selection) to identify whether a content provider is tracking the user. Although web browsers can currently block third-party content providers that track users, even after the option to not be tracked is selected, previous systems and methods could not identify content providers doing so. Summary of the Invention

[0008] The system and method discussed in this paper provide a machine learning model (e.g., neural networks, support vector machines, random forests, etc.) that can automatically determine the probability that a content provider is tracking a user, even if the user has selected an option to prevent them from being tracked otherwise. The input to the neural network can be common features of content snippets shown to the user on a webpage and other webpages, domains, and keywords previously viewed or entered into the browser by the user. Other inputs can be common webpages, domains, and keywords among users who viewed or provided input on different computers before receiving the same content snippet. Each of these inputs can be associated with different weights in the neural network to determine a binary classification representing the probability that a particular content provider is tracking the web user's online activity. The weights can be fine-tuned to improve the performance of the neural network, thereby accurately determining the probability that the content provider is tracking the user as the neural network receives more input and produces output. The system can obtain the probability determined by the neural network that the content provider is tracking the user and determine whether that probability is higher than a threshold selected by the administrator. If the probability is higher than the threshold, the system can prevent the content provider from serving content to the user. Any content requests from computer devices (sent to the content provider identified as tracking the user) can be redirected to other content providers.

[0009] Advantageously, by implementing the systems and methods discussed herein, the system can automatically identify third-party content providers who track web browser users after the user selects an option to prevent the provider from doing so. Systems that do not use the systems and methods discussed herein may rely on delegation, VPNs, user agent erasure, fingerprinting (e.g., unique identifiers in URL parameters), email matching and reduction, phone number matching and reduction, etc., to identify and prevent content providers from tracking users. Each of these techniques may be used depending on how the third party tracks users, but they are not effective against all tracking techniques. Furthermore, implementing these identification and prevention techniques is not always feasible because they may be difficult to scale, too expensive, or difficult to implement.

[0010] Fortunately, the system and method described in this paper can be used to accurately and automatically determine whether a third party is tracking a user, regardless of how the tracking is performed, and then prevent the third party from providing any further content to the web browser. The system can do this based on the characteristics of the content provided by the content provider and the characteristics of content previously viewed by the user. Using a neural network with binary classification output, the system can distinguish between content randomly selected and presented to the user and content presented as a targeted tracking product, even after the user has visited a website containing characteristics similar to the displayed content. Therefore, third-party content providers do not need to infringe on people's privacy by tracking their web browser activity, as they will be prevented from providing content to web users after being identified as tracking users who wish to hide their web browser activity.

[0011] In one aspect described herein, a method for detecting third-party re-identification of an anonymous computing device is provided. The method includes: retrieving a log of content items provided to the anonymous computing device by an analyzer of the computing system; identifying a first content item provided to a plurality of anonymous computing devices within a first predetermined time period by the analyzer; generating a set of identifiers for a second content item for each of the plurality of anonymous computing devices by the analyzer; determining that a signal or combination of signals with the highest predictability between the first and second identifier sets exceeds a threshold; identifying the provider of the first content item by the analyzer; and in response to the determination that a signal or combination of signals with the highest predictability between the first and second identifier sets exceeds the threshold, preventing the anonymous computing device's request for the content item from being transmitted to the identified provider.

[0012] In some implementations, the identifier of the second content item includes the identifier of the webpage accessed by each anonymous computing device. In some implementations, the identifier of the second content item includes the identifier of the domain accessed by each anonymous computing device. In some implementations, the identifier of the second content item includes the identifier of a keyword associated with the domain accessed by each anonymous computing device. In some implementations, the method further includes: the analyzer determining that the magnitude of the signal or signal combination with the highest predictability between the first identifier set and the second identifier set exceeds the magnitude of the signal or signal combination with the highest predictability between every other pair of identifier sets. In some implementations, the method further includes: the analyzer determining that the signal or signal combination with the highest predictability between the first identifier set and the second identifier set is common to a third identifier set. In some implementations, preventing the transmission request further includes: the computing device receiving a request for the content item from the anonymous computing device; and the computing system redirecting the request to a second provider.

[0013] In some implementations, the transmission prevention request also responds to: an analyzer identifying a third content item provided by the identified provider to a plurality of anonymous computing devices within a first predetermined time period; for each of the plurality of anonymous computing devices, before receiving the third content item within a second predetermined time period, the analyzer generating an identifier set of a fourth content item retrieved by the anonymous computing device; and the analyzer determining that a signal or combination of signals with the highest predictability between the first identifier set and the second identifier set of the fourth content item exceeds a threshold.

[0014] In some implementations, the method further includes incrementing a counter associated with the identified provider in response to determination that a signal or combination of signals with the highest predictability between the first set of identifiers and the second set of identifiers exceeds a threshold. In some implementations, preventing a transmission request also responds to the counter associated with the identified provider exceeding a second threshold.

[0015] In another aspect, a system for detecting third-party re-identification of anonymous computing devices is described. The system includes a computing system comprising a processor, a memory device, and a network interface, the processor executing an analyzer. The analyzer is configured to: retrieve a log of content items provided to the anonymous computing devices from the memory device; identify a first content item provided to a plurality of anonymous computing devices within a first predetermined time period; for each of the plurality of anonymous computing devices, before receiving the first content item within a second predetermined time period, generate an identifier set of second content items retrieved by the anonymous computing devices; determine that a signal or combination of signals with the highest predictability between the first identifier set and the second identifier set exceeds a threshold; and identify the provider of the first content item; and wherein the network interface is configured to prevent requests for content items from the anonymous computing devices from being transmitted to the identified provider in response to the determination that a signal or combination of signals with the highest predictability between the first identifier set and the second identifier set exceeds the threshold.

[0016] In some implementations, the identifier of the second content item includes the identifier of the webpage accessed by each anonymous computing device. In some implementations, the identifier of the second content item includes the identifier of the domain accessed by each anonymous computing device. In some implementations, the identifier of the second content item includes the identifier of a keyword associated with the domain accessed by each anonymous computing device. In some implementations, the analyzer is further configured to: determine that the magnitude of the signal or signal combination with the highest predictive power between the first identifier set and the second identifier set exceeds the magnitude of the signal or signal combination with the highest predictive power between every other pair of identifier sets.

[0017] In some implementations, the analyzer is further configured to: determine that the signal or combination of signals with the highest predictive power between the first and second identifier sets is common to the third identifier set. In some implementations, the network interface is further configured to: receive requests for content items from the anonymous computing device; and redirect the requests to a second provider.

[0018] In some implementations, the analyzer is further configured to: identify a third content item provided by the identified provider to a plurality of anonymous computing devices within a first predetermined time period; generate, for each of the plurality of anonymous computing devices, an identifier set of a fourth content item retrieved by the anonymous computing device before receiving the third content item within a second predetermined time period; and determine that a signal or combination of signals with the highest predictive power between the first identifier set and the second identifier set of the fourth content item exceeds a threshold.

[0019] In some implementations, the analyzer is further configured to increment a counter associated with the identified provider in response to determination that a signal or combination of signals with the highest predictive power between the first and second identifier sets exceeds a threshold. In some implementations, the network interface is further configured to prevent a transmission request in response to the counter associated with the identified provider exceeding a second threshold.

[0020] An optional feature of one aspect can be combined with any other aspect. Attached Figure Description

[0021] Details of one or more implementations are set forth in the following drawings and description. Other features, aspects, and advantages of this disclosure will become apparent from the description, drawings, and claims, wherein:

[0022] Figure 1A It is a block diagram of two sequences based on some implementation methods, each sequence including the user viewing a first webpage and a second webpage, and the content being provided on the second webpage;

[0023] Figure 1B It is a block diagram of the implementation of a system based on some implementation methods, which is used to determine whether a third party is tracking the activities of multiple users;

[0024] Figure 2 It is a block diagram of a machine learning model based on some implementation methods, which has inputs from the history of content viewed by the user and outputs indicating whether a third party is tracking the user's activity;

[0025] Figure 3 The flowchart illustrates a method for determining whether a third party is tracking user activity based on the output from a neural network, based on some implementation details.

[0026] Figure 4 The diagram illustrates a flowchart of another method based on some implementations, which is used to determine whether a third party is tracking user activity based on the output from a neural network.

[0027] In the various figures, the same reference numerals and names indicate the same elements. Detailed Implementation

[0028] Even after a user selects the option to disable cookies in their browser or enable LAT on their mobile device, a computer system or web browser may struggle to determine whether the user is being tracked by a third-party content provider. For example, a user might open a browser and select the option to disable cookies. The user might visit a webpage discussing cars and then visit another webpage unrelated to cars, where car-related content appears. The same content could be presented as a result of tracking the user from a website about cars, or it could be randomly selected and contextually chosen and presented by the server. Even with the user's option to disable cookies, previous systems and methods could not determine how content was selected (e.g., targeted, random, or contextual selection) to identify whether a content provider was tracking the user. While web browsers can currently block content providers that track users, previous systems and methods could not identify those doing so even after the option to not be tracked was selected. Therefore, a method is needed to automatically identify and prevent content providers from tracking users against their will.

[0029] For example, in some implementations, the first reference is... Figure 1A The diagram illustrates two sequences 102 and 116, each including a device retrieving and displaying a first webpage, then retrieving and displaying a second webpage, where content is provided. In some implementations, sequence 102 may be a sequence where a first user device 104 retrieves and displays website 106, then retrieves and displays another website 112. Content server 108 may monitor the different webpages accessed by the device. In some implementations, sequence 116 may be a sequence where another device 118 retrieves and displays website 120, then retrieves and displays another website 126. Content server 122 may provide random or context-oriented content to user device 118.

[0030] In some implementations, at sequence 102, user device 104 retrieves and displays website 106. This website may be related to purchasing different sports, hobbies, pets, or any other such content, or it may be unrelated to shopping. The website may also have features related to the website content, which are stored on the user device as cookies, as small files or website files. Cookies are typically first-party cookies generated and stored by the domain of the website being visited, but cookies can also be third-party cookies, which are cookies stored by a domain different from the domain accessed by the device. Third-party cookies are typically stored on the user device by the content provider to track web activity performed on a specific user device. Content providers can use cookies to select and present user-targeted content based on the tracked web activity of the device or the user associated with the device. In sequence 102, a third-party cookie may be stored on user device 104 to track that the device has visited website 106 discussing, for example, cars. In some implementations, when a website or web browser is opened, the user of the device may be presented with an option to enable or disable cookies. This option may apply to third-party and / or first-party cookies. If a user chooses to disable cookies, it likely means that the user does not want to be tracked and wants to preserve their privacy.

[0031] Content server 108 may be a server or processor configured to serve content to the user at a dedicated content space on the website after receiving a request from the user's device. Content server 108 may include behavior monitor 110. In some implementations, behavior monitor 110 may include an application, server, service, daemon, routine, or other executable logic to monitor the user's web browsing and / or search behavior across different user devices. In some implementations, behavior monitor 110 may monitor user behavior using deterministic or probabilistic methods, email-based identity synchronization, telephone-based identity synchronization, server-to-server synchronization, remote procedure call applications, IP address monitoring, etc. In some implementations, behavior monitor 110 may monitor user behavior when the user is using a privacy-preserving browser such as Google Chrome's incognito mode, for example, based on tracking requests from the device's IP address or similar data even though the device does not retain cookies. Behavior monitor 110 may track user web browsing behavior using methods other than third-party cookies, which users typically disable actively by choosing to disable or can be automatically disabled by the web browser unless the user chooses to enable third-party cookies.

[0032] Behavior monitor 110 can identify that a user has viewed website 106, such as example.com, and identify the website's content (e.g., by retrieving a copy of the website, identifying keywords associated with the website or domain, etc.). Therefore, behavior monitor 110 can provide content related to example.com on another website 112 (such as website.com). For example, if example.com is related to cars, the behavior monitor can determine that example.com is related to cars and provide additional car-related content when the user views another website (such as website.com). Additional content 114 can be an example of content provided by a content provider in response to identifying that a user has viewed a relevant website through behavior monitor 110 and is provided on another website as a result of this identification.

[0033] Sequence 116 may be similar to Sequence 102, but additional content is selected and served to the user randomly or contextually rather than as a result of any behavioral monitoring. In Sequence 116, device 118 may retrieve and display website 120, such as example.com, via user device 118. The user of the device may have opted out of disallowing third-party cookies. The device may subsequently retrieve and display another website 126, such as website.com. Additional content 128 may be served by content server 122 at website.com and may be randomly or contextually selected, rather than by tracking past browsing or search behavior.

[0034] Content server 122 may be a server or processor configured to provide additional content from a content provider to a website and / or user device upon receiving a request from a user device. Content server 122 may be similar to or the same as content server 108. Content server 122 is shown as including context selector 124. In some implementations, context selector 124 may include an application, server, service, daemon, routine, or other executable logic for contextually selecting additional content to be provided to a device accessing the website or other primary content. In some implementations, context selector 124 may use random number generation or pseudo-random number generation techniques to select content. In some implementations, context selector 124 may select content based on the context of the webpage on which the content will be provided. In the example shown, content server 122 may provide the same car-related content as content server 108 of sequence 102. After a user visits the same website as the user of sequence 102, content server 122 provides car-related content whose web activity is tracked by content server 108.

[0035] As seen in sequences 102 and 116, users can access the same website and be presented with the same content while viewing another website, regardless of whether they are being tracked. Sequence 102 involves a third party tracking the user, even though the user has opted out of options to disable third-party cookies and not be tracked. The third party selects and serves content based on the tracking. At sequence 116, the user visits the same webpage before receiving the same content from the content provider as in sequence 102, but at sequence 116, the content is selected and served contextually, rather than as a result of any tracking. Therefore, users and systems lacking the implementation discussed herein may find it difficult to identify when a third party is tracking them based on the content shown to them. Users opt out of third-party cookies, thus preventing third parties from tracking their web activity and protecting their privacy. For users and web browsers, it is difficult to determine when they are being tracked based solely on the content provided. Therefore, there is a need to automatically determine when a user is being tracked, so that a processor relaying content for the content provider can prevent third parties from tracking users and protect user privacy.

[0036] Fortunately, the system and method described in this paper can be used to accurately and automatically determine whether a third party is tracking a user, regardless of how the tracking is performed, and then prevent the third party from providing any further content to the user's device. The system can do this based on the characteristics of the content provided by the content provider and the characteristics of content previously viewed by the user. Using a neural network with binary classification output, the system can distinguish between content that is contextually selected and presented to the user and content presented as a targeted tracking product, even after the user has visited a website containing characteristics similar to the displayed content. Therefore, after being identified as tracking a user who wishes to hide their web browser activity, third-party content providers may be prevented from providing content to web users, thereby improving user privacy.

[0037] For example, refer to Figure 1BAccording to some implementations, an implementation of a system 134 for determining whether a third party is tracking the activities of multiple users is shown. In some implementations, system 134 is shown as including content providers 136 and 164, user devices 140, 142, and 144, and a re-identification server 146. Each of the content providers 136 and 164, user devices 140, 142, and 144, and the re-identification server 146 can communicate with each other and with other devices via network 139. Network 139 may include a synchronous or asynchronous network. Content providers 136 and 164 can provide content to user devices 140, 142, and 144 after receiving a request from one of the user devices. When providing content, using instructions stored in the re-identification server 146, the re-identification server 146 can determine whether content provider 136 or 164 is tracking the web activity of users at user devices 140, 142, and 144. The re-identification server 146 can do this using a neural network (or any other machine learning model) that automatically determines the probability that content provider 136 or 164 is tracking a user. The re-identification server 146 can determine if this probability is higher than a predetermined threshold, and if so, prevent content provider 136 or 164 from providing content to user devices 140, 142, and 144 by redirecting requests for content to other content providers or by blocking any content provided by content provider 136 or 164 from being transmitted to user devices 140, 142, and 144.

[0038] User devices 140, 142, and 144 (generally referred to as (multiple) user devices) may include any type and form of media device or computing device, including desktop computers, laptop computers, portable computers, tablet computers, wearable computers, embedded computers, smart TVs, set-top boxes, consoles, Internet of Things (IoT) devices or smart devices, or any other type and form of computing device. (Multiple) client devices may be referred to differently as clients, devices, client devices, computing devices, user devices, anonymous computing devices, or any other such term. Client devices and mediators may receive media streams via any suitable network, including local area networks (LANs), wide area networks (WANs) such as the Internet, satellite networks, cable networks, broadband networks, fiber optic networks, microwave networks, cellular networks, wireless networks, or any combination of these or other such networks. In many implementations, the network may include multiple subnets, which may be the same or different types, and may include multiple additional devices (not shown), including gateways, modems, firewalls, routers, switches, etc.

[0039] In some implementations, each operation performed by the re-identification server 146 can be performed by user devices 140, 142, and 144. User devices 140, 142, and 144 may include machine learning models (e.g., neural networks, random forests, support vector machines, etc.) that can determine the probability that a content provider will provide user-oriented content based on signal inputs from the machine learning model. The machine learning model can be implemented on the browsers of user devices 140, 142, and 144. Examples of inputs that user devices 140, 142, and 144 can use to determine whether the content received by a user is user-oriented content or whether the user is being tracked may include, but are not limited to, content viewed by the viewer before receiving the content, characteristics of previously viewed content, content, characteristics of the content, the webpage on which the content will be provided, characteristics of the webpage, etc. If the probability is higher than a threshold, user devices 140, 142, and 144 can prevent content providers that provide user-oriented content from providing content to the user devices in the future.

[0040] For example, a user device might access a webpage dedicated to purchasing shoes. When the user decides which shoes to buy, the device can retrieve multiple webpages displaying the shoes. The user can then stop shopping and navigate to a pets-related webpage. While viewing the pets webpage, the user might receive content similar to the shoes shown on previously viewed pages. The device's browser can implement a machine learning model to determine whether the shoe content is user-targeted content. The browser can use previously viewed webpages and domains, along with characteristics of those webpages, as input, along with the received content and the webpage on which it is provided. The machine learning model can receive the input and determine the probability that the content is being tracked and / or that the content provider is tracking the user device. The browser can compare this probability to a predetermined threshold. If the probability is greater than the threshold, the browser can determine that the content is user-targeted and / or that the content provider is tracking the user. The browser may then block the receiving of content from the content provider. Otherwise, if the probability is less than the threshold, the browser can determine the content to be provided based on the context of the webpage, rather than as a result of tracking the user.

[0041] In some implementations, user devices 140, 142, and 144 can retrieve logs from other user devices to determine the intersection and / or signal or combination of signals that user devices 140, 142, and 144 viewed before receiving the content, and that had the highest predictive power. If user devices 140, 142, and 144 viewed similar content before receiving the same content, then a strong signal can be associated with the commonly viewed content when that similar content is used as input in a machine learning model.

[0042] In some implementations, content provider 136 may be a third-party content provider that can track web activity performed at user devices 140, 142, and 144 and provide content related to that web activity to users at user devices 140, 142, and 144. Content provider 136 can still track web activity even if users implement tracking prevention technologies (e.g., disabling third-party cookies by website / domain or through a web browser, using proxy or VPN, user agent erasure, identifying a unique ID for content provider 136 and rejecting data associated with that ID, reducing emails from email addresses associated with content provider 136, reducing phone calls from numbers associated with content provider 136, etc.). Content provider 136 may use the data collected from the tracked web activity to continue providing user-targeted content related to that tracked web activity.

[0043] For example, a user can open a web browser on user device 140 and be immediately presented with the option to enable or disable third-party cookies. The user can choose to disable third-party cookies because they want to protect their privacy and do not want to be tracked while browsing the internet. Depending on whether the website or web browser presents the user with the option to block third-party cookies, the website or web browser can then block any third-party cookies that are to be installed on the user's device. However, content provider 136 can still use various techniques (such as deterministic or probabilistic methods, email-based identity synchronization, telephone-based identity synchronization, server-to-server synchronization, remote procedure call applications, etc.) to track activity on user device 140. Content provider 136 can use the tracked activity to identify characteristics of the user's associated activity and select content 138 associated with the identified characteristics at a location dedicated to receiving and presenting content from the content provider.

[0044] For example, a user at user device 140 can be tracked by content provider 136 using one of the techniques listed above. A user can access shoe-related web pages while viewing different web pages. Content provider 136 can identify that a user has accessed shoe-related web pages and identify shoe-related content 138. When a user visits websites and domains related to other topics (e.g., sports), user device 140 can send a request to content provider 136 to have content displayed in a dedicated content space, and based on the user's access to shoe-related web pages, content provider 136 can provide shoe-related content. In some implementations, the request to the content provider is first transmitted to re-identification server 146 before being transmitted to content provider 138.

[0045] In some implementations, the re-identification server 146 may include one or more servers or processors configured to determine whether content providers 136 or 138 are tracking the web activity of users at user devices 140, 142, and 144 and / or whether the content provided by content providers 136 or 138 is user-oriented content. In some implementations, the re-identification server 146 is shown as including processor 148 and memory 150. The re-identification server 146 may be configured to identify whether users have granted third-party consent to identify them and track their web activity, identify content providers that may be tracking users and providing content to users, implement a neural network with different inputs associated with users and their past web activity to determine whether content providers are tracking users and / or whether the content provided by content providers is user-oriented content, count the frequency at which content providers have been identified as tracking users or providing user-oriented content, and transmit requests for additional content from user devices 140, 142, and 144 to content providers based on whether content providers have been identified as tracking users or providing user-oriented content. Re-identifying one or more components within server 146 can facilitate communication between each component within server 146 and external components such as content providers 136 and 164 and user devices 140, 142, and 144. Re-identifying server 146 may include multiple connectivity devices, each providing some of the necessary operations (e.g., as a server group, a set of blade servers, or a multiprocessor system).

[0046] In some implementations, processor 148 may include one or more processors configured to execute instructions on modules stored in memory 150 within reidentification server 146. In some implementations, processor 148 executes an analyzer (not shown) to execute modules within memory 150, which may be configured to determine whether a third party is tracking the web activity of a user using the Internet. To this end, the analyzer may execute instructions stored in memory 150 of reidentification server 146. In some implementations, memory 150 is shown as including an authorizer 152, a content identifier 154, a provider identifier 156, a counter 158, an application 160, and a transmitter 162. By executing the analyzer to perform the operation of each component 152, 154, 156, 158, 160, and 162, processor 148 may automatically determine the probability that content provider 136 (or any other content provider) is serving user-oriented content to user devices 140, 142, and 144 on the Internet after the user has explicitly selected the option not to be tracked by a third party. Processor 148 can implement a neural network using application 160 to determine the probability that content provider 136 is using unauthorized re-identification techniques to serve user-oriented content. Processor 148 can compare the probability to a threshold to determine if the probability is high enough to limit the ability of content provider 136 to serve content to user devices over the network. In some implementations, processor 148 may include a coprocessor, or may communicate with a coprocessor such as a Tensor Processing Unit (TPU), which is dedicated solely to using machine learning techniques to determine the probability that content provider 136 is serving user-oriented content or that the content is being served due to user tracking.

[0047] In some implementations, the authorizer 152 may include an application, server, service, daemon, routine, or other executable logic to identify a user who has selected the option to disable third-party identification / tracking. In some implementations, the authorizer 152 provides the user with an option (e.g., via a user interface, a provided webpage, etc.) to enable or disable third-party tracking. This option is implemented using cookies and is used to select user-oriented content for each user. The user can select either option. The authorizer 152 may also provide the user with the option to enable or disable first-party cookies, which allows websites and / or domains to store user-related and website-specific data. For example, a shopping website may include a virtual shopping cart containing items that a shopper wishes to purchase. First-party cookies can allow items to remain in the shopping cart as the shopper continues shopping, rather than disappearing when the shopper leaves the page associated with that cart. First-party cookies can also be used to store the user's username and password on the user's device, so the user does not have to repeatedly enter their username and password when re-entering the website. Unfortunately, if first-party cookies are enabled, third parties can use them to track the user's web activity via server-to-server synchronization techniques.

[0048] Upon receiving an option to disable user tracking and thus disable third-party cookies, authorizer 152 can automatically block third-party cookies from being stored on the associated user device. Authorizer 152 can block third-party cookies specific to the domain the user is using when the option to disable tracking is presented, or those spanning all domains via a web browser. In some implementations, blocking third-party cookies can be a single domain or a default setting of the web browser. Therefore, in these implementations, third-party cookies can be automatically blocked by authorizer 152 unless the user manually changes the settings to enable third-party tracking. Although shown on a server, in many implementations, user authorization can be controlled by an authorizer performed by the user's computing device, such as a non-tracking flag controlled by the browser or other applications performed by the user's computing device.

[0049] In addition to presenting the user with the option to enable or disable third-party tracking, the authorizer 152 can store an instruction associated with each computing device, indicating that the user of the user device does not wish to be tracked. The authorizer 152 can receive an instruction from user devices 140, 142, or 144 indicating that the user has selected the option to disable third-party tracking, or that the user device can automatically block third-party tracking via cookies through a web browser. Upon receiving the instruction, the authorizer 152 can indicate to the content identifier 154 that the user does not wish to be tracked by third parties.

[0050] In some implementations, content identifier 154 may include an application, server, service, daemon, routine, or other executable logic to track the web activity of users who have been determined by licensor 152 to have opted out of third-party tracking, and to determine the content of the web activity. Content identifier 154 may receive instructions from licensor 152, either automatically via a web browser or by selecting an option to disable tracking, indicating users who have opted out of third-party tracking. When a user browses the Internet via a web browser, content identifier 154 may generate logs of websites, domains, and keyword searches accessed and / or entered into the browser, and store these logs in a database (not shown) within reidentification server 146.

[0051] When generating data logs associated with each user and storing these logs in a database, content identifier 154 can identify the content, domain, and keywords of each webpage accessed or entered into the corresponding user device by each user. Content identifier 154 can identify the content of a domain or webpage by comparing the domain with a table within the database of re-identification server 146. This table can include content information related to different domains or webpages associated with the Internet. For example, a domain might be associated with pet training. The table would have a domain name (or webpage URL) in one column and a content descriptor in another column indicating that the domain is for pet training. If a user accesses a domain associated with pet training, content identifier 154 can identify the domain from the associated URL and determine that the domain's content is related to pet training by finding the domain and associated content in a table containing relevant information. In some implementations, instead of using the domain's URL to determine content, content identifier 154 can use the most common content of the webpages associated with the domain to determine the content associated with the domain. Content identifier 154 can use any technology or method to determine the content of a domain.

[0052] Content identifier 154 can identify the content of different web pages associated with a domain visited by a user. To this end, content identifier 154 can identify common terms and / or terminology patterns appearing on each web page. For example, if a user is viewing a web page describing shoes, content identifier 154 can identify shoe-related terms such as high heels, different shoe brands, different shoe types, etc. Content identifier 154 can identify shoe-related words and determine whether a web page is associated with shoes if a threshold for the number of shoe-related words needs to be met.

[0053] In another implementation, content identifier 154 can identify content associated with a webpage, including video, audio, text, etc., by analyzing images or other content embedded in the webpage. For example, content identifier 154 can scan the webpage for media or other content items, identify any media on the webpage, and then determine the content of the webpage based on the content of the media. In some implementations, content identifier 154 can use object recognition technology to determine the content of images or pictures within a video. For example, content identifier 154 can identify the characteristics of images in the content and compare these characteristics with images already labeled with tags indicating the content associated with the images. If content identifier 154 can identify sufficient characteristics of images that are the same as or similar to the labeled images, then content identifier 154 can determine the content of the image and thus the content of the webpage. For example, if the webpage includes an image of shoes, content identifier 154 can identify the characteristics of shoes and compare the characteristics of shoe images in the database within the re-identification server 146. If content identifier 154 identifies sufficiently common characteristics (i.e., sufficient characteristics that meet a threshold), then content identifier 154 can determine that the content associated with the image is shoes, and the webpage is associated with shoes. Content identifier 154 can use any technology or method to identify content associated with a webpage.

[0054] To identify the content of keywords or terms used in searches or entered into a user's device, content identifier 154 can identify words in the search and compare them with words in a database (not shown) within re-identification server 146. The database may include terms labeled with content types, which indicate what content the words are associated with, similar to how domains are labeled with content types. Content identifier 154 can compare keywords with words in the database and determine the content associated with each keyword based on the content tags of matching words in the database. Content identifier 154 can use any technology or method to identify the content type associated with a keyword.

[0055] When identifying content associated with each keyword, webpage, and / or domain, the re-identification server 146 can tag each keyword, webpage, and / or domain with labels indicating what type of content it is associated with. For example, if a user performs a search in the Google search engine using the term "dog" and visits a domain associated with dogs, the content identifier 154 can automatically determine that the content associated with the keyword "dog" and the domain is dogs. The content identifier 154 can then tag the keyword and domain with the dog-associated label, add the keyword and domain to a log of keywords, webpages, and domains visited by the user within a set time period, and store the log in an association database within the re-identification server 146. The content of the keywords, webpages, and domains can be characteristics of the corresponding keywords, webpages, and domains and is used as input to the neural network described below.

[0056] In some implementations, content provided by content provider 136 can also be included in the content log and labeled with tags indicating what type of content it is associated with. For example, the content provider may send content related to the sale of cars. Content identifier 154 can identify the content, label it with tags indicating that the content is associated with cars, and add the content to the content log.

[0057] Content identifier 154 can generate, update, and store logs of data associated with a user's web activity at user devices 140, 142, or 144 for any length of time. In some implementations, content identifier 154 can generate and store a rolling window period that includes the most recent data from the period immediately preceding the current period. For example, content identifier 154 can store data in a log associated with the user's web activity for the previous 30 days. In some implementations, data from periods earlier than 30 days can be removed as each day passes, and new data from the current day can be added to the log. Therefore, content identifier 154 can track a user's current interests at user devices 140, 142, and 144, and the data can be unaffected by searches, web pages, and domains from periods the user does not intend to access. The rolling window period can be of any duration.

[0058] In some implementations, logs can be stored on user devices 140, 142, and 144. Logs can be provided to a re-identification server 146 upon request to determine whether the content is user-directed content and / or whether the content provider is tracking the user device receiving the content. Furthermore, if the methods described herein are performed on user devices, the logs can be used without transmitting them to another user device or server. Advantageously, by storing logs on user devices 140, 142, and 144, each user device can store private data without sharing data with the server. Users of user devices 140, 142, and 144 may not wish to share log data with other user devices or servers.

[0059] Still refer to Figure 1B Server 146 is re-identified as including application 160. Application 160 may include applications, servers, services, daemons, routines, or other executable logic to implement the following references. Figure 2 The neural network shown and described is used to determine whether content providers such as content providers 136 and 164 are serving user-targeted content after a user selects the option not to be tracked. Application 160 may receive input identified by content identifier 154, which is associated with the recent web activity of multiple users and the content presented to multiple users. Using the neural network, application 160 can use weights associated with the input to determine an output probability indicating the likelihood that a content provider is serving user-targeted content. In some implementations, the neural network may be a binary classifier that provides two outputs: one indicating that a content provider is serving user-targeted content, and the other indicating that the content provider is serving context-targeted content based on input associated with the web activity of different users. There can be any number of inputs, and each input can be associated with any weights. In other implementations, application 160 implements statistical analysis or linear regression models, random forests, gradient boosting decision trees, etc., instead of a neural network, to determine the probability that a content provider is tracking a user or that specific content is being targeted.

[0060] Application 160 can implement a log of content items generated or selected by content identifier 154 as one input to its neural network. This log includes content viewed by the user before receiving content from content provider 136. In some implementations, application 160 can specify a time period for receiving data from the log, such as, for example, 30 days before the user viewed the content. Application 160 can identify viewed content from any time period. By identifying tags associated with the content in the log of the content item, application 160 can identify the content and the content type associated with that content. As mentioned above, content and content type can be associated with keywords, web pages, domains, etc. Once a content item is identified, application 160 can feed the content item into its neural network to determine the probability that the content provider has provided user-targeted content to the user device without authorization. In some implementations, the content item can be compared to provided content items, with any content item similar to the provided content item having a stronger weight. As described herein, a stronger weight can also be referred to as a heavier or higher weight.

[0061] In some implementations, the input from the content item log may include the intersection and / or signals or combinations of signals between a first set of identifiers and a second set of identifiers from the content item log. The content of the intersection and / or signals or combinations of signals with the highest predictive power can be input into the neural network. The identifier set can be associated with the web activity of a specific user. For example, within 15 days prior to user A viewing a specific content snippet provided by a content provider, content identifier 154 can track content viewed or entered into the user's device by a specific user (user A). Each activity of user A (e.g., each webpage or domain visited or keyword entered into the user's device) can be generated as the first identifier set by application 160. Identifier sets can be generated for any number of users who view the same content provided by the content provider within any time period. Therefore, if a second user also views the same content provided by the content provider, application 160 can identify content items from the second user's web activity as the second identifier set. Application 160 can identify the intersection and / or signal or combination of signals with the highest predictive power between the first and second identifier sets as common content, which is viewed by the first and second users within a predetermined time period before viewing content provided by the content provider. The predetermined time period can be of any length.

[0062] For example, if users A and B are both presented with the same content associated with content provider 136, application 160 can identify all content viewed before users A and B were presented with the content. User A may have visited web pages C, D, E, and F, and user B may have visited web pages D, E, F, and G. Application 160 can determine the intersection and / or signal or combination of signals with the highest predictive power of the viewed content as web pages D, E, and F for both users. Application 160 can input web pages D, E, and F into a neural network as input associated with the intersection between users A and B. The neural network can use the input and the weights associated with the input to determine the probability that the content provider is tracking the web activity of users A and B or that a particular content segment is being targeted. Application 160 can obtain the probability and compare it to a predetermined threshold (such as 80%). If the probability exceeds a threshold, application 160 can determine that the content provider is tracking the user (or providing user-targeted content) and send a signal to transmitter 162, instructing transmitter 162 to stop providing content from the identified content provider to the user device that sent the content request. In other implementations, application 160 can use the identifiers of web pages C, D, E, F, and G as input, and the neural network can give greater weight to web pages shared by each user (e.g., D, E, and F) and less weight to web pages viewed by only one user (e.g., C and G). Therefore, the intersections and / or signals or combinations of signals with the highest predictive power may not need to be explicitly identified, but can be implicitly incorporated into the machine learning model once trained.

[0063] Application 160 can obtain data associated with the signal or combination of signals that has the highest predictive power. The signal or combination of signals with the highest predictive power can be a signal or combination of signals associated with the highest weight or score when used to determine whether content provided to a user is user-oriented and / or whether a content provider is tracking a user. In some embodiments, predictive power can be based on the number of users who viewed the same content before receiving it from the content provider. For example, if a large number of users visited the same webpage or domain before receiving the same content, these signals can have the highest predictive power. The larger the number, the higher the predictive power. In some embodiments, predictive power can be based on the similarity of characteristics between the received content and the content viewed before receiving the content. In some embodiments, the signal or combination of signals can include the intersection of content viewed by multiple users before receiving the same content, or the intersection of such content. The intersection can match signal inputs (e.g., signal inputs associated with the same content and / or content characteristics) to a trained machine learning model.

[0064] The highest predictive power of a signal or combination of signals can be positive or negative. In some embodiments, the highest predictive power can be positive if the signal or combination of signals indicates a high probability that the content is user-oriented or that the content provider is tracking the user. For example, content that has been viewed by multiple users before a user receives content from a content provider can be associated with a positive predictive power. In some embodiments, the highest predictive power can be negative if the signal or combination of signals indicates a high probability that the content is not user-oriented. For example, the signal or combination of signals can be associated with a webpage with a negative predictive power because a user who visits the webpage receives context-oriented content after visiting that webpage.

[0065] Application 160 can obtain data associated with the intersection and / or signal or combination of signals that has the highest predictive power for any number of users. For example, continuing the example above, if a third user (user C) views the same content fragments as users A and B, and application 160 determines via content identifier 154 that user C viewed web pages D, E, and G, then application 160 can determine the intersection and / or signal or combination of signals that has the highest predictive power for web pages D and E among users A, B, and C. The intersection and / or signal or combination of signals that has the highest predictive power among users A, B, and C can be fed into the neural network along with the intersection and / or signal or combination of signals that has the highest predictive power associated with A and B. In some implementations, when the neural network automatically determines the probability that a content provider is tracking the user or content fragment being targeted, the intersection and / or signal or combination of signals that has the highest predictive power among users A, B, and C can be associated with stronger weights or stronger signals in the neural network.

[0066] The intersection and / or signal or combination of signals with the highest predictive power can be determined for any number of viewers viewing the same content. The more common the sites viewed among more people, the greater the weight of common sites, and the higher the chance that the neural network determines that the content provider is tracking users or content segments being targeted. However, if there is no strong correlation between previous sites viewed by viewers of the same content, the probability that the neural network can determine that the content provider is tracking users or content segments being targeted is low.

[0067] For example, content can be provided to a large number of users. Among users viewing the content, 50% may have visited website A, 30% may have visited website B, and 20% may have visited website C. Based on the percentage of users visiting each site, the neural network weights the signals or combinations of signals associated with websites A, B, and C. Website A might be associated with the signal or combination of signals with the highest weight, followed by website B, and further down the list, website C. In some implementations, the weights can be directly related to the percentage of users visiting the website, although in other implementations, the weights may be unrelated to the percentage of users visiting the corresponding website (e.g., if a website is consistently associated with tracking users or providing user-targeted content, then its weight might be larger despite a smaller percentage of users visiting the site). Continuing the example above, the signal or combination of signals within the neural network associated with website A might be 2.5 times greater than the signal or combination of signals associated with website C, since 50% of users visited website A and 20% visited website C.

[0068] In some implementations, the weights associated with the input can also be based on when a user viewed content related to content provided by a content provider. The closer in time a user's viewing of content intersects with content viewed by other users to the content from the content provider, the stronger the weight or signal associated with the intersection and / or signal or combination of signals that has the highest predictive power. For example, if users A and B both viewed website C 30 minutes before the content was provided by the content provider, the neural network can give greater weight to the input associated with website C compared to websites visited by users A and B a week before the content was provided. Users A and B may have viewed website C at different times than the content was provided by the content provider; however, the application 160 considers these differences by taking the average of the times, summing the times, or any other method that normalizes the time difference between the two users. For example, the weights of the neural network input can be related to the time between viewing content and the website or to the order of accessing common websites, where the weights of the input are only related to the order in which users A and B accessed the websites.

[0069] In some implementations, features of the intersection and / or signal or combination of signals with the highest predictive power, based on similarity to content provided by the content provider, can be associated with different weights and used as inputs in a neural network. Features of intersection websites, keywords, or domains can be inputs, each with a unique weight associated with it. Examples of features include, but are not limited to, videos, images, text, colors, etc. Features can also include content types. For example, features can include different topics the content focuses on, such as animals, cars, shoes, sports, education, schools, etc. In some implementations, the weight of a feature can depend on the weights mentioned above, where features of the most frequently accessed content have a stronger weight than features of content that was not frequently accessed before the user was presented with content from the content provider. Furthermore, the more similar a feature is to the features of the provided content, the stronger the weight associated with the signal or combination of signals of that feature. Features of viewed content that are closer in time to the content from the content provider can also be associated with a stronger signal. The temporal relationship between the content and the provided content can also be inputs to the neural network.

[0070] In some implementations, Application 160 can identify web pages whose content is viewed by different users as inputs to the neural network. Since web page owners typically provide content from content providers associated with the web page, including content from unrelated content providers indicates a high probability that the provided content is being offered due to tracking by a third-party content provider. The neural network can identify the similarity between a web page and the provided content. The lower the similarity, the higher the weight associated with the web page. For example, if a user is viewing a web page describing cars, and content related to shoes is provided to the web page by a content provider, the neural network can associate the web page with a higher weight compared to a scenario where the content provided by the content provider is also related to cars. Furthermore, the characteristics of the web page can be inputs to the neural network, similar to the intersection and / or signal or combination of signal characteristics described above that have the highest predictive power.

[0071] In some implementations, another input to the neural network can be content provided by a content provider. Application 160 can identify the content type as the input to the neural network. This is beneficial because some types of content may be more likely to be provided by content providers tracking users compared to other types of content. For example, through an associated neural network, Application 160 can determine that content from a content provider related to shoes is more likely to be associated with unauthorized tracking of a user compared to content about sports. Therefore, in the neural network, the weights associated with content inputs about shoes can be higher than the weights associated with content inputs related to sports. Different types of content can have any weights.

[0072] As will refer to Figure 2 In more detail, by implementing a neural network that can have any number of hidden layers, the network can consider different weights based on how many users accessed or viewed the same content before it was served by a content provider. The weights associated with different inputs and hidden layers can be adjusted, so the time element of viewing content before receiving it from the content provider, and the number of users who viewed the same content before it was presented from the content provider, can be appropriately weighted. Using various training techniques for training neural networks, such as using reference material and backpropagation after determining the probability that a content provider is tracking users or that a content segment is user-targeted content, the neural network can automatically learn appropriate weights for different inputs, thus creating algorithms that can accurately determine the probability that an unauthorized content provider is tracking web browser users or serving user-targeted content to web browser users.

[0073] After determining the probability that a content provider is tracking a user or that content is being delivered to a user in a targeted manner (explicitly removing any potential consent from the content provider) and determining that the probability is above a predetermined threshold, application 160 may signal to sender 162, instructing sender 162 to prevent the content provider from receiving any future transmissions of requests from user devices to the content provider. Sender 162 may include an application, server, service, daemon, routine, or other executable logic to transmit content requests received from user devices 140, 142, and 144 and sent to the content provider. In some implementations, the content provider may send content back to user devices 140, 142, and 144 via sender 162. In other implementations, content provider 136 sends content directly to user devices 140, 142, and 144. Sender 162 may prevent the transmission of content requests from user devices 140, 142, and 144 to the content provider.

[0074] In some implementations, such as where the server acts as an intermediary between clients requesting additional content from a content provider, sender 162 can prevent the transmission of requests from user devices 140, 142, and 144 to the content provider by re-identifying server 146 as having provided unauthorized targeted content to user devices 140, 142, and 144. For example, when a user device is browsing the Internet, application 160 can determine that content provider 136 is providing targeted content to one of user devices 140, 142, and 144. Application 160 can signal to sender 162 that content provider 136 has provided user-targeted content to the user device, and sender 162 can redirect any future requests from the user device to content provider 164. In some implementations, sender 162 can redirect requests from all user devices, while in other implementations, sender 162 can redirect requests from a user device to which content provider 136 has been identified as providing user-targeted content.

[0075] In some implementations, before redirecting content requests from content provider 136, sender 162 may send a notification to the tracked user device indicating that the user device is being tracked. Sender 162 may also signal to the user device that the content received by the user device is user-directed content. The user at the tracked user device may be presented with the option to continue allowing content provider 136 to provide content to the user device or to block future content from content provider 136. For example, sender 162 may send the following message to the user device: “We believe content provider 136 is trying to track you against your consent. Would you like to block their cookies? Would you like to use our proxy / VPN service to help keep your privacy safe?” The user can select options associated with these questions, and re-identification server 146 can provide appropriate services. In other implementations, content providers identified as potentially tracking users and providing targeted content to users can be identified for each client device (e.g., in a blacklist or other list), and client devices can be configured not to transmit requests to such content providers, or may ignore or block content received from such content providers.

[0076] In some implementations, sender 162 can prevent content providers from delivering content to user devices by blocking all content sent from the content provider. Sender 162 can do this even if the user device requests content from a specific provider. In some implementations, sender 162 can generate and send reports to regulators instructing content providers to track users or deliver user-targeted content against the user's will. In some implementations, sender 162 can transmit messages to news media informing the public that a particular content provider is tracking a user, even if the user has explicitly chosen not to be tracked.

[0077] To prevent requests for content from being sent to the content provider, sender 162 may need to identify content providers that are tracking users or providing user-directed content to users. To do this, sender 162 may signal to provider identifier 156 to identify content providers performing unauthorized tracking. Provider identifier 156 may include applications, servers, services, daemons, routines, or other executable logic to identify content providers that track users and provide user-directed content, even if the user explicitly takes steps to block them. Provider identifier 156 can identify the provider by identifying the user-directed content and probing its source. Typically, the source of the user-directed content already leaves a fingerprint on the content, such as a tag indicating where the content came from, which provider identifier 156 can use to identify the provider. In some implementations, the content provider's fingerprint may be stored in the web browser that displays the content to the viewer. Provider identifier 156 can identify the provider from the fingerprint associated with the web browser.

[0078] In some implementations, a positive identification of unauthorized tracking or user-directed content may not be sufficient for application 160 to determine that a content provider is tracking a user or providing user-directed content against their will, even if the probability exceeds a predetermined threshold. In these implementations, application 160 may request content identifier 154 to identify a second content item provided by a content provider identified by provider identifier 156 to track a user or provide content against their will, and application 160 may again determine the probability that the identified content provider is tracking a user against their will or that the content is user-directed content. Application 160 may do this based on the intersection and / or signals or combinations of signals with the highest predictability of the content viewed by the user before receiving the second content item from the same identified content provider, a comparison of the provided content with the webpage on which the provided content is displayed, and the content of the provided content. If application 160 again determines that the content provider is performing unauthorized tracking on the user, or that the content is user-directed content, then sender 162 may redirect any requests from the user's device to the content provider that has been identified as tracking or providing user-directed content to the user against its will.

[0079] In some implementations, the re-identification server 146 implements a counter 158 to determine the number of times a particular content provider has been identified as tracking the user or providing user-directed content after the user has taken steps to avoid being tracked. The counter 158 may include an application, server, service, daemon, routine, or other executable logic to increment the counter in each instance where the application 160 determines that a content provider has been identified as tracking the user or providing user-directed content against their will. In each instance, the counter 158 may increment the counter associated with the particular content provider by 1. In some implementations, components of the re-identification server 146 may not prevent the content provider from receiving requests for content until the counter associated with the content provider reaches a predetermined threshold. Upon reaching the threshold, the sender 162 may execute the systems and methods described herein to prevent requests from user devices from reaching the content provider associated with the counter.

[0080] In some implementations, a user can reset the identification of content providers that deliver user-directed content or are tracking the user. The user can access a computing device that receives user-directed content or is being tracked by a content provider and choose the option to remove each content provider identified as either delivering user-directed content or tracking the user from an internal list of tracked content. The computing device can signal a re-identification server 146 to allow the transmission of requests and content to and from content providers on the list. The user can select all or some of the content providers on the list. In some implementations, if the method described herein is performed on a user device, the user device can request and allow content to be delivered from the selected content providers.

[0081] Now refer to Figure 2 According to some implementations, a block diagram of a neural network 200 is shown. This neural network 200 has inputs derived from the history of content viewed by the user and outputs indicating whether a third party is tracking the user's activity or whether the content is user-directed. In some implementations, the neural network 200 may be a reference... Figure 1B As part of the application 160 shown and described, and illustrated as a log including content item 202, input 204, hidden layer 208, and output layer 212. The neural network 200 may include any number of components. In some implementations, the neural network 200 may be implemented by a tensor processing unit dedicated to using machine learning to determine whether a content provider is performing unauthorized tracking or whether the content is user-oriented. The neural network 200 may include any number of components. The neural network 200 may operate as a log of content item 202 generated by the re-identification server 146, as shown and described with reference to FIG1, and is used as input to input 204. Signals or combinations of signals from input 204 are associated with weights and transmitted to hidden layer 208. Signals or combinations of signals from hidden layer 208 may be associated with weights and transmitted to output layer 212. Components 202, 204, 208, and 212 of the neural network 200 may be implemented to determine the probability that a content provider is performing unauthorized tracking against different users or that the content is user-oriented. Although shown as having one hidden layer, in some implementations, neural network 200 may include more than one hidden layer.

[0082] In some implementations, content log 202 may be a log of content items provided by anonymous computing devices. Re-identification server 146 can retrieve content log 202 from a web browser, which is associated with different computing devices that have been provided with the same content fragments. Content log 202 may include browsing history associated with different devices including web pages viewed by different users; keywords entered, selected, or associated with web pages or domains by users; and domains visited by users. In some implementations, content log 202 may include characteristics of each of these web pages identified by re-identification server 146. Content log 202 may also include intersections and / or signals or combinations of signals with the highest predictability of the same content viewed by multiple users before receiving the same content item from the content provider. Content log 202 may also include content items provided by the content provider to multiple users and the web pages on which the content items are displayed. Further, content log 202 may include a login page associated with the content item logged in when a user clicks on a content item. Content log 202 may include any number of content items and any type of content.

[0083] Input 204 is the first layer of neural network 200, representing the input layer of neural network 200. Input 204 can receive the content of content log 202 at node 206 as input to neural network 200. Each input can be a node associated with the input of content log 202, which sends a signal to each node in node 210 of hidden layer 208. Based on the numerical identification of the input in the database within reidentification server 146, the input from content log 202 can be converted by reidentification server 146 into numerical values, binary codes, matrices, vectors, etc. Input can be the provided content item; the characteristics of the provided content item; web page; domain; keywords; what the input is (i.e., the intersection and / or signal or combination of signals with the highest predictive power, the provided content, content landing page, etc.) and the characteristics of web page, domain, and other inputs. For example, the intersection web page of content log 202 can be associated with the number 12, for example, based on the characteristics of the web page (i.e., the web page can include characteristics common to the content provided by the content provider; the more similar common characteristics, the higher the value). Based on the input-related numbers in the database within reidentification server 146, reidentification server 146 can convert all inputs from content log 202 into numbers. Reidentification server 146 can then normalize the numbers to values ​​between -1 and 1 using any technique, so the operation can be performed by nodes 210 of hidden layer 208 with respect to the numbers. Reidentification server 146 can normalize numbers to any range of values. After reidentification server 146 has converted the numbers to values ​​between -1 and 1, neural network 200 can implement weights associated with each input and signal from each node 206 in node 206, and the signal or combination of signals can be transmitted to hidden layer 208.

[0084] Hidden layer 208 is a layer of node 210 that receives input signals or combinations of signals from input 204, performs one or more operations on the input signals or combinations of signals, and provides signals or combinations of signals to output layer 212. While one hidden layer is shown, any number of hidden layers can exist. In some implementations, hidden layer 208 may be associated with the number of users who viewed similar content before viewing the content provided by the content provider. Based on the value of node 206 and the weights associated with the signals or combinations of signals transmitted between node 206 and node 210 in hidden layer 208, neural network 200 can perform operations at hidden layer 208, such as multiplication, linear operations, sigmoid, hyperbolic tangent, etc. Signals or combinations of signals from node 210 can be sent to output layer 212, and each of these signals or combinations of signals can be associated with weights.

[0085] Output layer 212 may be a layer of neural network 200 dedicated to providing the probability of a content provider tracking a user without their consent or of whether the content is user-targeted. Output layer 212 is shown as comprising two nodes, unauthorized re-identification 214 and no unauthorized re-identification 216. Each node is associated with a probability determined by neural network 200 based on input from content log 202. If the input indicates a high probability that a content provider is tracking a user or that the content is user-targeted, unauthorized re-identification 214 may be associated with a high probability (e.g., probability greater than 50%), and no unauthorized re-identification 216 may be associated with a low probability (e.g., probability less than 50%). After the probabilities are associated with each of unauthorized re-identification 214 and no unauthorized re-identification 216, re-identification server 146 may compare the probabilities to a threshold determined by an administrator to determine whether a content provider is tracking a user against their will or whether the content is user-targeted.

[0086] The weights associated with the signal or combination of signals propagating between input 204 and hidden layer 208, and then between hidden layer 208 and output layer 212, can be automatically determined based on training data provided by the administrator. The training data can include content and content features that serve as input to neural network 200, as well as the expected output based on the input. Neural network 200 may initially have random weights associated with each of its signals or combinations of signals, but after a sufficient amount of training data has been input into neural network 200, the weights can be determined to obtain a sufficient degree of determinism as identified by the administrator. In some implementations of the systems and methods described herein, the input and signal or combination of signals associated with the highest weight can be features associated with content provided by the content provider, the landing page, and the domain in which the content is displayed when it is shown. Other inputs can include similarities to content provided by the content provider. To use the training data, the neural network can be a supervised system that implements backpropagation. After the training data is used as input and the neural network identifies the probability of the output, the neural network can identify the expected output from the training data and identify the difference between the actual output and the expected output. The neural network can identify the difference as a differential and modify the weights so that the actual output is closer to the expected output. Neural networks can use a learning rate, which identifies the degree of change in each weight, to modify the weights of their signals or combinations of signals for iteration on training data implemented in the neural network 200. As more and more training data is fed into the neural network, the weights of the signals or combinations of signals may change, and the difference may become smaller. Therefore, in some implementations, the results may become more accurate.

[0087] In some implementations, the neural network 200 can be a semi-supervised system, where the training data used as input to the system includes labeled data with and without output data, as well as data with input data. This is advantageous when a large amount of data is available, but humans would need to spend a significant amount of time labeling data with the correct output. In a semi-supervised system, the neural network 200 can receive only the input data to determine the output and label the data based on the output. The newly labeled data can then be fed into the neural network 200 along with the labeled dataset to train the neural network 200 using backpropagation. Using a semi-supervised system, the neural network can continuously update as it identifies content and determine whether the content provider is tracking the user against their will or whether the content is user-directed.

[0088] Now refer to Figure 3 According to some embodiments, a flowchart of method 300 is shown, which is used to determine whether a third party is tracking user activity based on the output from a neural network, or whether the content is user-targeted content. Method 300 may include any number of operations. In operation 302, the re-identification server may retrieve a log of content items. The log of content items may include any number of items, including but not limited to identified content items, characteristics of the content items, the web page on which the content items are provided, the content landing page, the characteristics of the content landing page, and data from the browsing history of users who have viewed the content items (i.e., domain, web page, keywords associated with the web page and domain, characteristics of the domain and web page, time periods indicating that items related to the content items in the browsing history were viewed, intersections and / or signals or combinations of signals with the highest predictive power of content viewed among different users, etc.). Each item in the content item log may be associated with a number based on the position of each content item in a table that associates the content item with numbers in a database within the re-identification server.

[0089] In operation 304, the re-identification server can identify a first content item from the content item's log. The first content item can be content provided by a content provider, which the re-identification server is using to determine whether the content provider is tracking the user after the user selects the option not to be tracked, or whether the content is user-directed content. In operation 306, the re-identification server can generate a set of identifiers for content items from the content item's log. Identifiers can be generated by user devices based on stored content associated with different users' web activities. The stored content records can be browsing history associated with the user's browser. Identifiers can be retrieved from any number of user devices. In some implementations, identifiers can be associated with content viewed during the time period in which the first content item was viewed.

[0090] In operation 308, the re-identification server may identify the intersection and / or signal or combination of signals with the highest predictive power among the browsing histories of different users in the content log. In some implementations, the intersection and / or signal or combination of signals with the highest predictive power are common websites, domains, and keywords associated with web pages and domains across multiple users. The re-identification server may identify common content items as the unique input to the re-identification server's neural network. The neural network can use these inputs and determine the probability that the content provider offering the first content item tracked the user who viewed the first content item or the content is user-directed content before providing the content. If the probability is not higher than a predetermined threshold, in operation 310, the re-identification server may determine that it is unlikely that the content was provided due to unauthorized identification and continue to transmit content from the provider to the computing device.

[0091] However, if the re-identification server determines the probability is above a threshold, then in operation 312, the re-identification server can identify the content provider that delivered the first content item. The re-identification server can do this using information from the browser receiving the first content item, thereby identifying the source of the first content item. In some implementations, the re-identification server can identify the content provider based on the web pages most frequently visited by the user before the first content item was presented. For example, if most users visited website A before the first content item was presented, the re-identification server can determine that the tracking likely started or occurred at website A. This is advantageous if the content provider uses an incorrect identifier when delivering the content to avoid any identification by the browser displaying the content item.

[0092] In operation 314, the re-identification server may increment the counter associated with the content provider identified in operation 312. The re-identification server may store and increment any number of counters for content providers, and increment the corresponding counter for each content provider in each instance where the content provider is determined to be tracking users against their will (i.e., after the user explicitly or passively disables third-party cookies or takes other security measures) or if the content is user-directed content. In operation 316, after the counter associated with the identified content provider is incremented, the re-identification server may determine whether the counter exceeds a predetermined threshold. The predetermined threshold may be determined by an administrator and may be any number. If the counter is determined to be no higher than the threshold, then in operation 310, any content provided by the content provider may be transmitted to the user device.

[0093] However, if the counter is determined to be above a predetermined threshold, in operation 318, the reidentification server can receive a request for content from the identified content provider from the user device, and in operation 320, redirect the request to a second content provider that has not yet been established to track users against their will, or if the content is user-directed content. Based on each determination related to the content provider, the reidentification server can update its neural network. Furthermore, the reidentification server can repeatedly perform the above operations for any number of content items provided to different users.

[0094] Now refer to Figure 4 According to some embodiments, a flowchart of another method 400 is shown, which is used to determine whether a third party is tracking user activity based on the output from a neural network, or whether the content is user-oriented content. Method 400 can be performed by a re-identification server or any server. Operations 402, 404, 406, 408, 410, and 412 can be compared with reference to... Figure 3 The corresponding operations 302, 304, 306, 308, 310, and 312 shown and described are the same or similar. After identifying a content provider that may be tracking the user against their will, or if the content is user-directed content, in operation 414, the re-identification server may identify a third content item provided by the identified content provider. Based on the browsing history of the user's device, the re-identification server may identify the third content item from content logs provided by multiple user devices. In some implementations, the re-identification server may identify the third content item from the content logs retrieved in operation 402. In some implementations, the re-identification server may identify the third content item after retrieving another content log from multiple user devices.

[0095] In operation 416, the reidentification server can generate an identifier from the content log associated with the third content item. The reidentification server can identify the content and its characteristics, similar to how it does so in operation 406. The reidentification server can identify the intersection and / or signal or combination of signals with the highest predictive power, representing common content viewed by multiple users before viewing the third content item, and input this intersection and / or signal or combination of signals into the reidentification server's neural network. In operation 418, the reidentification server can determine, through the neural network, whether the output of the neural network based on the input of the intersection and / or signal or combination of signals with the highest predictive power exceeds a threshold. If the output does not exceed the threshold, then in operation 420, the reidentification server can continue to transmit the request for content from the user device to the content provider.

[0096] However, if the probability exceeds a threshold, in operation 422, the reidentification server can receive requests for content from user devices based on the identified content provider. In operation 424, the reidentification server can redirect the request to a second content provider that has not yet been established to track users against its will, or that has not provided targeted content. Based on each determination related to a content provider, the reidentification server can update its neural network. Furthermore, the reidentification server can repeatedly perform the above operations for any number of content items provided to different users.

[0097] Advantageously, by implementing the systems and methods described herein, the system can determine whether a third-party content provider is tracking the web activity of users they do not wish to be tracked, or whether the content is user-targeted. Previous methods have been unsuccessful in determining whether a third-party content provider is tracking users or whether the content is user-targeted, because the same content can be displayed to the user regardless of whether they are being tracked. However, by feeding input into a neural network that analyzes the browser history of different users before the user is served content, the systems and methods described herein can automatically identify whether the content is user-targeted or when the content provider is tracking the user, and prevent them from doing so. Therefore, users can feel secure in terms of privacy when searching for web pages on the Internet.

[0098] Regarding the systems discussed in this paper that collect or utilize personal information related to users, users can be provided with the opportunity to control whether a program or feature can collect personal information (such as information related to a user's social networks, social actions or activities, user preferences, or user location), or to control whether and / or how content that may be more relevant to the user is received from content servers or other data processing systems. Furthermore, before specific data is stored or used, it can be anonymized in one or more ways, such that personally identifiable information is removed when generating parameters. For example, a user's identity can be anonymized, making it impossible to identify the user personally, or the user's geographic location can be generalized, where location information is obtained (such as generalized to a city, postal code, or state), making it impossible to determine the user's specific location. Therefore, users can control how information about them is collected and used by content servers.

[0099] The implementations of the subject matter and operations described in this specification can be implemented in digital electronic circuit systems or computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents) or combinations thereof. Implementations of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on one or more computer storage media for execution by or control of the operation of a data processing device. Alternatively or additionally, the program instructions can be encoded on artificially generated propagating signals (e.g., machine-generated electrical, optical, or electromagnetic signals generated to encode information for transmission to a suitable receiver device for execution by the data processing device). The computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or combinations thereof, or be included therein. Furthermore, when the computer storage medium is not a propagating signal, it can be a source or destination of computer program instructions encoded in artificially generated propagating signals. The computer storage medium can also be one or more separate components or media (e.g., multiple CDs, disks, or other storage devices), or be included therein. Therefore, the computer storage medium can be tangible.

[0100] The operations described in this specification can be implemented as operations performed by a data processing device on data stored on one or more computer-readable storage devices or received from other sources.

[0101] The terms "client" or "server" include all kinds of devices, equipment, and machines for processing data, such as programmable processors, computers, systems-on-a-chip, or a combination thereof. Devices may include special-purpose logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the device may also include code that creates an execution environment for the computer program under consideration, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The device and execution environment can implement various computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0102] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any form of programming language (including compiled or interpreted languages, declarative or programming languages), and can be deployed in any form (including as standalone programs, or as modules, components, subroutines, objects, or other units suitable for a computing environment). Computer programs may, but are not required to, correspond to files in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), or as a single file dedicated to the program under development, or as multiple collaborating files (e.g., files storing one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on a single computer, or on multiple computers located at a single site or distributed across multiple sites and interconnected via a communication network.

[0103] The processes and logic flows described in this specification can be executed by one or more programmable processors, which execute one or more computer programs to perform actions by manipulating input data and generating output. The processes and logic flows can also be executed by a dedicated logic circuit system (e.g., an FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit)), and the apparatus can also be implemented as such a dedicated logic circuit system.

[0104] Processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. Essential components of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., disks, magneto-optical disks, or optical disks) for storing data, or the computer will be operatively coupled to receive data from or transfer data to such mass storage devices, or both. However, a computer does not necessarily need to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including: semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices); magnetic disks (such as internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by dedicated logic circuitry or incorporated into such dedicated logic circuitry.

[0105] To provide interaction with the user, the implementation of the subject matter described in this specification can be carried out on a computer having: a display device for displaying information to the user, such as a CRT (cathode ray tube), LCD (liquid crystal display), OLED (organic light-emitting diode), TFT (thin-film transistor), plasma, other flexible configuration, or any other monitor; and a keyboard, pointing device, such as a mouse, trackball, etc.; or a touch screen, touchpad, etc., through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form (including acoustic input, voice input, or tactile input). Additionally, the computer can interact with the user by sending documents to and receiving documents from a device used by the user; and by sending web pages to a web browser on the user's client device in response to a request received from a web browser.

[0106] The implementations of the subject matter described in this specification can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., client computers with graphical user interfaces or web browsers through which users can interact with the implementations of the subject matter described in this specification), or computing systems that include one or more such backend components, middleware components, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Communication networks can include local area networks (“LANs”) and wide area networks (“WANs”), interconnected networks (e.g., the Internet), and peer-to-peer networks (e.g., self-organizing peer-to-peer networks).

[0107] While this specification contains numerous specific implementation details, these details should not be construed as limiting the scope of any invention or potentially claimed content, but rather as descriptions of features specific to particular implementations of a particular invention. Certain features described in the context of individual implementations in this specification may also be implemented in combination within a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations. Furthermore, although features may be described above as functioning in certain combinations, or even as originally claimed, one or more features from a claimed combination may, in some cases, be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.

[0108] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the indicated specific order or in a sequential order, or that all illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above implementations should not be interpreted as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated together in a single software product or encapsulated in multiple software products.

[0109] Therefore, specific implementations of this subject matter have been described. Other implementations are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some implementations, multitasking or parallel processing can be utilized.

Claims

1. A method for detecting third-party re-identification of an anonymous computing device, comprising: receiving, by a client device, a plurality of first content items over a predetermined time period; subsequently transmitting, by the client device, a first request for a content item from a content provider; in response to the first request, receiving, by the client device, a second content item from the content provider; and transmitting, by the client device, a second request for a content item from the content provider, the second request being blocked by an intermediary server in response to the intermediary server determining that a signal or combination of signals having a highest predictive power between a first set of identifiers of a third content item retrieved by a first anonymous computing device and a second set of identifiers of a fourth content item retrieved by a second anonymous computing device exceeds a threshold, the third content item and the fourth content item being retrieved prior to the respective first and second anonymous computing devices receiving the second content item.

2. The method of claim 1, wherein, the identifiers of the fourth content item include identifiers of web pages accessed by the second anonymous computing device.

3. The method of claim 1, wherein, the identifiers of the fourth content item include identifiers of domains accessed by the second anonymous computing device.

4. The method of claim 1, wherein, the identifiers of the fourth content item include identifiers of keywords associated with domains accessed by the second anonymous computing device.

5. The method of claim 1, wherein, the intermediary server is configured to determine that a size of the signal or combination of signals having the highest predictive power between the first set of identifiers and the second set of identifiers exceeds a size of a signal or combination of signals having the highest predictive power between another pair of sets of identifiers.

6. The method of claim 1, wherein, the intermediary server determines that the signal or combination of signals having the highest predictive power between the first set of identifiers and the second set of identifiers is common to a third set of identifiers.

7. The method of claim 1, wherein, the intermediary server is configured to block the second request by redirecting the second request to a second content provider.

8. The method of claim 1, wherein, the intermediary server is configured to block the second request by: identifying a fifth content item, the fifth content item being provided by the content provider to a third anonymous computing device and a fourth anonymous computing device; for the third and fourth anonymous computing devices, generating a set of identifiers of a sixth content item retrieved by the third and fourth anonymous computing devices prior to receiving the fifth content item; and determining that the signal or combination of signals having the highest predictive power between a first set of identifiers of the sixth content item and a second set of identifiers of the sixth content item exceeds the threshold.

9. The method of claim 1, wherein, the intermediary server is configured to increment a counter associated with the content provider in response to determining that the signal or combination of signals having the highest predictive power between the first set of identifiers and the second set of identifiers exceeds the threshold.

10. The method of claim 9, wherein, the intermediary server is configured to block transmission of the second request further in response to a counter associated with the content provider exceeding a second threshold.

11. A system for detecting third-party re-identification of an anonymous computing device, comprising: A client device comprising a processor, a memory device, and a network interface, the processor configured to: receive, via the network interface, a plurality of first content items over a predetermined time period; subsequently transmit, via the network interface, a first request for a content item from a content provider; in response to the first request, receive, via the network interface, a second content item from the content provider; and transmit, via the network interface, a second request for a content item from the content provider, the second request blocked by an intermediary server in response to the intermediary server determining that a signal or combination of signals having a highest predictive power between a first set of identifications of third content items retrieved by a first anonymized computing device and a second set of identifications of fourth content items retrieved by a second anonymized computing device exceeds a threshold, the third content items and the fourth content items retrieved prior to the respective first anonymized computing device and second anonymized computing device receiving the second content item.

12. The system of claim 11, wherein, The identifications of the fourth content items comprise identifications of web pages accessed by the second anonymized computing device.

13. The system of claim 11, wherein, The identifications of the fourth content items comprise identifications of domains accessed by the second anonymized computing device.

14. The system of claim 11, wherein, The identifications of the fourth content items comprise identifications of keywords associated with domains accessed by the second anonymized computing device.

15. The system of claim 11, wherein, The intermediary server is configured to determine that a size of the signal or combination of signals having the highest predictive power between the first set of identifications and the second set of identifications exceeds a size of a signal or combination of signals having the highest predictive power between another pair of sets of identifications.

16. The system of claim 11, wherein, The intermediary server determines that the signal or combination of signals having the highest predictive power between the first set of identifications and the second set of identifications is common to a third set of identifications.

17. The system of claim 11, wherein, The intermediary server is configured to block the second request by redirecting the second request to a second content provider.

18. The system of claim 11, wherein, The intermediary server is configured to block the second request by: identifying a fifth content item, the fifth content item provided by the content provider to a third anonymized computing device and a fourth anonymized computing device; for the third anonymized computing device and the fourth anonymized computing device, generating a set of identifications of sixth content items retrieved by the third anonymized computing device and the fourth anonymized computing device prior to receiving the fifth content item; and determining that the signal or combination of signals having the highest predictive power between a first set of identifications of sixth content items and a second set of identifications of sixth content items exceeds the threshold.

19. The system of claim 11, wherein, The intermediary server is configured to increment a counter associated with the content provider in response to determining that the signal or combination of signals having the highest predictive power between the first set of identifications and the second set of identifications exceeds the threshold.

20. The system of claim 19, wherein, The intermediary server is configured to block transmission of the second request further in response to a counter associated with the content provider exceeding a second threshold.

Citation Information

Patent Citations

  • Automatic passive and anonymous feedback system

    CN102572539A

  • Data perturbation and anonymization using one-way hash

    CN103562851A