Low-Entropy Browsing History for Pseudo-Personalization of Content

Anonymized content retrieval via pseudo-personalization using aggregated browsing history clusters addresses privacy and resource concerns in personalized content delivery, maintaining relevance without exposing individual devices.

JP7813317B2Active Publication Date: 2026-02-12GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024097488
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-27
Filing Date
2024-06-17
Publication Date
2026-02-12
Estimated Expiration
2040-03-31

AI Technical Summary

Technical Problem

Personalized content delivery based on user and device identifying information leads to privacy risks and excessive computing resource consumption.

Method used

Anonymized content retrieval through pseudo-personalization using aggregated browsing history of devices, employing sparse matrix construction and dimensionality reduction to create pseudo-personalized clusters without exposing individual device details.

Benefits of technology

Maintains content relevance while ensuring user privacy and reducing computing resources by aggregating similar browsing histories into clusters, allowing content providers to infer user interests without identifying individual devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007813317000001
    Figure 0007813317000001
  • Figure 0007813317000002
    Figure 0007813317000002
  • Figure 0007813317000003
    Figure 0007813317000003
Patent Text Reader

Abstract

To provide a system and method for content retrieval.SOLUTION: The present disclosure provides systems and methods for content quasi-personalization or anonymized content retrieval via aggregated browsing history of a large number of devices, such as millions or billions of devices. A sparse matrix may be constructed from the aggregated browsing history, and dimensionally reduced, resulting in reducing entropy and providing anonymity for individual devices. Relevant content may be selected via quasi-personalized clusters representing similar browsing histories, without exposing individual device details to content providers.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. patent application Ser. No. 16 / 698,548, entitled "Low Entropy Browsing History for Content Quasi-Personalization," filed Nov. 27, 2019, which claims priority to U.S. patent application Ser. No. 16 / 535,912, entitled "Low Entropy Browsing History for Content Quasi-Personalization," filed Aug. 8, 2019, and U.S. patent application Ser. No. 62 / 887,902, entitled "Low Entropy Browsing History for Content Quasi-Personalization," filed Aug. 16, 2019, each of which is incorporated by reference herein in its entirety. [Background technology]

[0002] In a computer networked environment such as the Internet, content providers may provide content items that are inserted into information resources (e.g., web pages) that are processed and rendered by applications (e.g., web browsers) running on client devices. Summary of the Invention [Problem to be solved by the invention]

[0003] Personalized content delivery is typically based on capturing user and / or device identifying information, such as the device's browsing or access history, which can result in the collection of personally traceable data, expose users to potential security risks, and consume excessive computing resources. [Means for solving the problem]

[0004] The systems and methods discussed herein enable pseudo-personalization of content or anonymized content retrieval via the aggregated browsing history of large numbers of devices, such as millions or billions of devices. A sparse matrix can be constructed from the aggregated browsing history and dimensionally reduced, reducing entropy and providing anonymity for individual devices. Without exposing individual device details to content providers, relevant content can be selected via pseudo-personalized clusters representing similar browsing histories.

[0005] In one aspect, the present disclosure is directed to a method for anonymized content retrieval. The method includes generating, by a browser application of a computing device, a profile based on a browsing history of the computing device. The method also includes encoding, by the browser application, the profile as an n-dimensional vector. The method also includes calculating, by the browser application, a dimensionality-reduced vector from the n-dimensional vector. The method also includes determining, by the browser application, a first cluster corresponding to the dimensionality-reduced vector. The method also includes sending, by the browser application, a request for an item of content to a content server, the request comprising identification information of the first cluster. The method also includes receiving, by the browser application, from the content server, the item of content selected according to the identification information of the first cluster.

[0006] In some implementations, the method includes generating a profile based on a browsing history of a user of a computing device by identifying, from a browsing history log, a number n of accesses to each of a plurality of addresses within a predetermined time period. In some implementations, the method includes generating a string with a value representing each of one or more accesses to an address associated with a corresponding position in the string within the predetermined time period.

[0007] In some implementations, the method includes performing a singular value decomposition of the n-dimensional vector. In further implementations, the method includes receiving a set of singular vectors of the singular value decomposition from a second computing device. In still further implementations, the method includes transmitting the n-dimensional vector to the second computing device, the second computing device calculating the set of singular vectors based on an aggregation of the n-dimensional vector of the computing device and the n-dimensional vector of at least one other computing device.

[0008] In some implementations, the method includes receiving, from the second computing device, boundaries of each of the plurality of clusters. In further implementations, the method includes selecting a first cluster of the plurality of clusters in response to the dimensionality-reduced vector being within a boundary of the first cluster.

[0009] In some implementations, the method includes receiving, by the browser application from the second computing device, weights of a neural net model determined based on an aggregation of the n-dimensional vector of the computing device and the n-dimensional vector of at least one other computing device; applying, by a machine learning system of the browser application, the neural net model to the dimensionality-reduced vector to generate a ranking of a predetermined set of clusters; and selecting, by the browser application, the first cluster as the highest-ranked cluster of the predetermined set of clusters.

[0010] In another aspect, the present disclosure is directed to a method for anonymized content retrieval. The method includes receiving, by a server computing device, from each of a plurality of client computing devices, a profile based on the browsing history of the corresponding client computing device, each profile comprising an n-dimensional vector. The method also includes aggregating, by the server computing device, the n-dimensional vectors of the plurality of profiles into a matrix. The method also includes computing, by the server computing device, a singular value decomposition of the matrix to generate a set of singular values. The method also includes transmitting, by the server computing device, the set of singular values ​​to each of the plurality of client computing devices and at least one content provider device.

[0011] In some implementations, the method includes determining, by a server computing device, a boundary of each cluster in the set of clusters of the matrix. In further implementations, the method includes transmitting, by the server computing device to each of the plurality of client computing devices and to the at least one content provider device, a boundary of each cluster in the set of clusters of the matrix.

[0012] In some implementations, the method includes identifying, by a server computing device, each cluster of the set of clusters from the matrix via a neural net model. In further implementations, the method includes transmitting weights of the neural net model to each of the plurality of client computing devices and the at least one content provider device.

[0013] In yet another aspect, the present disclosure is directed to a system for anonymized content retrieval. The system includes a computing device having a network interface in communication with a content server, a memory that stores a browsing history of the computing device, and a browser application. The browser application is configured to generate a profile based on the browsing history of the computing device, encode the profile as an n-dimensional vector, calculate a dimensionality-reduced vector from the n-dimensional vector, determine a first cluster corresponding to the dimensionality-reduced vector, send a request for an item of content to the content server via the network interface, the request including identification information of the first cluster, and receive from the content server via the network interface the item of content selected according to the identification information of the first cluster.

[0014] In some implementations, the browser application is further configured to generate the string with a value representing each of one or more accesses to an address associated with a corresponding position in the string within a predetermined time period. In some implementations, the browser application is further configured to perform singular value decomposition of the n-dimensional vector. In further implementations, the browser application is further configured to receive a set of singular vectors of the singular value decomposition from a second computing device via the network interface. In still further implementations, the browser application is further configured to transmit the n-dimensional vector to the second computing device via the network interface, and the second computing device calculates the set of singular vectors based on an aggregation of the n-dimensional vector of the computing device and the n-dimensional vectors of at least one other computing device.

[0015] In some implementations, the browser application is further configured to receive, from the second computing device via the network interface, weights of a neural net model determined based on an aggregation of the n-dimensional vector of the computing device and the n-dimensional vector of at least one other computing device, apply the neural net model to the dimensionality-reduced vector to generate a ranking of a predetermined set of clusters, and select the first cluster as the highest-ranked cluster of the predetermined set of clusters.

[0016] At least one aspect is directed to a method for encoding an identifier for content selection. A first application executing on a client device can identify a browsing history maintained on the client device. The browsing history can record information resources accessed by the client device through the first application. The first application can apply a classification model to the browsing history of the first application to identify a class into which the first application should be categorized. The first application can assign the first application to a class identifier corresponding to the identified class. The class identifier of the first application can be identical to the class identifier of the second application. The first application can receive, from a content publisher device, an information resource comprising primary content and a content slot available for receiving content from a content selection service. The first application can generate a request for content for a content slot of the information resource, the request including the class identifier of the first application. The first application can send the request for content to the content selection service. The content selection service can select a content item to insert into the content slot of the information resource using the class identifier of the first application and the class identifier of the second application.

[0017] In some implementations, the first application can receive from the application administrator service a set of parameters for updating a classification model for categorizing applications into one of a plurality of classes. In some implementations, the first application can modify the classification model maintained on the client device based on the set of parameters received from the application administrator service. In some implementations, the first application can train the classification model maintained on the client device using a distributed learning protocol in collaboration with multiple applications executing on a corresponding plurality of client devices.

[0018] In some implementations, the first application can use a dimensionality reduction process to generate a reduced set of feature vectors from the browsing history determined from the client device, the feature vectors having a smaller file size than the browsing history. In some implementations, applying the classification model can include applying the classification model to the reduced set of feature vectors generated from the browsing history.

[0019] In some implementations, the first application may determine that a classification model should be applied to the browsing history in accordance with the identifier assignment policy. In some implementations, applying the classification model may include applying the classification model in response to determining that the classification model should be applied in accordance with the identifier assignment policy.

[0020] In some implementations, applying the classification model may include applying the classification model to identify a class from a plurality of classes. Each class of the plurality of classes may have at least a predetermined number of classes to be categorized into the class. In some implementations, assigning the first application to a class identifier may further include assigning the first application to a class identifier of the plurality of class identifiers. Each class identifier may correspond to one of the plurality of classes.

[0021] In some implementations, generating the request for the content may include generating the request for the content without a unique tracking identifier corresponding to an account associated with the first application, the first application, or the client device provided by the content selection service. In some implementations, generating the request for the content may include generating a request for the content comprising a secure cookie for transmission over a connection with the content selection service. The secure cookie may have a class identifier of the first application. In some implementations, identifying a browsing history may include identifying a browsing history over a predetermined time range to which the classification model should be applied.

[0022] At least one aspect is directed to a system for encoding an identifier for content selection. The system may include a first application executable on a client device having one or more processors. The first application executing on the client device may identify a browsing history maintained on the client device. The browsing history may record information resources accessed by the client device via the first application. The first application may apply a classification model to the browsing history of the first application to identify a class into which the first application should be categorized. The first application may assign the first application to a class identifier corresponding to the identified class. The class identifier of the first application may be identical to the class identifier of the second application. The first application may receive, from a content publisher device, an information resource comprising primary content and an available content slot for receiving content from a content selection service. The first application may generate a request for content for the content slot of the information resource, the request including the class identifier of the first application. The first application may send the request for content to the content selection service. The content selection service can use the class identifier of the first application and the class identifier of the second application to select a content item to insert into a content slot of the information resource.

[0023] In some implementations, the first application can receive from the application administrator service a set of parameters for updating a classification model for categorizing applications into one of a plurality of classes. In some implementations, the first application can modify the classification model maintained on the client device based on the set of parameters received from the application administrator service. In some implementations, the first application can train the classification model maintained on the client device using a distributed learning protocol in collaboration with multiple applications executing on a corresponding plurality of client devices.

[0024] In some implementations, the first application can use a dimensionality reduction process to generate a reduced set of feature vectors from the browsing history determined from the client device, where the feature vectors have a smaller file size than the browsing history. In some implementations, the first application can apply a classification model to the reduced set of feature vectors generated from the browsing history.

[0025] In some implementations, the first application can determine that a classification model should be applied to the browsing history in accordance with the identifier assignment policy. In some implementations, the first application can apply the classification model in response to determining that a classification model should be applied in accordance with the identifier assignment policy.

[0026] In some implementations, the first application can apply a classification model to identify a class from a plurality of classes. Each class of the plurality of classes can have at least a predetermined number of clients to be categorized into the class. In some implementations, the first application can assign the first application to a class identifier of the plurality of class identifiers. Each class identifier can correspond to one of the plurality of classes.

[0027] In some implementations, the first application generates a request for content without a unique tracking identifier corresponding to an account associated with the first application, the first application, or the client device provided by the content selection service. In some implementations, the first application generates a request for content that comprises a secure cookie for transmission over a connection with the content selection service. The secure cookie may have a class identifier of the first application. In some implementations, the first application can identify browsing history over a predetermined time range to which to apply the classification model.

[0028] The present disclosure also provides a computer program comprising instructions that, when executed by a computing device, cause the computing device to perform any of the methods disclosed herein. The present disclosure also provides a computer-readable medium comprising instructions that, when executed by a computing device, cause the computing device to perform any of the methods disclosed herein.

[0029] At least one aspect is directed to a method for encoding an identifier for content selection. The method may include identifying a plurality of information resources accessed via a first application executing on a client device. The method may include extracting, from each information resource of the plurality of information resources, features corresponding to at least a portion of the content of the information resource. The method may include applying a classification model to the features extracted from the plurality of information resources to identify a set of classes into which the first application should be categorized. The method may include determining that a class from the set of classes satisfies a threshold number of applications assigned to that class. The method may include assigning the first application to a class identifier corresponding to the class in response to determining that the class satisfies the threshold number. The class identifier of the first application may be identical to the class identifier of the second application. The method may include receiving, from a content publisher device for presentation via the first application, an information resource comprising primary content and a content slot available for receiving content from a content selection service. The method may include generating a request for content, including the class identifier of the first application, for a content slot of the information resource. The method can include sending a request for content to a content selection service, where the content selection service can use the class identifier of the first application and the class identifier of the second application to select a content item to insert into a content slot of the information resource.

[0030] In some implementations, the method may include, in response to receiving the information resource, selecting a class from the set of classes according to an obfuscation policy, which may specify conditions under which a corresponding class identifier is allowed to be included in a request for content associated with a content slot of the information resource.

[0031] In some implementations, the method may include, in response to receiving the second information resource, selecting, in accordance with the obfuscation policy, a second class identifier of the set of classes identified from applying the classification model. The second class identifier may be different from the class assigned to the first application. In some implementations, the method may include generating, for a content slot of the second information resource, a second request for content including a second class identifier corresponding to the second class in place of the class identifier corresponding to the class.

[0032] In some implementations, the method may include, in response to receiving the second information resource, determining, according to the obfuscation policy, not to include any class identifier in a second request for content to be inserted into a content slot of the second information resource. In some implementations, the method may include, in response to determining not to include any class identifier, sending a second request for content to a content selection service, the second request for content lacking any class identifier.

[0033] In some implementations, determining that the class meets the threshold number may include executing a threshold encryption protocol in collaboration with a class authorization service using an encrypted class identifier generated using a class identifier corresponding to the class. In some implementations, determining that the class meets the threshold number may include checking the class identifier against a probabilistic data structure for class identifiers maintained by the class authorization service.

[0034] In some implementations, transmitting the request for the content may include transmitting the request for the content. The content selection service may use the class identifier to maintain an aggregate browsing history for the first application and the second application. In some implementations, applying the classification model may include applying the classification model to identify a set of classes from the plurality of classes that are within a proximity threshold of each other in a feature space defined by the classification model.

[0035] In some implementations, the method may include using a dimensionality reduction process to generate a reduced set of feature vectors from the browsing history identified from the client device, the feature vectors having a smaller file size than the browsing history. In some implementations, applying the classification model may include applying the classification model to the reduced set of feature vectors generated from the browsing history. In some implementations, extracting features further comprises extracting features from at least a portion of content of the information resource, the portion of content including at least one of textual data, visual data, or audio data.

[0036] At least one aspect is directed to a system for encoding identifiers for content selection. The system may include a first application executable on a client device having one or more processors. The first application may identify a plurality of information resources to be accessed via the first application executing on the client device. The first application may extract features from each of the plurality of information resources corresponding to at least a portion of the content of the information resource. The first application may apply a classification model to the features extracted from the plurality of information resources to identify a set of classes into which the first application should be categorized. The first application may determine that a class from the set of classes satisfies a threshold number of applications assigned to that class. In response to determining that the class satisfies the threshold number, the first application may assign the first application to a class identifier corresponding to the class. The class identifier of the first application may be identical to the class identifier of the second application. The first application may receive, from a content publisher device for presentation via the first application, an information resource comprising primary content and an available content slot for receiving content from a content selection service. The first application can generate a request for content for a content slot of the information resource, the request including a class identifier of the first application. The first application can send the request for content to a content selection service. The content selection service can use the class identifier of the first application and the class identifier of the second application to select a content item to insert into the content slot of the information resource.

[0037] In some implementations, in response to receiving the information resource, the first application can select a class from the set of classes according to an obfuscation policy, which can specify conditions under which a corresponding class identifier is allowed to be included in a request for content associated with a content slot of the information resource.

[0038] In some implementations, in response to receiving the second information resource, the first application can select a second class identifier of the set of classes identified from applying the classification model according to the obfuscation policy. The second class identifier can be different from the class assigned to the first application. In some implementations, the first application can generate a second request for content for a content slot of the second information resource, the second request including a second class identifier corresponding to the second class instead of the class identifier corresponding to the class.

[0039] In some implementations, in response to receiving the second information resource, the first application can determine, in accordance with the obfuscation policy, not to include any class identifiers in a second request for content to be inserted into a content slot of the second information resource. In some implementations, in response to determining not to include any class identifiers, the first application can send a second request for content to the content selection service, the second request for content lacking any class identifiers.

[0040] In some implementations, the first application can determine that the class meets the threshold number by performing a threshold encryption protocol in collaboration with the class authorization service using an encrypted class identifier generated using a class identifier corresponding to the class. In some implementations, the first application can determine that the class meets the threshold number by checking the class identifier against a probabilistic data structure for class identifiers maintained by the class authorization service.

[0041] In some implementations, the first application can send a request for content. The content selection service can use the class identifier to maintain an aggregate browsing history of the first application and the second application. In some implementations, the first application can apply a classification model to identify a set of classes from a plurality of classes that are within a proximity threshold of each other in a feature space defined by the classification model.

[0042] In some implementations, the first application can use a dimensionality reduction process to generate a reduced set of feature vectors from the browsing history identified from the client device, the feature vectors having a smaller file size than the browsing history. In some implementations, the first application can apply a classification model to the reduced set of feature vectors generated from the browsing history. In some implementations, the first application can extract features from at least a portion of content of the information resource, the portion of content including at least one of textual data, visual data, or audio data.

[0043] Any optional feature of one aspect may be combined with any other aspect.

[0044] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the present disclosure will become apparent from the description, drawings, and claims. [Brief explanation of the drawings]

[0045] [Figure 1A] FIG. 1 is a diagram of an example profile vector, according to some implementations. [Figure 1B] FIG. 1 is a diagram of aggregation of profile vectors into a sparse matrix according to some implementations. [Figure 1C] FIG. 1 is a diagram of a process for anonymization to provide pseudo-personalized clustering, according to some implementations. [Figure 2] FIG. 1 is a block diagram of a system for anonymization to provide pseudo-personalized clustering, according to some implementations. [Figure 3] 1 is a flowchart of a method for anonymization to provide pseudo-personalized clustering, according to some implementations. [Figure 4] FIG. 1 is a block diagram illustrating a system for encoding identifiers for content selection using classification models, according to an example implementation. [Figure 5] FIG. 1 is a block diagram illustrating a client device and an application management service in a system for encoding identifiers for content selection using classification models, according to an example implementation. [Figure 6] FIG. 1 is a block diagram illustrating a client device, a content provider, a content publisher, and a content selection service in a system for encoding identifiers for content selection using a classification model, according to an exemplary implementation. [Figure 7] 1 is a block diagram illustrating a client device and a content selection service in a system for encoding identifiers for content selection using a classification model, according to an example implementation. [Figure 8] FIG. 1 is a flow diagram illustrating a method for encoding an identifier for content selection using a classification model, according to an example implementation. [Figure 9] FIG. 1 is a block diagram illustrating the general architecture of a computer system that may be utilized to implement elements of the systems and methods described and illustrated herein, according to an exemplary implementation. DETAILED DESCRIPTION OF THE INVENTION

[0046] Like reference numbers and designations in the various drawings indicate like elements.

[0047] Personalized content delivery is typically based on capturing user and / or device identification information, such as the device's browsing or access history. For example, a device may provide identification information, such as a device identifier, account name, cookies, or other such information, and content providers may store this information for use in selecting personalized content. As a result, content providers may obtain a large amount of data about individuals. This has significant implications for the privacy and security of devices and users. Opt-out and do-not-track policies give users some control over their privacy (as long as content providers comply with those policies). However, these policies impair the content provider's ability to provide relevant content. Furthermore, requests generated according to such policies may be completely devoid of any user requirements, preventing proper analysis of such requests.

[0048] The systems and methods discussed herein enable pseudo-personalization of content or anonymized content retrieval via the aggregated browsing history of large numbers of devices, such as millions or billions of devices. For example, the browsing history of each device can be encoded as a long data string, or an n-dimensional vector. FIG. 1A is a diagram of an exemplary profile vector 100 according to some implementations. The profile vector may comprise an identification of visits or accesses by the device to domains, websites, or web pages within a given period of time. In the illustrated example, the vector specifies the number of views or visits to each of a number of domains 1-n for each day of the week. While shown as a one-dimensional vector, in similar implementations, the vector may comprise an n-by-7 array (e.g., one row for each day of the week). In other implementations, additional data may be included (e.g., each day of the month, time of day, etc.). Thus, the vector may comprise a large n-dimensional vector or array. Additional data, such as an identification of the time of creation, location, IP address, or any other such information, may be included in the vector.

[0049] As noted above, in many implementations, the vector may be very large. For example, there are approximately 1.8 billion active websites and over 330 million registered domains on the Internet. In some implementations, the vector may record visits or accesses to any of these domains or websites. In other implementations, the vector may record visits or accesses to only a subset of the domains or websites. For example, fewer than 1 million websites account for approximately 50% of web traffic. Thus, in many implementations, the vector may only record or represent accesses or visits to a subset of the domains or websites. Nevertheless, even in many such implementations, the vector may be large, e.g., 2 26 That's the order.

[0050] As shown in the diagram of FIG. 1B, the vector may be provided to a server device, which may aggregate the vector with vectors from other devices. The profile vectors 100 from each of multiple devices 102 may be aggregated to create a very large matrix. For example, there are billions of users (e.g., 2 million) on the web per month. 30 There are (order of magnitude) active users or devices. Therefore, the matrix that combines the profile vectors 100 of each of these devices is 2 30 ×2 26 , or in some implementations may be larger.

[0051] However, this matrix is ​​very sparse: a typical user may visit hundreds of millions of potential domains in a given period, such as a week. Due to this highly sparse nature and redundancy of the browsing history of any given device, it is possible to reduce the dimensionality of this matrix.

[0052] In some implementations, a linear dimensionality reduction technique such as singular vector decomposition (SVD) may be used to calculate a rank X matrix that is the best approximation to the matrix (e.g., minimizes the squared error). Each profile vector 100 may be projected into X dimensions, where X is smaller than the original dimension of the matrix. For each dimension, the projection result may be quantized into 2 to the power of Ni buckets, where i∈[1,X], and the number of buckets is proportional to the singular value of the corresponding dimension. K=ΣNi bits may be used to represent the quantized projection results for all X dimensions. The bits concatenated together may be a cluster identifier for the device. In other implementations, a clustering algorithm (e.g., nearest neighbor) may be used to cluster devices in the reduced-dimensional space.

[0053] Because the singular value vectors are orthogonal to one another, and assuming that there is a nonlinear dependency between the profile vectors as a result of the quantization process, a statistically similar number of devices (e.g., approximately equal, assuming the total number of devices is large) may be within or identified as belonging to each cluster. Each cluster may be represented by an identifier, which may be referred to as a cluster identifier, browsing group identifier, or similar terminology.

[0054] In other implementations, other linear decomposition methods may be used, such as principal component analysis. In still other implementations, non-linear dimensionality reduction algorithms may be used to reduce the dimensionality of the matrix. Various classification techniques may be used, including nearest neighbor search, latent class analysis, etc.

[0055] 1C is a diagram of this process according to some implementations. As shown, in step 120, profile vectors from multiple devices may be aggregated into a large matrix. In step 122, the matrix may be dimensionally reduced. In step 124, clusters may be identified as discussed above.

[0056] In some implementations, a classification system may be trained as part of and / or from the cluster identification information. For example, in some implementations, a neural network may be used to classify devices as part of a predetermined number of clusters. Such a network may use the reduced-dimensionality profile vector as input and the cluster identifier as output. In various implementations, the network may be trained via supervised or unsupervised learning.

[0057] In some implementations, the neural net model or weights for the model may be provided to the client device, or other parameters for classification may be provided. The singular vectors generated from the dimensionality reduction may also be provided to the client device. Thus, after receiving the vectors and / or model, the clients may generate their own classifications using the local profile vectors without requiring further communication to the server. The server may regenerate the singular vectors and / or classification model parameters periodically, for example, monthly or quarterly. In some implementations, the data transfer may be very large (e.g., 2 26 2 in each of the dimensions 4 (This is on the order of singular vectors, which in some implementations would require approximately 2 GB of data.) In some implementations, to reduce data transfer to other devices, the server may compare the parameters and vectors with previously generated parameters and vectors and provide a new set only if there are significant differences (e.g., exceeding a threshold difference), or provide only a subset of parameters and / or vectors that have changed significantly. In various implementations, the client may use these parameters and vectors to locally update the classification more frequently, for example, daily, hourly, or with each content request.

[0058] Similarly, in some implementations, the singular vectors and / or model parameters may be provided to the content provider. When a client device requests an item of content, the request may include the cluster identification information. The cluster identification information may be embedded in the payload or header of the request, such as in an HTTPS request or in an optional field of an HTTP header. In some implementations, the content provider may use a neural net model or a provider-specific neural net model to infer attributes and / or user interests for devices in each cluster without being able to identify individual device or user characteristics (e.g., by determining approximate profile vectors corresponding to the cluster identification information based on the singular vectors of the dimensionality reduction, and then providing the vectors to a machine learning system to infer characteristics).

[0059] Thus, these implementations provide device anonymity through aggregation, i.e., aggregating devices with similar browsing histories or patterns together. The clustering algorithm attempts to maintain similar cluster sizes so that every cluster contains roughly the same number of users. Assuming a certain number of active devices on the Internet, the entropy of the cluster identifiers governs the cluster size (e.g., the higher the entropy of the cluster identifiers, the fewer devices there are in each cluster and the lower the privacy protection). By adjusting the entropy of the cluster identifiers (e.g., by providing fewer or more clusters), the system can achieve a desired level of anonymity and privacy protection while still maintaining the effectiveness of content personalization.

[0060] These implementations improve upon systems that do not utilize aggregation and pseudo-personalization while improving privacy and maintaining effectiveness. In such unimproved systems, a user or device identifier may be used to record a device's browsing history and infer the corresponding user's attribute information and interests based on the accumulated history. This inferred information may be used to predict the effectiveness of personalized content selection, such as click-through rate, attention, or other such measures. Instead, through the systems and methods discussed herein, browsing history may be accumulated only for a group of devices associated with a given cluster identifier or group. The inferred group's attributes and interests, along with the inferred effectiveness measure, can still be used for content selection, while content providers remain unable to distinguish between the characteristics of devices or users within a group or cluster.

[0061] In some implementations, 2K typical browsing history patterns are defined based on the aggregated browsing histories of billions of Internet users. Each typical browsing history pattern has a unique cluster identifier or browsing history identifier, which may be represented by a K-bit integer, where K is some small number, so that each cluster identifier is shared by multiple devices. When a user chooses to opt out of personalized content selection or opt in to pseudo-personalization, a browser application on the user's device may select a typical browsing history pattern that best matches the user's actual browsing history. The browser may provide the corresponding cluster identifier to a content provider for content personalization purposes.

[0062] Typical browsing history patterns and encoding of cluster identifiers are chosen so that an approximately equal number of devices are associated with each cluster identifier. By controlling the value of K and the entropy of other signals each content provider may obtain from the browser along with the content request (e.g., IP address, user agent identifier, etc.), the browser can significantly reduce the risk of user re-identification while enabling pseudo-personalization of content.

[0063] 2 is a block diagram of a system for anonymization to provide pseudo-personalized clustering according to some implementations. A client device 200, which may comprise a desktop computer, laptop computer, tablet computer, wearable computer, smartphone, embedded computer, smart car, or any other type and form of computing device, may communicate with one or more servers, such as a classifier server 230 and / or a content server 250, over a network 225.

[0064] In many implementations, the client device 200 may include a processor 202 and a memory device 206. The memory device 206 may store machine instructions that, when executed by the processor, cause the processor to perform one or more of the operations described herein. The processor 202 may include a microprocessor, an ASIC, an FPGA, or the like, or a combination thereof. In many implementations, the processor may be a multi-core processor or an array of processors. The memory device 206 may include, but is not limited to, an electrical, optical, magnetic, or any other storage device capable of providing program instructions to the processor. The memory device may include a floppy disk, a CD-ROM, a DVD, a magnetic disk, a memory chip, a ROM, a RAM, an EEPROM, an EPROM, a flash memory, an optical media, or any other suitable memory from which a processor can read instructions. The instructions may include code from any suitable computer programming language, such as, but not limited to, C, C++, C#, Java, JavaScript, Perl, HTML, XML, Python, and Visual Basic.

[0065] The client device 200 may include one or more network interfaces 204. The network interface 204 may include any type and format of interface, including Ethernet, including 10 Base T, 100 Base T, or 1000 Base T (“Gigabit”); any of the various 802.11 wireless standards, such as 802.11a, 802.11b, 802.11g, 802.11n, or 802.11ac; cellular, including CDMA, LTE, 3G, or 4G cellular; Bluetooth or other short-range wireless connections; or any combination of these or other interfaces for communicating with a network. In many implementations, the client device 200 may include multiple network interfaces 204 of different types, allowing connections to various networks 225. Correspondingly, network 225 may comprise a local area network (LAN), a wide area network (WAN) such as the Internet, a cellular network, a broadband network, a Bluetooth network, an 802.11 (WiFi) network, a satellite network, or any combination of these or other networks, and may include one or more additional devices (e.g., routers, switches, firewalls, hubs, network accelerators, caches, etc.).

[0066] A client device may include one or more user interface devices. A user interface device may be any electronic device (e.g., a keyboard, a mouse, a pointing device, a touchscreen display, a microphone, etc.) that conveys data to a user through sensory information (e.g., visualization on a display, one or more sounds, haptic feedback, etc.) and / or converts sensory information received from a user into electrical signals. According to various implementations, the one or more user interface devices may be internal to the housing of the client device, such as an internal display, a touchscreen, a microphone, or external to the housing of the client device, such as a monitor connected to the client device, speakers connected to the client device, etc.

[0067] The memory 206 may include an application 208 for execution by the processor 202. The application 208 may include any type and format of application, such as a media application, a web browser, a productivity application, or any other such application, and may be generally referred to herein as a browser application. The application 208 may receive content from a content server and display it via a user interface for a user of the client device.

[0068] The memory 206 may store an access log 210 (shown as log 210a for the client device 200), which may be part of or maintained by the application 208 (e.g., as part of a profile, preference file, history file, or other such file). The access log 210 may be stored in any format usable by the application 208. The access log may comprise identification of websites, domains, web pages, content, or other data accessed, retrieved, displayed, or otherwise obtained by the application 208. The access log 210 may also store a profile vector 100, as discussed above, which may be generated from the access history of the application and / or device. The profile vector 100 may comprise an n-dimensional string or array of values ​​representing accesses to one or more domains, web pages, websites, or other such data during a predetermined period (e.g., day, week, time period, etc.). As discussed above, the profile vector may be provided to the classifier server 230 (step A). The profile vector 100 may be generated by the application 208 or the log reducer 214, which may comprise an application, service, daemon, routine, plug-in, or other executable logic for generating a profile vector from an access log. In many implementations, the log reducer 214 may comprise part of the application 208.

[0069] The memory 206 may also store the singular vectors 212. As discussed above, the singular vectors 212 may be obtained from the classifier server 230 (step B), which may be calculated from the dimensionality reduction of the matrix of profile vectors of the multiple client devices 200 as discussed above. The singular vectors 212 may be stored in any suitable format, such as a flat file, a data array, or other structure, and in many implementations may be compressed.

[0070] The memory 206 may also store parameters of the neural net model 216. As discussed above, the neural net model 216 may be parameters or weights generated by the classifier server and provided to the client device 200 (step C). The classifier 218 of the client device 200, which may comprise an application, service, server, daemon, routine, or other executable logic for executing a machine learning algorithm, may utilize the parameters of the neural net model 216 to generate cluster identifiers 220 from the client device's reduced-dimensionality profile vector. In some implementations, the classifier 218 may comprise hardware circuitry, such as a tensor processing unit, or other such hardware. In other implementations, the classifier 218 may comprise software executed by the device's processor 202.

[0071] The memory 206 may also store cluster identifiers 220. The cluster identifiers 220 may comprise a cookie or other string associated with the cluster identifier and / or may encode or comprise information identifying characteristics of the cluster (e.g., XML code or parameter, parameter-value pairs, etc.). The cluster identifiers 220 may be predetermined or generated by the server 230 and provided to the client device 200. The client device's classifier 218 may process the client device's reduced-dimensionality profile vector using a neural net model to generate a rank or score for each cluster identifier 220 and may select the identifier with the highest rank or score for transmission to the content server during a content request (step D).

[0072] The classifier server 230 may comprise one or more server computing devices, one or more physical computing devices, or one or more virtual computing devices executed by one or more physical computing devices (e.g., a cloud, cluster, or server farm). The classifier server 230 may be generally referred to as a server, a measurement server, an aggregation server, or other such terminology.

[0073] The classifier server 230 may comprise one or more processors 202, a network interface 204, and a memory device 206, as well as other devices not shown. The classifier server 230 may store in memory the access logs and / or profile vectors 210a-210n obtained from the multiple client devices 200. As discussed above, the aggregator 232 of the classifier server 230, which may comprise an application, server, service, routine, or other executable logic executed by the processor 202, may aggregate the profile vectors 210a-210n into a matrix or an n-dimensional array. The aggregator 232 may also compute a matrix decomposition or dimensionality reduction into singular vectors 212, which may be provided to the client devices 200 (and, in some implementations, the content server 250).

[0074] The classifier server 230 may also store the classifier 218 in memory. The classifier 218 may be part of the aggregator 232 or may be a separate application, service, server, routine, or other executable logic executed by the processor 202 and / or a hardware processor such as a tensor processing unit for identifying clusters in the dimensionality-reduced matrix. In many implementations, the classifier 218 may comprise a neural network or similar artificial intelligence for classifying the dimensionality-reduced vector into one of multiple clusters. Once trained, the parameters of the neural network may be provided to the client device 200 to similarly generate cluster identification information or rankings as discussed above.

[0075] The content server 250 may comprise any type and form of content-providing server or service, including a content delivery network, a web server, a media server, a social media service, or any other type and form of computing system. The content server 250 may comprise one or more processors 202, a network interface 204, and a memory device 206. In many implementations, the content server 250 may store multiple content items 252, which may comprise any type and form of content, including text, audio, video, animation, images, executable scripts, web pages, or any other such data.

[0076] The content server 250 may include a content selector 254, which may be part of a web server or file server, or may be a separate application, service, server, daemon, routine, or other executable logic for selecting content to deliver to client devices. The content may be selected based on inferred characteristics of the device. The content server may receive a cluster identifier along with a request for content from a device and may select content based on inferred characteristics for the devices in that cluster. In some implementations, the content server may use the singular vectors obtained from the classifier server 230 to expand the cluster identifier into a corresponding profile vector representing the aggregate cluster. The profile vector of a cluster may not match the profile vector of any particular device, but may be an approximation or average of the vectors of all devices associated with the cluster.

[0077] 3 is a flowchart of a method for anonymization to provide pseudo-personalized clustering, according to some implementations. In step 302, client device 200 may provide an access log or a profile vector generated from the access log to classifier server 230. As discussed above, the profile vector may be based on the browsing or access history of the client device and may comprise an n-dimensional vector or string thereof, with a value representing each of one or more accesses to an address associated with a corresponding position in the string or array within a predetermined time period.

[0078] In step 304, the classifier server 230 may aggregate the profile vectors or logs from the client devices 200 into a matrix with profile vectors or logs obtained from one or more additional client devices 200. The profile vectors may be obtained by the classifier server 230 periodically or upon login to a service provided by the classifier server or an associated server. Steps 302-304 may be repeated for multiple client devices, which may be a small fraction of all devices that can utilize the singular vectors in 312 to perform dimensionality reduction in 314. In some implementations, steps 302-312 may be replaced by federated learning on the client devices, and the classifier server 230 may be optional or eliminated.

[0079] In step 306, the classifier server 230 may compute a dimensionality reduction or decomposition of the matrix. In some implementations, the classifier server may use a singular value decomposition algorithm and may generate multiple singular vectors and a dimensionality reduced matrix.

[0080] The classifier server may optionally identify cluster boundaries or parameters for clusters in the dimensionality reduced matrix in step 308. The classifier server may use any suitable algorithm, such as principal component analysis, or a machine learning system, such as a neural network, to identify the clusters.

[0081] In some implementations, a classifier model may be tuned or trained based on the identified clusters and the dimension-reduced profile vectors in step 310. In various implementations, the training may be supervised or unsupervised.

[0082] In step 312, the singular vectors, and in some implementations, the weights of the neural net model or other learning algorithm, may be provided to the client device 200, and in some implementations, to one or more content servers 250.

[0083] At step 314, the client device 200 may use the singular vectors received from the classifier server to calculate a dimensionality reduction of the profile vector or the device's access log. The dimensionality reduced vector may be classified via artificial intelligence or a neural network at step 316 using the model parameters received from the classifier server and using the classification determined at step 318. Determining the classification may comprise calculating a score or rank for each of a plurality of clusters (e.g., in some implementations, identified by the classifier server and provided via parameters) for the device's profile vector.

[0084] In step 320, the client device may send a request for an item of content to one or more content servers 250. The request may include an identification of a cluster corresponding to the device's profile vector. The request may be generated in response to the execution of a script on a web page, after completion of playback of an item of media or a portion of an item of media, or any other such situation.

[0085] In step 322, the content server may determine characteristics of the cluster based on the classifier model parameters and the singular vectors received from the classifier server. In some implementations, the content server may use the singular vectors to calculate a profile vector representing the aggregate browsing history of all devices in the cluster, and then infer characteristics of the cluster based on aspects of the history (e.g., keywords / topics associated with web pages or domains or other content, other related domains or web pages, etc.). In step 324, the content server may use the cluster identification information in the request (or inferred characteristics associated with the cluster as discussed above) to select an item of content. In step 326, the content may be sent to the client device, and in step 328, the client device may render or display the content item.

[0086] Thus, the systems and methods discussed herein enable pseudo-personalization of content or anonymized content retrieval via the aggregated browsing history of a large number of devices, such as millions or billions of devices. A sparse matrix can be constructed from the aggregated browsing history and dimensionally reduced, reducing entropy and providing anonymity for individual devices. Without exposing individual device details to content providers, relevant content can be selected via pseudo-personalized clusters representing similar browsing histories.

[0087] In a networked environment, an application (e.g., a web browser) executing on a client device can receive an information resource (e.g., a web page) with primary content provided by a content publisher and one or more content slots (e.g., inline frames) for supplemental content. The information resource can include code snippets or scripts (e.g., content selection tags) that specify retrieval of content items from a content provider from a content selection service to be inserted into the content slots. Upon parsing the script of the information resource, the application can generate a request for content to be inserted into the content slots and can send the request to the content selection service. In response to receiving the request, the content selection service can select one of the content items supplied by the content provider to be embedded into the content slots of the information resource.

[0088] Selection of content items by the content selection service may depend on the use of a deterministic tracking identifier unique to the user (or via an account), the client device operated by the user, or applications running on the client device. The identifier may be maintained on the client device and accessed by the content selection service via a cookie. The cookie may be, for example, a third-party cookie associated with a domain of the content selection service that is different from the domain of the content publisher for the information resource. When a content slot of an information resource specifies retrieval of content from the content selection service, a cookie containing the identifier may be passed from the client device to the content selection service. Using the cookie, the content selection service can track which information resources are accessed by the user through applications running on the client device. Additionally, the content selection service can identify content items determined to be relevant to the user operating the application on the client device based on the tracked information resources.

[0089] While the use of these unique tracking identifiers enables the selection of content items that are customized for a particular user, it can have a number of drawbacks, particularly with regard to data security and integrity. For one, users may be exposed to data security risks regarding user data exchanged between client devices and the content selection service. For example, administrators of the content selection service may intentionally provide private user data to third parties without the user's consent. In addition, unauthorized, malicious entities may intercept third-party cookies in transit and use the unique tracking identifiers to monitor the online activity of affected client devices and applications. For another, the aggregation of information resources accessed using such identifiers may pose a risk of data leakage on the part of the content selection service. For example, the accidental disclosure of data, some of which can be individually traced back to a particular user, or malicious attacks attempting to exfiltrate such collected data, may lead to a loss of user data privacy.

[0090] One approach to addressing issues with unique tracking identifiers may include disabling third-party cookies with unique tracking identifiers from client devices. Applications running on client devices may be configured to prohibit the generation, maintenance, or transmission of unique tracking identifiers to a content selection service. However, restricting third-party cookies may cause countless other problems. Disabling third-party cookies may prevent the content selection service from tracking which information resources are accessed by the client device through the application. Thus, when a request for content is received, the content selection service may not be able to use such information in determining the relevance of the content item to a user operating the application. As a result, the selected content item may be less likely to be interactive with the user of the client device than a content item selected using a tracking identifier. As a result, the information resource on which the content item is inserted for display may suffer from a degradation in the quality of human-computer interaction (HCI).

[0091] To address the technical challenges of prohibiting the use of unique identifiers to track individual client devices or applications in selecting content, each application can categorize itself into one of a number of clusters based on locally maintained browsing history. Applications with similar browsing patterns, and by extension, users who operate the applications, can be categorized into the same cluster. Users with similar browsing patterns and categorized into the same cluster can be correlated as having similar traits and interests and therefore may be more likely to have similar responses to the same content items. Because each cluster can have a large number of associated users (e.g., greater than 1000), the categorization of users into clusters may not be a characteristic unique to individual users.

[0092] When assigning itself to one of the clusters, an application can convert or encode the browsing history into a vector with a preset number of dimensions. For example, one element in a feature vector can indicate whether a user visited a particular domain, a section of a website, a web page in a particular category (e.g., vacation), or even a particular URL in a particular time slot (e.g., a particular time of day and a particular day of the week). The application can then apply a clustering or classification algorithm (e.g., k-nearest neighbor algorithm, linear classification, support vector machine, and pattern recognition) to the feature vector to identify the cluster to which the application, and by extension, the user, should be assigned. Clustering algorithms can be provided and updated for an application by an application manager (e.g., a browser vendor).

[0093] Once the clusters are discovered using a clustering algorithm, the application can identify a cluster identifier (also called a class identifier or browsing history identifier) ​​for the cluster. The cluster identifier may be assigned to each cluster by the application manager and provided to the application and content selection service. In contrast to a unique deterministic tracking identifier, a cluster identifier may not be specific to one individual user, application, or client device. Multiple users may be categorized into the same cluster, and a cluster identifier may also be common to multiple users, applications, or client devices with similar browsing patterns. Because a cluster identifier is shared by multiple users, the cluster identifier may have lower entropy than a unique tracking identifier assigned to an individual user. For example, a unique deterministic tracking identifier for all Internet users may have an entropy of over 30 bits, while a cluster identifier may be set to an entropy of 18 to 22 bits. The lower entropy may allow the cluster identifier itself to be shorter and smaller in size than a unique tracking identifier.

[0094] When an information resource with a content slot is received, the application can include the cluster identifier in a request for content for the content slot, instead of the unique tracking identifier, and send the request to the content selection service. Additional protection measures can be employed to improve data security and integrity and reduce the probability of theft of the cluster identifier and associated information. To protect against interception of the cluster identifier during transmission over the network, the application can use an encryption protocol, such as the Hypertext Transfer Protocol Secure (HTTPS) protocol. Additionally, to protect the cluster identifier maintained on the client device, the application can restrict other client-side processes (e.g., JavaScript on the information resource) from accessing the cluster identifier. For example, the cluster identifier can be included in a secure flag or HTTP-only flag cookie maintained on the client device to prevent access to the cluster identifier. This can be in contrast to a third-party cookie, which does not have such access control rights.

[0095] In response to receiving the request, the content selection service can select one of the content items using the cluster identifier. The content selection service can accumulate browsing history for users categorized as the cluster using previous requests for content that includes the cluster's cluster identifier. By applying a profiling model to the accumulated history for the cluster, the content selection service can infer the traits and interests of users within the cluster. The results of the profiling model enable the content selection service to discover content items determined to be relevant to the cluster into which the user associated with the request is categorized.

[0096] By using the cluster identifier, the browsing history of an application maintained on a client device may be prevented from being accessed by the content selection service such that the browsing history could be traced back to a particular user, application, or client device. Additionally, the content selection service may be unable to track individual users, applications, or client devices across different domains to assemble a detailed browsing history. Instead, the content selection service may aggregate browsing history for a particular cluster of users associated with the cluster identifier received from the application. When aggregating, the content selection service can protect the data privacy of individual users by merging browser histories from different users in the same cluster. The degree of data privacy may also be controlled by setting the number of users to be assigned to each cluster.

[0097] Furthermore, under the assumption that users in the same cluster have similar browsing patterns, the content selection service can infer and determine similar traits and interests for users of the same cluster based on the browsing history aggregated for the cluster. The content selection service can further select content items with the expectation that users of the same cluster will respond similarly. In this way, data security, integrity, and privacy for individual users' browsing histories can be maintained. At the same time, being able to determine relevance in content item selection can maintain the quality of human-computer interaction (HCI) for the selected content items or the entire information resource.

[0098] Referring now to FIG. 4, a block diagram illustrates one implementation of a computer networked environment or system 400 for encoding identifiers for content selection using a classification model. Broadly, the system 400 may include at least one network 405 for communication between components of the system 400. The system 400 may include at least one application manager service 410 (also referred to herein as a browser vendor) to provide services for at least one application (e.g., a browser). The system 400 may include at least one content provider 415 to provide content items. The system 400 may include at least one content publisher 420 to provide information resources (e.g., web pages). The system 400 may include at least one content selection service 425 to select content items. The system 400 may include one or more client devices 430A-N (generally referred to herein as client device 430). Each client device 430 may include at least one application 435A-N (generally referred to herein as application 435). Each of the components of system 400 (e.g., network 405, application manager service 410 and its components, content provider 415 and its components, content publisher 420 and its components, content selection service 425 and its components, and client device 430 and its components) may be implemented using components of computing system 900 detailed herein in connection with FIG. 9 .

[0099] In further detail, the network 405 of the system 400 can communicatively couple the application manager service 410, the content provider 415, the content publisher 420, the content selection service 425, and the client device 430 to one another. The application manager service 410, the content provider 415, the content publisher 420, and the content selection service 425 of the system 400 can each include multiple servers located in at least one data center or server farm that are communicatively coupled to one another via the network 405. The application manager service 410 can communicate with the content provider 415, the content publisher 420, the content selection service 425, and the client device 430 via the network connection 405. The content provider 415 can communicate with the application manager service 410, the content publisher 420, the content selection service 425, and the client device 430 via the network 405. The content publisher 420 can communicate with the application manager service 410, the content publishers 420, the content selection service 425, and the client devices 430 over the network 405. The content selection service 425 can communicate with the application manager service 410, the content providers 415, the content publishers 420, and the client devices 430 over the network 405. Each client device 430 can communicate with the application manager service 410, the content providers 415, the content publishers 420, and the content selection service 425 over the network 405.

[0100] The application manager service 410 may include a server or other computing device operated by an application vendor (sometimes referred to herein as a browser vendor) to provide resources and updates to applications 435 executing on client devices 430. For example, the application manager service 410 may provide applications 435 for installation on client devices 430. The application manager service 410 may also provide updates to applications 435 installed on client devices 430. The updates may affect at least one of the subcomponents of the applications 435. The application manager service 410 may also provide plug-ins or add-ons to applications 435 executing on client devices 430 to enhance the functionality of the applications 435. The application manager service 410 may communicate with a content selection service 425 to provide information on the applications 435 executing on client devices 430. The preparation of applications 435 and associated files or data may be communicated by the application manager service 410 over the network 405.

[0101] The content provider 415 may include a server or other computing device operated by a content provider entity to provide content items for display on an information resource at the client device 430. The content provided by the content provider 415 may take any convenient form. For example, third-party content may include content related to other displayed content, such as a website page related to the displayed content. The content may include third-party content items or creations (e.g., advertisements) for display on an information resource, such as an information resource containing primary content provided by the content publisher 420. Content items may also be displayed on search result web pages. For example, the content provider 415 may provide or be the source of content items 455 for display in a content slot (e.g., an iframe element) of the information resource 450, such as a company's web page if the primary content of the web page is provided by the company, or for display on a search result landing page provided by a search engine. Content items associated with the content provider 415 may be displayed on information resources other than web pages, such as content displayed as part of the execution of an application on a smartphone or other client device 430.

[0102] Content publisher 420 may include a server or other computing device operated by a content publishing entity to provide information resources containing primary content for display over network 405. For example, content publisher 420 may include a web page operator that provides primary content for display on an information resource. The information resource may include content other than that provided by content publisher 420, and the information resource may include content slots configured for display of content items from content provider 415. For example, content publisher 420 may operate a company's website and provide content about the company for display on a web page of the website. The web page may include content slots configured for display of content items provided by content provider 415 or by content publisher 420 itself. In some implementations, content publisher 420 includes a search engine computing device (e.g., a server) of a search engine operator that operates a search engine website. The primary content of a search engine web page (e.g., landing web page results) may include search results displayed in content slots of information resources, such as content items from content providers 415, as well as third-party content items. In some implementations, content publishers 420 may include one or more servers for providing video content.

[0103] The content selection service 425 may include a server or other computing device operated by a content placement entity to select or identify content items to be inserted into content slots of an information resource over the network 405. In some implementations, the content selection service 425 may include a content placement system (e.g., an online advertising server). The content selection service 425 may maintain an inventory of content items from which to select content items to provide over the network 405 for insertion into content slots of an information resource. The inventory may be maintained on a database accessible to the content selection service 425. The content items or identifiers (e.g., addresses) of the content items may be provided by the content provider 415.

[0104] Each client device 430 may be a computing device for communicating over the network 405 to display data. The displayed data may include content provided by content publishers 420 (e.g., information resources) and content provided by content providers 415 (e.g., content items for display in content slots of information resources) as identified by the content selection service 425. The client devices 430 may include desktop computers, laptop computers, tablet computers, smartphones, personal digital assistants, mobile devices, consumer computing devices, servers, clients, digital video recorders, television set-top boxes, video game consoles, or any other computing devices configured to communicate over the network 405. The client devices 430 may be communications devices through which end users can issue requests to receive content. The request may be a request to a search engine, and the request may include a search query. In some implementations, the request may include a request to access a web page.

[0105] The applications 435 executing on the client device 430 may include, for example, an internet browser, a mobile application, or any other computer program capable of executing or otherwise invoking computer-executable instructions provided to the client device 430, such as computer-executable instructions contained in information resources and content items. The contained information resources may correspond to script, logic, markup, or instructions (e.g., Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), Extensible Markup Language (XML), Cascading Style Sheets (CSS), and JavaScript, or any combination thereof). Content items may be inserted into content slots of information resources.

[0106] 5, a block diagram is shown illustrating one implementation of a client device 430 and an application management service 410 in a system 400. Broadly, the application manager service 410 may include at least one classification model 500 for categorizing applications 435 based on browsing patterns. The application manager service 410 may include at least one model trainer 505 for training and maintaining the at least one classification model 500. The application manager service 410 may include at least one model updater 510 for modifying or adjusting the classification model 500. The application manager service 410 may include at least one database 515 for storing and maintaining a set of class identifiers 520A-N (generally referred to herein as class identifiers 520, and sometimes referred to herein as "web zone identifiers" or "web zip codes"). The application manager server 410 may include at least one instance of a class authorizer 550 (sometimes referred to herein as a class authorization service) for authorizing the inclusion of a class identifier 520 in requests transmitted over the network 405.

[0107] On each client device 430, the application 435 may include at least one classification model 500 to categorize the application 435 based on browsing patterns. The application 435 may include at least one model trainer 505 to train and maintain the classification model 500. The application 435 may include at least one content extractor 525 to select features from information resources accessed via the application 435. The application 435 may include at least one model applicator 530 to categorize the application 435 using the classification model 500. The application 435 may also include at least one instance of a class authorizer 550. The model trainer 505, content extractor 525, model applicator 530, and class authorizer 550 may be integral parts of the application 435, add-ons or plug-ins to the application 435, or separate applications that interface with the application 435. The application 435 may include at least one browsing history 535 for maintaining and storing one or more entries 540A-N (generically referred to herein as entries 540). The application 435 may include at least one identifier cache 545 for maintaining and storing at least one class identifier 520 for the application 435.

[0108] In further detail, the classification model 500 can classify, group, or otherwise categorize each application 435 (or each client device 430 executing the application 435 or account associated with the application 435) based on the browsing history 535. The classification of the application 435 on the client device 430 can indicate or represent the classification of the instance of the application 435 executing on the client device 430. For example, there may be one instance of the application 435 (e.g., a web browser) executing on one client device 430 and another instance of the application 435 (e.g., the same type of web browser) executing on another client device 430. Both instances may be classified into the same category or different categories. By extension, the classification of the application 435 may correspond to or include the classification of the user associated with the account operating the application 435 or the client device 430 operated by the user.

[0109] The classification model 500 may be a classification or clustering model or algorithm. The classification model 500 may include classification algorithms such as linear classifiers (e.g., linear regression, logistic regression, and naive Bayes classifier), support vector machines (SVMs), quadratic classifiers, k-nearest neighbor algorithms, and artificial neural networks (ANNs), among others. The classification model 500 may include clustering models such as geometric center-based clustering algorithms (e.g., k-means algorithm or expectation-maximization (EM) algorithm), density-based clustering algorithms (e.g., density-based spatial clustering of applications with noise (DBSCAN)), distribution-based clustering algorithms (e.g., Gaussian mixture models), and artificial neural networks (ANNs), among others. The classification model 500 may also include hash functions such as locality-sensitive hashing schemes (e.g., MinHash, SimHash, and Nilsimsa Hash), among others.

[0110] The classification model 500 may include a set of inputs, a set of parameters, and a set of outputs according to either a classification or clustering model and algorithm. The set of inputs may incorporate or include browsing history 535 entries 540. In some implementations, the set of inputs may incorporate or include a dimensionally reduced representation of browsing history 535 entries 540. In some implementations, the set of inputs may include a complete representation of browsing history 535 entries 540. A set of parameters (or weights) may connect or associate the set of inputs with the set of outputs. In some implementations, the set of parameters may include the number of classes and a value for each class. For example, the set of parameters may be a geometric center in k-means clustering for each class. In another example, the set of parameters may include a range of values ​​associated with each class. The number of classes may be equal to the number of class identifiers 520. The number of classes and the number of class identifiers 520 may be predetermined (e.g., to a fixed value) or may be dynamically determined. The set of outputs may create or include a class corresponding to one of the class identifiers 520. The set of outputs may include one of the class identifiers 520 itself. For example, the class identifier 520 may be a hash value calculated using a hash function. Each class identifier 520 may be or include a set of alphanumeric or numeric values ​​(e.g., integers or binary numbers).

[0111] A model trainer 505 executing on the application manager service 410 can train the classification model 500. The model trainer 505 can use a training dataset to train the classification model 500. Training the classification model 500 can follow unsupervised learning techniques. The training dataset can include sample browsing histories 530 from a sample set of applications 435 executing on a sample set of client devices 430. In some implementations, the model trainer 505 can obtain and accumulate the sample browsing histories 530 from a content provider 415, a content publisher 420, a content selection service 425, or an application 435 executing on the client device 430. Prior to training, the model trainer 505 can identify the number of classes for the classification model 500. In some embodiments, the number of classes can be predetermined or assigned by an administrator of the application manager service 410. In some implementations, the number of classes can be determined by the model trainer 505 based on the number of client devices 430 with the applications 435. For example, the number of classes may be set so that a set number of client devices 430 (eg, 800 to 1000 devices) are assigned to each class.

[0112] When training the classification model 500, the model trainer 505 can use the training dataset to change, adjust, or otherwise set the values ​​of the parameters in the classification model 500 (e.g., the values ​​of each class). At each iteration, the model trainer 505 can determine whether the classification model 500 has converged compared to the previous iteration based on the change in the set of values ​​of the parameters. In some implementations, the model trainer 505 can compare the change in the set of values ​​in the parameters of the classification model 500 to a convergence threshold. If the change is less than the convergence threshold, the model trainer 505 can determine that the classification model 500 has converged. Otherwise, if the change is greater than the convergence threshold, the model trainer 505 can determine that the classification model 500 has not converged. When it is determined that the classification model 500 has not converged, the model trainer 505 can continue training the classification model 500.

[0113] Alternatively, when it is determined that the classification model 500 has converged, the model trainer 505 can finish, terminate, or stop training of the classification model 500. The model trainer 505 can store the classification model 500 in the application manager service 410. Additionally, the model trainer 505 can transmit or send the classification model 500 to applications 435 executing on client devices 430. In some implementations, the model trainer 505 can transmit or send a set of parameters for the classification model 500. For each class in the classification model 500, the model trainer 505 can identify the class and assign or otherwise associate it with a corresponding class identifier 520. The class identifier 520 can be a set of alphanumeric characters for referencing the class. The classification model 500 can associate each class with a corresponding class identifier 520. The model trainer 505 can transmit and provide the set of class identifiers 520 to the applications 435 executing on each client device 430 and to the content selection service 425.

[0114] A model trainer 505 of an application 435 executing on a client device 430 can also train the classification model 500. In some implementations, the model trainer 505 can use a distributed learning protocol to train the classification model 500. The distributed learning protocol may cooperate with applications 435 executing on other client devices 430 and with the application manager service 410 communicating over the network 405. The distributed learning protocol may include, for example, federated learning using an optimization algorithm (e.g., stochastic gradient descent (SGD) or averaging) to train the classification model 500. The number of classes and the number of class identifiers 520 may be predetermined (e.g., to a fixed value) or dynamically determined, as described above. During each iteration, each model trainer 505 can use the training dataset to change, adjust, or otherwise set parameter values ​​(e.g., the value of each class) in the classification model 500. At the end of an iteration, each model trainer 505 can provide each other (the model trainer 505 instances on the other client devices 430) with values ​​of the parameters (e.g., values ​​for each class) in the classification model 500. The model trainer 505 can use the received values ​​of the parameters to adjust, change, or set parameters of the locally maintained classification model 500. The model trainer 505 can repeat the iterations until it determines that the classification model 500 has converged compared to the previous iteration based on changes in the set of parameter values ​​as discussed above.

[0115] A model updater 510 executing on the application manager service 410 can provide or send updates to the classification model 500 maintained on each client device 430 running an application 435. The model updater 510 can determine whether to update the classification model 500 according to a model update policy. The model update policy can specify a set of conditions under which the classification model 500 should be updated. In some implementations, the model update policy can include a schedule indicating when the classification model 500 should be updated. The model update policy can specify that the classification model 500 should be updated when the number of applications 435 assigned to each class is not evenly distributed (e.g., the difference in class size is between 5% and 100% of one class). The model update policy can specify that the classification model 500 should be updated when the amount of accumulated additional browsing history 535 meets a threshold amount. If the determination is that the classification model 500 should not be updated, the model updater 510 can maintain the classification model 500.

[0116] On the other hand, if the decision is to update, the model updater 510 can invoke a model trainer 505 (on the application manager service 410 or across applications 435 running on the client device 430) to retrain the classification model 500. In some implementations, the model updater 510 can accumulate browsing history 530 for training datasets from content providers 415, content publishers 420, content selection services 425, or applications 435 running on the client device 430. The model updater 510 can pass the accumulated browsing history 530 to the model trainer 505 to retrain the classification model 500. Upon determining that the classification model 500 has converged, the model trainer 505 can finish, terminate, or stop training of the classification model 500. The model updater 510 can transmit or send the newly trained classification model 500 (or a set of parameters for the classification model 500) to each application 435 to update the classification model 500. The model updater 510 can transmit and provide the set of class identifiers 520 to the applications 435 running on each client device 430 and to the content selection service 425.

[0117] A model applicator 530 of an application 435 executing on a client device 430 can receive the classification model 500 from the application manager service 410 over the network 405. Upon receipt, the model applicator 530 can store and maintain the classification model 500 on the client device 430. In some implementations, the model applicator 530 can receive a set of parameters for the classification model 500 from the application manager service 410. Receiving the set of parameters can be for updating the classification model 500. Upon receipt, the model applicator 530 can change, configure, or otherwise modify the classification model 500 using the received set of parameters. Additionally, the model applicator 530 can receive a set of class identifiers 520 for classes in the classification model 500 from the application manager service 410. Receiving the class identifiers 520 can be part of updating the classification model 500.

[0118] The configuration of the classification model 500 enables the model applicator 530 to identify a browsing history 535 maintained on the client device 430 by the application 435. The application 435 may maintain the browsing history 535 by creating an entry 540 each time an information resource is accessed. The browsing history 535 may record information resources (e.g., web pages) and other online content accessed through the application 435. The browsing history 535 may include a set of entries 540. Each entry 540 may include the address of the accessed information resource (e.g., a Uniform Resource Locator (URL) including the hostname and pathname of the web page) and a timestamp indicating the time the information resource was accessed. The set of entries 540 may be indexed by the timestamp or address of the information resource within the browsing history 535. In some implementations, the model applicator 530 may identify a portion of the browsing history 535 spanning a range of time to use to facilitate processing by the model applicator 530. That portion of browsing history 535 may include a subset of entries 540 with timestamps within that time range. The time range may be, for example, a week's worth of entries 540 from the current time.

[0119] In conjunction therewith, a content extractor 525 of an application 435 executing on the client device 430 can identify, select, or otherwise extract one or more features from each information resource for a corresponding entry 540 in the browsing history 535. The features may be extracted from at least a portion of the content on the information resource. The features may include, for example, textual data, visual data, or audio data. In some implementations, the content extractor 525 can identify the features from the content of the information resource because the application 435 accesses the information resource via the network 405. In some implementations, the content extractor 525 can access each information resource via a corresponding entry 540 in the browsing history 535 to extract the features. This access may be separate from or subsequent to the presentation of the information resource via the application 435.

[0120] When extracting from information resources, the content extractor 525 can apply one or more attribute selection algorithms to the content of each information resource accessed by the application 535. To extract text data, the content extractor 525 can apply at least one natural language processing algorithm, such as term extraction, named entity recognition, relationship extraction, automatic summarization, and word frequency-inverse document frequency (TF-IDF), among others. The text data identified using a natural language processing algorithm may include a subset of the text content on the information resource. To extract visual data, the content extractor 525 can apply at least one computer vision algorithm, such as object recognition and optical character recognition (OCR). The visual data identified using a computer vision algorithm may include a subset (or identifiers corresponding to the subset) of graphics on the information resource. To extract audio data, the content extractor 525 can apply at least one audio signal analysis algorithm and speech recognition algorithm. The extracted audio data may include, for example, words recognized from the audio, among others.

[0121] Upon identifying the browsing history 535, the model applicator 530 can form or generate a set of feature vectors using the entries 540 of the browsing history 535. In some implementations, the model applicator 530 can generate the set of feature vectors using features extracted from information resources accessed by the application 435. The set of feature vectors can be used as input for the classification model 500. The set of feature vectors can include or be defined by dimensions. The dimensions can include time ranges (e.g., time of day or day of the week) and address attributes (e.g., domain name, website section, topic category, or the address itself). The dimensions can also include text data, image data, and audio data corresponding to the extracted features. Each feature vector can be associated with at least one of the time ranges based on a timestamp associated with the corresponding entry 540. Each feature vector can be associated with at least one of the address attributes for the information resource based on the address for the information resource identified in the corresponding entry 540.

[0122] In some implementations, the model applicator 530 may generate a set of feature vectors by projecting entries 540 of the browsing history 535 onto feature vector dimensions defined by the time range and address attributes. For an entry 540 of the browsing history 535, the model applicator 530 may identify whether an existing feature vector exists based on the address and timestamp in the entry 540. To identify, the model applicator 530 may determine whether the entry 540 matches any of the existing feature vectors. When an existing feature vector exists, the model applicator 530 may add the entry 540 to the feature vector. Conversely, when no existing feature vector exists, the model applicator 530 may create a new feature vector for the entry 540.

[0123] In some implementations, the model applicator 530 can use a dimensionality reduction process to form or generate a reduced set of feature vectors. The dimensionality reduction process can include, among other things, linear reduction techniques (e.g., principal component analysis (PCA), singular value decomposition (SVD), non-negative matrix factorization (NMF)), non-linear dimensionality reduction (e.g., generalized discriminant analysis (GDA), locally-linear embedding, and Sammon mapping), or others (e.g., Johnson-Lindenstrauss lemma and multifactor dimensionality reduction). In some implementations, the model applicator 530 can apply a dimensionality reduction process when projecting the entries 540 of the browsing history 535 onto the dimensions of the feature vectors. In some implementations, the model applicator 530 can apply a dimensionality reduction process to the generated set of feature vectors. By applying a dimensionality reduction process, the model applicator 530 can reduce the number of dimensions in the original set of feature vectors to generate a reduced set of feature vectors. The reduced set of feature vectors may have fewer dimensions and data points than the initial set of reduced feature vectors. For example, the reduced set of feature vectors may omit time range or address attributes that do not have any associated entries 540. In some implementations, the model applicator 530 may omit the dimensionality reduction process and continue processing the feature vectors without the dimensionality reduction process.

[0124] The model applicator 530 can apply the classification model 500 to the browsing history 535 (or any subset or representation of the entries 540 of the browsing history 535, such as feature vectors) to identify one or more classes into which the application 435 should be categorized. To do so, the model applicator 530 can provide the browsing history 535 as a set of inputs to the classification model 500. In some implementations, the model applicator 530 can provide a feature vector or a reduced set of feature vectors as inputs to the classification model 500. Once provided, the model applicator 530 can use the classification model 500 to compare the inputs to parameters that define the classes and generate or create a set of outputs based on the comparison. The set of outputs may include one or more classes into which the browsing patterns as indicated in the browsing history 535 maintained by the application 435 should be categorized.

[0125] From the output of the classification model 500, the model applicator 530 can identify one or more classes. In some implementations, the model applicator 530 can identify a single class (sometimes referred to herein as the closest or closest class) from the output. This single class may correspond to the portion of the feature space defined by the classification model 500 that is closest in distance to the input feature vector. The feature space defined by the classification model 300 may have the same parameters and their values ​​as the input.

[0126] In some implementations, the model applicator 530 can identify a set of classes (sometimes referred to herein as closest or most closely related classes) from the output. The identified set can be a subset of the classes defined by the classification model 500. The member classes of the set can be within a proximity threshold of each other in the feature space defined by the classification model 300. The proximity threshold can define a distance in the feature space within which one or more classes should be selected. In some implementations, the model applicator 530 can identify a set of classes that are within a proximity threshold of the feature vector input to the classification model 500. In some implementations, the model applicator 530 can rank the set of classes by their distance from the input feature vector in the feature space.

[0127] Each identified class may correspond to one of a number of classes as defined by classification model 500. The identified class for an application 435 executing on a client device 430 may be common, shared, or identical to the identified classes for at least some other applications 435 executing on other client devices 430. As discussed above, each class defined by classification model 500 may have a certain number of client devices to be categorized into the class.

[0128] Upon identifying one or more classes for an application 435, the class authorizer 550 can determine whether each class meets (e.g., meets or exceeds) a threshold number of applications 435 assigned to the class. The threshold number can stipulate or define the number of applications 435 assigned to a class that can be used to generate requests for content. The threshold number can be set to match or achieve a target entropy. There can be multiple applications 435 assigned to the same class by each individual instance of the model applicator 530. However, the class cannot be used to generate requests for content until the number of such applications 435 exceeds the threshold number. Determination by the application 435 on the client device 435 may be coordinated with or in coordination with a class authorization service, such as the class authorizer 550 running on the application manager service 410. The functionality of the class authorizer 550 can be divided between the application 435 and the application manager service 410 (or some other server device). The determination can be according to a threshold encryption protocol or a check against a probabilistic data structure.

[0129] In some implementations, the class authorizer 550 of the application 435 and the class authorizer 550 of the application manager service 410 can perform a threshold encryption protocol when making their decision. For each identified class, the class authorizer 550 on the application 435 can identify a class identifier 520 corresponding to the class. Upon identification, the class authorizer 550 can generate an encrypted copy of the class identifier 520. In some implementations, the class authorizer 550 can generate the encrypted copy using at least a portion of a private encryption key according to an asymmetric encryption algorithm. The asymmetric encryption algorithm may include, for example, digital signature, Diffie-Hellman key exchange, elliptic curve cryptography, or the Rivest-Shamir-Adleman (RSA) algorithm, among others. The class authorizer 550 of the application 435 can generate an authentication request that includes the encrypted copy of the class identifier 520. Once generated, the class authorizer 550 of the application 435 can send the request over the network 405 to the class authorizer 550 of the application manager service 410.

[0130] Upon receiving from the client device 430, the class authorizer 550 executing on the application manager service 410 can analyze the authentication request to identify an encrypted copy of the class identifier 520. This identification allows the class authorizer 550 to attempt to decrypt the original class identifier 520 from the encrypted copy of the class identifier 520. The decryption may be according to an asymmetric encryption algorithm. Under a threshold encryption protocol, the class authorizer 550 may not be able to successfully decrypt the encrypted copy until the number of received authentication requests with encrypted copies of the same class identifier 520 meets a threshold number. The inability to decrypt may be because, for example, the class authorizer 550 may not have received a sufficient portion of the request's private key from a different application 435. Once a sufficient portion of the private key is received, the class authorizer 550 may be able to successfully decrypt the encrypted copy to recover the original class identifier 520.

[0131] The class authorizer 550 may generate an output from the decryption of the encrypted copy. If the number of requests with encrypted copies for the same class identifier 520 does not meet (e.g., is less than) a threshold, the class authorizer 550 may be unable to recover the original class identifier 550 from the decryption. Thus, the output generated from the attempted decryption may not match the original class identifier 520. On the other hand, if the number of requests with encrypted copies for the same class identifier 520 meets (e.g., is greater than or equal to) a threshold, the class authorizer 550 may be able to recover the original class identifier 550 from the decryption. Thus, the output generated from the attempted decryption may match the original class identifier 520. Using the obtained output, the class authorizer 550 may generate a response including the output. Once generated, the class authorizer 550 on the application manager service 510 may send the response with the obtained output to the client device 430 via the network 405.

[0132] The class authorizer 550 of the application 435 may then receive a response from the class authorizer 550 running on the application manager service 410. Upon receipt, the class authorizer 550 may parse the response to identify the resulting output from the decryption attempt. The class authorizer 550 may compare the resulting output with the class identifier 520 corresponding to the class included in the authentication request. When the resulting output is determined to match, the class authorizer 550 may determine that the class satisfies the threshold number of the application 435. The model applicator 530 may continue to use the class and the corresponding class identifier 520. Conversely, when the resulting output is determined to not match, the class authorizer 550 may determine that the class does not satisfy the threshold number of the application 435. The model applicator 530 may discard the class and the corresponding class identifier 520 for further use. Additionally, the class authorizer 550 may repeat the threshold encryption protocol for another class to discover a class to use in a request for content.

[0133] In some implementations, for each identified class, the class authorizer 550 of the application 435 can check the class identifier 520 corresponding to the class against at least one data structure to make a determination. The data structure can be maintained by the class authorizer 550 running on the application manager service 410 (e.g., on the database 515). The data structure can indicate whether any application 435 is assigned to the class. The data structure can also indicate whether the number of applications 435 assigned to the class meets a threshold number. In some implementations, the data structure can be a probabilistic data structure. The probabilistic data structure can include, for example, a counting Bloom filter, a quotient filter, a Cuckoo filter, a count-min sketch, among others.

[0134] To perform the verification, the class authorizer 550 of the application 435 can generate an authentication request that includes a class identifier 520 corresponding to the class. Once generated, the class authorizer 550 can send the authentication request to a class authorizer 550 running on the application manager service 410. The class authorizer 550 on the application manager service 410 can then receive the authentication request. The class authorizer 550 can parse the authentication request to identify the class identifier 520. The class authorizer 550 can apply a data structure to determine whether the number of applications 435 assigned to the class meets a threshold number. In applying the data structure, the class authorizer 550 can feed the class identifier 520 included in the request into the data structure and identify an output from the data structure. Additionally, each time an authentication request is received, the class authorizer 550 can update the data structure for the class maintained by the application manager service 410.

[0135] If the data structure indicates that the number meets (e.g., is greater than or equal to) the threshold number, the class authorizer 550 on the application manager service 410 can generate a success response. This response can indicate that the number of applications 435 meets the threshold. The class authorizer 550 of the application manager service 410 can send the success response to the class authorizer 550 of the application 435. Based on receiving the success response, the class authorizer 550 of the application 435 can identify the response as indicating that the threshold number is met. Additionally, the model applicator 530 can continue to use the class and corresponding class identifier 520.

[0136] On the other hand, if the data structure indicates that the number does not meet (e.g., is less than) the threshold number, the class authorizer 550 of the application manager service 410 can generate a failure response. The response can indicate that the number of applications 435 does not meet the threshold. The class authorizer 550 of the application manager service 410 can send the failure response to the class authorizer 550 of the application 435. Based on receiving the failure response, the class authorizer 550 of the application 435 can identify the response as indicating that the threshold number is not met. The model applicator 530 can discard the class and corresponding class identifier 520 for further use. Additionally, the class authorizer 550 can repeat the check for another class to find a class to use in a request for content.

[0137] Upon determining that at least one of the classes meets a threshold number of applications 435 assigned to the class, the model applicator 530 can assign the application 435 to the class's corresponding class identifier 520. In some implementations, the model applicator 530 can identify the class identifier 520 that corresponds to the closest class determined to meet the threshold number. By default, the class identifier 520 assigned by the model applicator 530 to the application 435 may correspond to the closest class. In some implementations, the model applicator 530 can use the classification model 500 to identify a class identifier 520 associated with the identified class. With this identification, the model applicator 530 can then assign the class identifier 520 to the application 435. The assignment of a class identifier 520 to the application 435 can indicate that the browsing history 535 for the application 435 is similar to other browsing histories 530 on other applications 435 with the same class identifier 520. The class identifier 520 assigned to an application 435 executing on a client device 430 may be common, shared, or identical to the class identifier 520 for at least some other applications 435 executing on other client devices 430.

[0138] This assignment allows the model applicator 530 to store and maintain the class identifier 520 in the identifier cache 545. For example, as shown, the model applicator 530 of the first application 435A may have identified the first application 435A as having a browsing pattern similar to that of other applications 435B-N with a class corresponding to the first class identifier 520A. The identifier cache 545 can control access to the class identifier 520 by scripts executed by the application 435. In some implementations, the model applicator 530 can store the class identifier 520 in a secure cookie maintained in the identifier cache 545. The secure cookie may include a cookie with a secure flag or an HTTP-only flag set. The secure cookie can prevent scripts on the information resource from accessing the class identifier 520 maintained in the identifier cache 545. Additionally, the secure cookie can permit authorized entities to access the class identifier 520 stored in the identifier cache 545. The secure cookie may identify the content selection service 425 or the application manager service 410 as being allowed to access the class identifier 520 on the identifier cache 545 .

[0139] The model applicator 535 can determine whether to apply the classification model 500 according to an identifier assignment policy. The identifier assignment policy can specify a set of conditions under which the classification model 500 should be applied for reassignment of class identifiers 520. In some implementations, the identifier assignment policy can include a schedule indicating when the classification model 500 should be applied. The identifier assignment policy can specify that the classification model 500 should be updated when new updates to the application 435 are provided by the application manager service 410. The identifier assignment policy can specify that the classification model 500 should be applied when the amount of accumulated additional browsing history 535 meets a threshold amount. If the decision is not to reapply the classification model 500, the model applicator 530 can maintain the class identifiers 520. On the other hand, if the decision is to reapply the classification model 500, the model applicator 530 can repeat the functions described above. For example, the model applicator 530 may identify a browsing history 535, use a dimensionality reduction process to generate a reduced set of feature vectors, apply the classification model 500 to the reduced set of feature vectors to identify classes, and assign a classifier identifier 520 associated with the identified class to the application 435.

[0140] 6, a block diagram is shown illustrating one implementation of a client device 430, a content provider 415, a content publisher 420, and a content selection service 425 in the system 400. Broadly, the application 435 of each client device 430 (e.g., as illustrated for the first client device 430A) may include at least one resource processor 615 to handle reading and parsing at least one information resource 600 and other data communicated with the content provider 415, content publisher 420, or content selection service 425. The application 435 may also include at least one identifier selector 620 to determine which class identifier 520 to insert into a request for content transmitted over the network 405.

[0141] In further detail, the resource processor 615 executing on the client device 430 can receive the information resource 600 from the content publisher 420. The receipt of the information resource 600 may be in response to a request for the information resource 600 sent by the application 435 to the content publisher 420 and may be for presentation at the client device 430. The received information resource 600 (e.g., a web page) may include at least one primary content 605 (e.g., the body, text, and images of the web page) and at least one content slot 610 (e.g., an inline frame of the web page). The primary content 605 may correspond to a portion of the information resource 600 provided by the content publisher 420. The content slot 610 may be capable of receiving content from the content provider 415 or the content selection service 425. The content to be inserted into the content slot 610 may have a hostname that is different from the hostname of the information resource 600. Once received, the resource processor 615 can parse the information resource 600 , including the primary content 605 and the content slots 610 .

[0142] For a content slot 610 of the information resource 600, the resource processor 615 can generate requests 625A-N (generally referred to herein as requests for content 625) to send to the content selection service 425. The generation of the request for content 625 can be according to a script (e.g., an advertisement tag or a content selection tag) for inserting content into the content slot 610. The script for the content slot 610 can be embedded or contained in the content slot 610 itself or in another portion of the information resource 600. In generating the request for content 625, the resource processor 615 can include addresses 630A-N (generally referred to herein as addresses 630) in the request for content 625. The address 630 can reference the content selection service 425 specified by the script for the content slot 610, such as a URL for the content selection service 425. The address 630 can indicate a destination address to which the request for content 625 is to be routed. Additionally, the resource processor 615 can include a source address referencing the client device 430 in the request for content 625. The resource processor 615 can also include an address corresponding to the content publisher 420 that provided the information resource 600 to the application 435.

[0143] Additionally, the identifier selector 620 can identify or otherwise select which of the class identifiers 520 to insert into the request for content 625 according to at least one obfuscation policy (sometimes referred to herein as a publication policy). The obfuscation policy can specify one or more conditions under which the class identifiers 520 are allowed or restricted from being included in the request for content 625. In some implementations, the conditions of the obfuscation policy can be specific to one or more information resources 600, the main content 605 of the information resource 610, or the content slots 610 of the information resource 600, or any combination thereof. In some implementations, the conditions can depend on the entries 540 of the browsing history 535. For example, the conditions of the obfuscation policy can specify that a different class identifier 520 or no class identifier 520 should be included in the request for content 625 for a first-time access of the information resource 600 via the application 435. In another example, a condition of the obfuscation policy may specify that a different class identifier 520 should be used for information resources 600 that are rarely visited by applications 435 (or accessed less than a threshold number of times as indicated by browsing history 535). In another example, the obfuscation policy may specify that the class identifier 520 should not be used unless the information resource 600 is received in accordance with the Hypertext Transfer Protocol Secure (HTTPS) protocol. In this manner, the obfuscation policy may further protect data privacy for accessing the information resource 600 via applications 435 over the network 405.

[0144] When selecting according to the obfuscation policy, the identifier selector 620 can identify one or more conditions for comparison against the obfuscation policy. In some implementations, the identifier selector 620 can identify the information resource 600, including the content and individual parts of the address (e.g., protocol, domain name, and pathname). In some implementations, the identifier selector 620 can identify individual primary content 605 on the information resource 600. In some implementations, the identifier selector 620 can identify a content slot 610 for which a request for content 625 should be generated. In some implementations, the identifier selector 620 can identify an entry 540 in the browsing history 535 of the application 435. With these identifications, the identifier selector 620 can compare against the conditions specified by the obfuscation policy. If the conditions are determined not to match, the identifier selector 620 can retain the class identifier 520 to be included in the request for content 625.

[0145] On the other hand, if the condition is determined to be met, the identifier selector 620 can determine whether the obfuscation policy specifies allowing or restricting the class identifier 520. When the obfuscation policy specifies that the class identifier 520 is allowed, the identifier selector 620 can keep the class identifier 520 for inclusion in the request for content 625. Conversely, when the obfuscation policy specifies that the class identifier 520 is restricted, the identifier selector 620 can discover another class identifier 520 or cannot discover the class identifier 520. In some implementations, the obfuscation policy can specify that another class identifier 520 should be used under such conditions. Thus, the identifier selector 620 can use another class identifier 520 that corresponds to the set of classes identified by the model applicator 530. In some implementations, the obfuscation policy can specify that the class identifier 620 should not be included under such conditions. Thus, the identifier selector 620 can prevent, exclude, or otherwise restrict any class identifier 620 from being included in the request for content 625.

[0146] Upon selecting the class identifier 520, the resource processor 615 can include the selected class identifier 520 for the application 435 in the request for content 625. In some implementations, the resource processor 615 can access the identifier cache 545 to retrieve the class identifier 520. Once retrieved, the resource processor 615 can include the class identifier 520 for inclusion in the request for content 625. In some implementations, the request for content 625 originally generated by the application 435 when parsing the script for the content slot 610 can originally include a unique tracking identifier. The resource processor 615 can remove or otherwise exclude from the request for content 625 any unique tracking identifier corresponding to the application 435 (or the client device 430 executing the application 435 or the account associated with the application 435). The unique tracking identifier can include, for example, a cookie user identifier corresponding to an account associated with the application 435 or a cookie device identifier corresponding to the client device 430 executing the application 435. The unique tracking identifier may have been provided by the content provider 415 or another content disposition service. Once removed, the resource processor 615 can include the class identifier 520 in the request for content 625. In some implementations, the resource processor 615 can replace the unique tracking identifier included in the request for content 625 with the class identifier 520. In some implementations, the resource processor 615 can remove any identifiers, including either the tracking identifier or the class identifier 520, in response to a determination according to an obfuscation policy.

[0147] In some implementations, the resource processor 615 can package or include the class identifier 520 in a designated portion of the request for content 625. In some implementations, the resource processor 615 can include the class identifier 520 in at least one header field of the request for content 625. In some implementations, the resource processor 615 can include the class identifier 520 in the body of the request for content 625. In some implementations, the resource processor 615 can include the class identifier 520 in a cookie. In some implementations, the cookie can be generated along with the request for content 625. In some implementations, the cookie can be retrieved from the application 435 (e.g., the identifier cache 545). The cookie can have a secure flag or an HTTP-only flag set to prevent unauthorized entities from intercepting and accessing the class identifier 520. The unauthorized entities can include entities other than the content selection service 425 or the application manager service 410. By setting a secure or HTTP-only flag, the cookie may also restrict access to the class identifier 520 over a secure communication channel (e.g., Hypertext Transfer Protocol Secure (HTTPS)) through the network 405. The resource processor 615 may include the cookie containing the class identifier 520 in a request for content 625. The cookie may also include an address corresponding to the content publisher 420 that provided the information resource 600 to the application 435. Once generated, the resource processor 615 may send the request for content 625 to the content selection service 425 over the network 405. In some implementations, the application 435 may establish a secure communication channel (e.g., according to HTTP) between the client device 430 and the content selection service 425 to send the request for content 625.Establishing a channel can allow the content selection service 425 access to the class identifier 520 included in the cookie of the request 625 for content.

[0148] The class identifiers 520A-N assigned by each model applicator 530 to different applications 435A-N executing on client devices 430A-N may not be specific to one application 435A-N and may not uniquely identify the applications 435A-N. For example, as shown, a first class identifier 520A may be assigned to a first application 435A on a first client device 430A and to a second application 435B executing on a second client device 430B. In contrast, a second class identifier 520B may be assigned to an nth application 435N executing on an nth client device 430N. This may be in contrast to a unique tracking identifier, such as a user or client identifier provided by a content provider 415 or other content disposition service, that specifically identifies an application 435A-N or client device 430A-N. Additionally, the class identifiers 520A-N may have lower entropy than such unique tracking identifiers because the class identifiers 520A-N cannot uniquely identify each application 435 running on the client device 430. For example, the entropy of a unique tracking identifier may have greater than 63 bits of entropy, while the entropy of a class identifier 520A-N may have between 18 and 52 bits of entropy. Thus, the class identifiers 520A-N may be smaller in size than these unique tracking identifiers, thereby reducing the size of the request 625 for content sent over the network 405.

[0149] 7, a block diagram illustrating one implementation of a client device 430 and a content selection service 425 in the system 400 is shown. The content selection service 425 may include at least one history aggregator 700 to store and maintain browsing history. The content selection service 425 may include at least one class characterizer 705 to determine selection parameters for each class. The content selection service 425 may include at least one content locater 710 to identify content items 725A-N (hereinafter generally referred to as content items 725) for the application 435 associated with the request 620 using the identified classes. The content selection service 425 may include at least one history database 715 to maintain and store browsing history entries 720A-N (hereinafter generally referred to as entries 720) for the class identifiers 520.

[0150] In further detail, a history aggregator 700 executing on the content selection service 425 can collect, aggregate, or otherwise maintain a history database 715 using cookies 630 included in requests for content 625 received from client devices 430. The history database 715 may include a set of entries 720 indexed by class identifiers 520 defined by the application manager service 410. Each entry 720 may include an address for the accessed information resource 600 and a timestamp indicating the time the information resource 600 was accessed. Instead of aggregating the browsing history of client devices 430 using a unique tracking identifier (e.g., a cookie identifier), the history aggregator 700 can aggregate the browsing history by class identifier 520. Unlike the browsing history 530 maintained for each individual application 435, the history database 715 may not individually identify the application 435 (or the user associated with the application 435) from which the entry 720 is generated. Each time a request for content 625 is received, the history aggregator 700 can identify an address corresponding to the information resource 600 on which the content should be returned. The history aggregator 700 can further identify a class identifier 520 included in the request for content 625. These identifications enable the history aggregator 700 to add an entry 720 (not shown in FIG. 4) including an address and a timestamp to the set of entries 720 for the class identifier 520 included in the request for content 625.

[0151] A class characterizer 705 executing on the content selection service 425 can determine one or more characteristics of each class based on the entries 720 for the class's class identifier 520. The characteristics may include, for example, common traits, profiles, behaviors, or interests of the class corresponding to the class identifier 520. In some implementations, the class characterizer 705 can use a class profile model to determine the characteristics of a class based on a set of entries 720 in the history database 715 for the class identifier 520. The class profile model can be any model such as linear regression, logistic regression, an artificial neural network (ANN), a support vector machine (SVM), and a naive Bayes classifier, among others. The class profile model may have been trained using a sample data set that correlates browsing history labeled by the class identifier 520 with certain characteristics. The class profile model can convert the entries 720 in the history database 715 for each class identifier 520 into characteristics for the corresponding class. In some implementations, the class characterizer 705 can store and maintain the characteristics of each class identifier 520. A content placer 710 executing on the content selection service 425 can select or identify a content item 725 from a set of content items 725 in response to a request for content 625 using a class identifier 520. The use of a class identifier 520 can be contrasted with using a unique identifier of a particular user associated with a request for content, in that the selection of a content item 725 may not be based on an identifier unique to a particular user. Each content item 725 can include an object or element to be embedded, inserted, or otherwise added to a content slot 610 of the information resource 600. Each content item 725 can be provided by one or more of the content providers 415. Upon receipt, the content placer 710 can parse the request for content 625 to identify the class identifier 520.Upon identification, the content placer 710 can identify characteristics of the class corresponding to the class identifier 520. The content placer 710 can identify or select a content item 725 associated with the characteristics of the class. In some implementations, the content placer 710 can select the content item 725 using a content placement process. The content placement process can use models such as linear regression, logistic regression, artificial neural networks (ANNs), support vector machines (SVMs), and naive Bayes classifiers, among others. For each content item 725, the content placement process can calculate, determine, or generate a predicted probability of interaction by a user in the class corresponding to the class identifier 520 included in the request for content 625. The content placer 710 can identify the content item 725 with the highest probability of interaction by a user in the class. Once selected, the content placer 710 can transmit the content item 725 to the client device 430 that issued the request for content 625. In some implementations, the content placer 710 can send the address of the content item 725 to the client device 430 for the application 435 to retrieve the selected content item 725 from the content provider 415.

[0152] 6 in conjunction with FIG. 7, the resource processor 615 may receive the content item 725 identified by the content selection service 425. The resource processor 615 may embed, insert, or add the content item 725 into a content slot 610 of the information resource 600. In some implementations, the resource processor 615 may receive an address of the content item 725. The address of the content item 725 may reference a content provider 415. The resource processor 615 may send another request to the content provider 415 to retrieve the content item 725 and insert the content item 725 into the content slot 610 of the information resource 600.

[0153] In this way, the content selection service 425 can select content items 725 with the expectation that users associated with the same class identifier 520 are expected to have similar responses. Furthermore, the security, integrity, and privacy of data about individual users' browsing histories 530 can be protected. At the same time, being able to determine the relevance of content items 725 to individual classes in the selection of content items 725 can maintain the quality of human-computer interaction (HCI) with the information resource 600 as a whole.

[0154] Referring now to FIG. 8, a flow diagram illustrates an implementation of a method 800 for encoding an identifier for content selection using a classification model. The functionality described herein with respect to method 800 may be implemented or otherwise performed by system 400 as shown in FIG. 4 or a computing device as shown in FIG. 9. Broadly, an application running on a client device may identify an accessed information resource (805). The application may reduce dimensionality (810). The application may apply a classification model (815). The application may identify a class (820). The application may assign a class identifier (825). The application may determine whether the class identifier is allowed (830). If not, the application may assign a different class identifier (835). The application may receive the information resource (840). The application may generate a request for content (845). The application may determine whether to obfuscate the class identifier (850). If obfuscated, the application may use a different class identifier (855). The application may include the class identifier (860). The application may send a request for content (865). The application may receive the selected content item (870). The application may decide whether to reallocate (875). If not, the application may keep the class identifier (880).

[0155] In further detail, an application (e.g., application 435) on a client device (e.g., client device 430) can identify accessed information resources (805). In some implementations, the application can identify the accessed information resources from a browsing history (e.g., browsing history 530). The browsing history can include a set of entries (e.g., entries 540). Each entry can include an address of an accessed information resource and, in some implementations, a timestamp that identifies the time the information resource was accessed. For each accessed information resource, the application can extract features from the content on the information resource. The application can generate a set of feature vectors from the set of browsing history entries. The feature vectors can be a projection of the browsing history onto a set of dimensions. The dimensions can include, among other things, a time range and an address attribute. The application can reduce dimensions (810). Using a dimension reduction process, the application can generate a set of reduced feature vectors from an initial set of feature vectors. In some implementations, step (810) may be performed across application 435 and another server. In some implementations, step (810) may be omitted. In some implementations, step (810) may be replaced by or combined with training of the classification model. For example, the classification model may be trained using a distributed learning protocol, such as federated learning using an optimization algorithm (e.g., stochastic gradient descent (SGD) or averaging). During each interaction, each application may set or adjust values ​​for the classification model using the training dataset and may provide values ​​to each other according to the distributed learning protocol.

[0156] The application may apply a classification model (e.g., classification model 500) to features extracted from the accessed information resource (815). The classification model may include a set of inputs, a set of parameters, and a set of outputs. The classification model may be, for example, a classification algorithm, a clustering model, or a locality-sensitive hash function, among others. The set of inputs may include features extracted from the accessed information resource, browsing history entries, or a representation of browsing history. The set of parameters may associate the inputs with outputs. The set of outputs may include classes into which users interacting with the application should be classified based on the user's browsing history on the application. The application may provide a set of reduced-dimensionality feature vectors as inputs to the classification model. The classification model may apply the parameters to the inputs. The application may identify classes (820). By applying the classification model, the classification model may generate an output identifying one or more classes into which users interacting with the application should be classified. By extension, the classification of the application may correspond to or include a classification of users associated with the application operated by the user or the account operating the client device. The application can assign 825 a class identifier (e.g., class identifier 520). The classification model can associate each class with one of the class identifiers. Once the class identifiers are identified, the application can identify the class identifier that corresponds to the class and assign the application to the class identifier.

[0157] The application may determine whether the class identifier is authorized (830). This determination may be in coordination with an authentication service (e.g., class authorizer 550 on application manager service 410) according to a threshold encryption scheme. The application may send an encrypted copy of the class identifier corresponding to the class. Under the threshold encryption scheme, the authentication service may be unable to decrypt the class identifier unless the number of requests with encrypted copies of the same identifier exceeds a threshold number. The inability to decrypt may be because, for example, the authentication service may not have received a sufficient portion of the decryption key (e.g., a private key) from the request. Once a sufficient portion is received, the authentication service may be able to successfully decrypt the encrypted copy. The authentication service may return an output of the decryption attempt. The application may compare the output with the original class identifier. If there is a match, the application may determine that the class is authorized. Otherwise, the application may determine that the class is not authorized. If the class identifier is not authorized, the application may assign another class identifier (835). The application may find another class identifier from the identified set of classes and repeat the function of (830).

[0158] An application may receive (840) an information resource (e.g., information resource 600). The information resource may include primary content (e.g., primary content 605) and content slots (e.g., content slots 610). The primary content may be provided by a content publisher (e.g., content publisher 420). The content slots may be available for insertion of content from a content provider (e.g., content provider 415) or a content selection service (e.g., content selection service 425). Upon receipt, the application may parse the information resource. The application may generate (845) a request for content (e.g., request for content 625). The generation of the request for content may be coordinated with the analysis of the information resource.

[0159] The application may determine whether to obfuscate the class identifier (850). The decision may be based on an obfuscation policy. The obfuscation policy may specify one or more conditions under which the class identifier is constrained to be included in a request for content. For example, the conditions may include the security protocol (e.g., HTTPS) under which the information resource is received. The obfuscation policy, in this example, may specify that the class identifier should not be included in a request for content when the information resource is not encrypted under HTTPS. The application may specify conditions on the information resource to compare against the conditions specified by the obfuscation policy. If the conditions do not match, the application may determine that the class identifier should not be obfuscated and may maintain the current class identifier. Otherwise, if the conditions do not match, the application may determine that the class identifier should be obfuscated. If it is determined that the class identifier should be obfuscated, the application may use a different class identifier (855).

[0160] The application may include a class identifier in a request for content (860). The request for content may include a class identifier corresponding to the class identified using the classification model. The application may also remove any unique tracking identifier associated with the user of the application, the application itself, or the client device running on the application. The unique tracking identifier may have been provided as part of a third-party cookie by the content provider or another content disposition platform. The class identifier may be included as part of a secure cookie included in the request for content. The application may send a request for content to a content selection service or another content provider (865). The transmission may be over a secure communication channel established between the client device and the content selection service. The request for content may be received by the content selection service. The content selection service may use the class identifier included in the request to identify a content item (e.g., content item 725) from a set of content items. Upon selection, the content selection service may send the content item to the application. The application may receive the selected content item (870). The application may insert the content item into a content slot defined on the information resource.

[0161] The application can decide whether to reassign the class identifier (875). The reassignment can be in accordance with an identifier assignment policy. The policy can specify a set of conditions under which the classification model should be reapplied to the browsing history to discover a new class identifier. For example, the reassignment policy can specify that the classification model should be reapplied when the amount of additional entries in the browsing history since the previous assignment exceeds a threshold amount. If the decision is to reassign, the application can repeat functions from (805) through (835) onward. On the other hand, if the decision is not to reassign, the application can retain the class identifier (880).

[0162] Thus, the systems and methods described herein enable the selection of content items relevant to a user without tracking the user's activities individually. In this way, data security, integrity, and privacy about an individual user's browsing history can be protected. At the same time, the ability to determine the relevance to individual classes in the selection of content items can maintain the quality of human-computer interaction (HCI) with the entire information resource.

[0163] 9 illustrates the general architecture of an exemplary computer system 900 that may be utilized to implement any of the computer systems discussed herein (application manager service 410 and its components, content provider 415 and its components, content publisher 420 and its components, content selection service 425 and its components, and client device 430 and its components) according to some implementations. Computer system 900 may be used to provide information for display over a network 930. Computer system 900 comprises one or more processors 920 communicatively coupled to memory 925, one or more communication interfaces 905 communicatively coupled with at least one network 930 (e.g., network 405), and one or more output devices 910 (e.g., one or more display units) and one or more input devices 915.

[0164] The processor 920 may include a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or the like, or a combination thereof. The memory may include, but is not limited to, an electrical, optical, magnetic, or any other storage or transmission device capable of providing program instructions to the processor. The memory 925 may comprise any computer-readable storage medium and may store computer instructions, such as processor-executable instructions for implementing various functions described herein for each system, as well as any data generated by or received through a communication interface or input device (if any). The memory 925 may include a floppy disk, a CD-ROM, a DVD, a magnetic disk, a memory chip, an ASIC, an FPGA, a read-only memory (ROM), a random-access memory (RAM), an electrically erasable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical medium, or any other suitable memory from which the processor can read instructions. The instructions may include code from any suitable computer programming language.

[0165] The processor 920 shown in FIG. 9 may be used to execute instructions stored in memory 925, and in so doing may read from or write to memory various information that is processed or generated in accordance with the execution of the instructions. The processor 920 (collectively referred to herein as a processing unit) coupled to the memory 925 may be included in the application manager service 410. For example, the application manager service 410 may include the memory 925 as the database 515. The processor 920 (collectively referred to herein as a processing unit) coupled to the memory 925 may be included in the content provider 415. For example, the content provider 415 may include the memory 925 to store the content items 725. The processor 920 (collectively referred to herein as a processing unit) coupled to the memory 925 may be included in the content publisher 420. For example, the content publisher 420 may include the memory 925 to store the information resource 600. The processor 920 (collectively referred to herein as a processing unit) coupled with the memory 925 may be included in the content selection service 425. For example, the content selection service 425 may include the memory 925 as the history database 720. The processor 920 (collectively referred to herein as a processing unit) coupled with the memory 925 may be included in the client device 430. For example, the client device 430 may include the memory 925 as the browsing history 535 and the identifier cache 545.

[0166] The processor 920 of the computer system 900 may also be communicatively coupled to or control the communication interface 905 to send or receive various information pursuant to the execution of instructions. For example, the communication interface 905 may be coupled to a wired or wireless network, bus, or other communication means, thereby enabling the computer system 900 to send information to or receive information from other devices (e.g., other computer systems). Although not explicitly shown in the systems of FIGS. 4-7 or 9, one or more communication interfaces facilitate the flow of information between components of the system 900. In some implementations, the communication interface may be configured (e.g., via various hardware or software components) to provide a website as an access portal to at least some aspects of the computer system 900. Examples of the communication interface 905 include user interfaces (e.g., applications 435, information resources 600, primary content 605, content slots 610, and content items 725) through which a user can communicate with other devices in the system 400.

[0167] 9 may be provided, for example, to allow various information to be seen or otherwise perceived in connection with the execution of instructions. Input devices 915 may be provided, for example, to allow a user to make manual adjustments, make selections, input data, or otherwise interact in any of a variety of ways with the processor during the execution of instructions. Additional information regarding general computer system architectures that may be utilized for the various systems discussed herein is provided further herein.

[0168] Network 930 may include a computer network such as the Internet, a local area network, a wide area network, a metro area network, or other area network, an intranet, a satellite network, a voice or data cellular communication network, or other computer network, and combinations thereof. Network 930 may be any type of computer network that relays information between components of system 400, such as application manager service 410, content providers 415, content publishers 420, content selection service 425, and client devices 430. For example, network 930 may include the Internet and / or other types of data networks, such as a local area network (LAN), a wide area network (WAN), a cellular network, a satellite network, or other types of data networks. Network 930 may also include any number of computing devices (e.g., computers, servers, routers, network switches, etc.) configured to receive and / or transmit data within network 930. Network 930 may further include any number of wired and / or wireless connections. The client device 430 may communicate wirelessly (e.g., via WiFi, cellular, radio, etc.) with a transceiver that is wired (e.g., via fiber optic cable, CAT5 cable, etc.) to other computing devices in the network 930.

[0169] Implementations of the subject matter and operations described herein may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware embodied on tangible media, including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. Implementations of the subject matter described herein may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a computer storage medium for execution by or to control the operation of a data processing device. The program instructions may be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium may be or be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or one or more combinations thereof. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium may include a source or destination of computer program instructions encoded on an artificially generated propagated signal. The computer storage medium may also be, or may be contained in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).

[0170] The features disclosed herein may be implemented in a smart television module (or connected television module, hybrid television module, etc.), which may include a processing module configured to integrate more traditional television program sources (e.g., received via cable, satellite, airwaves, or other signals) with an Internet connection. The smart television module may be physically built into a television set or may include a separate device, such as a set-top box, Blu-ray or other digital media player, game console, hotel television system, or other companion device. The smart television module may be configured to allow viewers to search and discover videos, movies, photos, and other content on the web, on local cable TV channels, on satellite TV channels, or stored on a local hard drive. A set-top box (STB) or set-top unit (STU) may include an information appliance device, which includes a tuner and can connect to a television set and external sources of signals to convert the signals into content that is then displayed on a television screen or other display device. The smart television module may be configured to provide a home or top-level screen that includes icons for multiple different applications, such as a web browser and multiple streaming media services, connected cable or satellite media sources, other web "channels," etc. The smart television module may further be configured to provide an electronic program guide to the user. A companion application to the smart television module may be operable on a mobile computing device to provide the user with additional information about available programs, allow the user to control the smart television module, etc.In some implementations, features may be implemented on a laptop computer or other personal computer, a smartphone, other mobile phone, a handheld computer, a tablet PC, or other computing device. In some implementations, features disclosed herein may be implemented on a wearable device or component (e.g., a smart watch) that may include a processing module configured to integrate Internet connectivity (e.g., with another computing device or network 930).

[0171] The operations disclosed herein may be implemented as operations performed by a data processing apparatus on data stored in one or more computer-readable storage devices or data received from other sources.

[0172] The terms "data processing apparatus," "data processing system," "user device," or "computing device" encompass all types of apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, a system on a chip or chips, or a combination of the foregoing. An apparatus may include special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, an apparatus may also include code that creates an execution environment for a subject computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or one or more combinations thereof. The apparatus and execution environment may implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0173] A computer program (also known as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted, declarative or procedural, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program may be deployed to run on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.

[0174] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform activities by manipulating input data and generating output. The processes and logic flows may also be performed by, and an apparatus may be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0175] Processors suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for performing activities in accordance with instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices, such as magnetic disks, magneto-optical disks, or optical disks, for storing data, or is operatively coupled to receive data from, transfer data to, or both. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive). Suitable devices for storing computer program instructions and data include all types of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory may be augmented by, or incorporated in, special purpose logic circuitry.

[0176] To enable interaction with a user, implementations of the subject matter described herein may be implemented on a computer having a display device, such as a CRT (cathode ray tube), plasma, or LCD (liquid crystal display) monitor, for displaying information to the user, as well as a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to enable interaction with the user. For example, feedback provided to the user may include any form of sensory feedback, such as visual feedback, audible feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, speech, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from a device used by the user, for example, by sending web pages to a web browser on the user's client device in response to a request received from the web browser.

[0177] Implementations of the subject matter described herein may be implemented in a computing system that includes back-end components, e.g., as data servers, or that include middleware components, e.g., application servers, or that include front-end components, e.g., client computers having graphical user interfaces or web browsers through which users can interact with implementations of the subject matter described herein, or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include local area networks (“LANs”) and wide area networks (“WANs”), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0178] A computing system such as system 900 or system 400 may include clients and servers. For example, application manager service 410, content provider 415, content publisher 420, and content selection service 425 of system 400 may each include one or more servers in one or more data centers or server farms. Clients (e.g., client device 430) and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some implementations, a server sends data (e.g., HTML pages) to a client device (e.g., for the purpose of displaying data to and receiving user input from a user interacting with the client device). Data generated at a client device (e.g., as a result of user interaction) may be received from the client device at the server.

[0179] Implementations of the subject matter described herein may be implemented in a computing system that includes back-end components, e.g., as data servers, or that include middleware components, e.g., application servers, or that include front-end components, e.g., client computers having graphical user interfaces or web browsers through which users can interact with implementations of the subject matter described herein, or that include any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, e.g., a communications network. Communications networks may include local area networks (“LANs”) and wide area networks (“WANs”), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0180] In situations where the systems discussed herein may collect or utilize personal information about a user, the user may be given the opportunity to control whether a program or feature may collect personal information (e.g., information about the user's social network, social behavior or activities, the user's preferences, or the user's location) or to control whether or how content that may be more relevant to the user is received from a content server or other data processing system. Additionally, some data may be anonymized in one or more ways before it is stored or used, so that personally identifiable information is removed when generating parameters. For example, a user's identification information may be anonymized so that personally identifiable information cannot be determined about the user, or so that the user's geographic location from which location information is obtained may be generalized (such as to the city level, zip code level, or state level) so that the user's specific location cannot be determined. Thus, users have control over how information about them is collected and used by content servers.

[0181] While this specification contains details of many specific implementations, these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features specific to particular implementations of a particular invention. Some features described herein in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented in multiple implementations separately or in any suitable subcombination. Moreover, while features may be described as operating in a certain combination, and even initially claimed as such, one or more features from a claimed combination may in some cases be omitted from the combination, and the claimed combination may be directed to a subcombination or variations of the subcombination.

[0182] Similarly, while operations are shown in the figures in a particular order, this should not be understood as requiring that the operations be performed in the particular or sequential order shown, or that all of the shown operations be performed, to achieve desirable results. In some situations, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the above-described implementations should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products.

[0183] Thus, specific implementations of the present subject matter have been described. Other implementations are within the scope of the following claims. In some cases, the activities recited in the claims may be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown or sequential order to achieve desirable results. In some implementations, multitasking or parallel processing may be utilized. [Explanation of symbols]

[0184] 100 Profile Vectors 102 devices 200 client devices 202 processors 204 Network Interface 206 memory 208 Applications 210 Access Log 212 Singular Vectors 214 Log Reducer 216 Neural Network Model 218 Classifier 220 Cluster Identifiers 225 Network 230 Classifier Server 232 Concentrator 250 content servers 252 content items 254 Content Selector 400 System 405 Network 410 Application Manager Service 415 Content Provider 420 Content Publisher 425 Content Selection Service 430 Client Device 435 Applications 500 classification models 505 Model Trainer 510 Model Updater 515 databases 520 Class Identifier 525 Content Extractor 530 Model Applicator 535 Browsing History 540 entries 545 Identifier Cache 550 Class Authorizer 600 Information Resources 605 Main Content 610 content slots 615 Resource Processor 620 Identifier Selector 625 request 630 Address 700 History Aggregator 705 Class Characterizer 710 Content Placer 715 History Database 720 entries 725 content items 900 Computer Systems 905 Communication Interface 910 Output Device 915 Input Devices 920 processor 925 memory 930 Network

Claims

1. providing, by a client device, access data to a server corresponding to an information source access history of said client device; receiving singular vectors and model weights from the server based on the access data provided to the server by the client device; generating, by the client device, a dimensionality-reduced vector of the accessed data using the singular vectors received from the server; classifying, by the client device, the dimensionality-reduced vectors into classes based on model parameters received from the server, the classes having class identifiers; sending, by the client device, a request for content to a content selection service that includes the class identifier; and presenting, by the client device, content provided by the content selection service in response to the request for content including the class identifier. method.

2. The method described in claim 1, further comprising a step of determining, by the server, parameters of clusters of the dimensionally reduced matrix.

3. The method described in claim 2, further comprising a step of adjusting a classifier model by the server based on the parameters of the clusters of the dimensionality-reduced matrix.

4. The method of claim 3, further comprising a step of determining a score for the cluster by the client device.

5. The method described in claim 4, wherein the step of classifying the dimensionality reduced vector by the client device includes a step of classifying the dimensionality reduced vector based on the model parameters and the scores of the clusters.

6. The method of claim 1 , further comprising categorizing, at the client device, a plurality of applications of the client device into clusters based on the information source access history of the client device.

7. A method as described in claim 6, wherein the step of determining by the client device that two applications have matching browsing patterns includes the step of categorizing the plurality of applications into several clusters based on the matching browsing patterns.

8. A system including a client device, the client device comprising: providing access data corresponding to an information source access history of the client device to a server; receiving singular vectors and model weights from the server based on the access data provided to the server; generating a dimensionality-reduced vector of the access data using the singular vectors received from the server; classifying the dimensionality-reduced vectors into classes based on model parameters received from the server, the classes having class identifiers; sending a request for content to a content selection service that includes the class identifier; presenting content provided by the content selection service in response to the request for content including the class identifier; configured to: system.

9. The system further comprising: a server configured to determine parameters of clusters of the dimensionality-reduced matrix. The system of claim 8.

10. The system described in Claim 9, wherein the server is further configured to adjust a classifier model based on the parameters of the clusters of the dimensionality-reduced matrix.

11. The system of claim 10, wherein the client device is further configured to determine a score for the cluster.

12. The system of claim 11 , wherein classifying the reduced-dimensionality vector comprises classifying the reduced-dimensionality vector based on the model parameters and the scores of the clusters.

13. The system of claim 8, wherein the client device is further configured to categorize multiple applications of the client device into clusters based on the information source access history of the client device.

14. The client device, and determining that two applications have matching browsing patterns, wherein categorizing the plurality of applications includes categorizing the two applications into a number of clusters based on the matching browsing patterns. The system of claim 13.

15. One or more computer-readable storage media storing instructions that, when executed by a client device, cause the client device to: providing access data corresponding to an information source access history of the client device to a server; receiving singular vectors and model weights from the server based on the access data provided to the server; generating a dimensionality-reduced vector of the access data using the singular vectors received from the server; classifying the dimensionality-reduced vectors into classes based on model parameters received from the server, the classes having class identifiers; sending a request for content to a content selection service that includes the class identifier; presenting content provided by the content selection service in response to the request for content including the class identifier; to carry out A computer-readable storage medium.

16. The instructions, when executed by a server, cause the server to: The computer-readable storage medium of claim 15 , further comprising: determining parameters of clusters of a dimensionally reduced matrix.

17. The instructions, when executed by a server, cause the server to: further adjusting a classifier model based on the parameters of the clusters of the dimensionality reduced matrix.

17. The computer-readable storage medium of claim 16.

18. The instructions, when executed by the client device, cause the client device to: The computer-readable storage medium of claim 17 , further comprising determining a score for the cluster.

19. 20. The computer-readable storage medium of claim 18, wherein classifying the reduced-dimensionality vector comprises classifying the reduced-dimensionality vector based on the model parameters and the scores of the clusters.

20. the instructions, when executed by the client device, further cause the client device to categorize a plurality of applications of the client device into clusters based on the information source access history of the client device.

16. The computer-readable storage medium of claim 15.

Citation Information

Patent Citations

  • Item recommendation program, item recommendation method, and item recommendation apparatus

    JP2017182724A

  • System for detecting associations between items

    US20090089273A1