Time-series user interest in content of computer networks
The system addresses the challenge of accurately predicting user interest in digital content by grouping users into cohorts and using time-series forecasting, effectively balancing accuracy and privacy to enhance content delivery and advertisement strategies.
Patent Information
- Application Number
- PCT/IB2024/061523
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-11-18
- Publication Date
- 2025-06-05
AI Technical Summary
Existing technologies face challenges in accurately tracking and predicting user interest in digital content across computer networks, balancing accuracy with user privacy concerns.
A system that monitors content requests, categorizes them, groups users into cohorts based on similar interests, and uses time-series clustering and ETS models to generate periodic forecasts of content requests, detecting anomalies when actual requests deviate from the forecast.
The system provides accurate and privacy-compliant insights into user behavior, enabling targeted content delivery and advertisement strategies, while reducing unnecessary content distribution by identifying deviations from predicted user interest.
Smart Images

Figure IB2024061523_05062025_PF_FP_ABST
Abstract
Description
Time-Series User Interest in Content of Computer NetworksBackground
[0001] Tracking user behavior in the consumption of digital content, so as to enable the delivery relevant content, is a growing area of technological development. Accuracy and privacy are two prime concerns that are often in tension.Brief Description of the Figures
[0002] FIG. 1 is a block diagram of an example system for generating cohorts of user profiles.
[0003] FIG. 2 is a plot of an example rate of engagement with content.
[0004] FIG. 3 is a plot of five prototypical cohorts as generated through 41 real users across roughly 10 minutes of actual web activity.
[0005] FIG. 4 is a plot of an example time-series forecast.
[0006] FIG. 5 is a plot of example actual interest compared to the time-series forecast of FIG. 4.
[0007] FIG. 6 is a plot of an example detection of an anomaly in the time-series forecast of FIG. 4.Summary
[0008] An aspect of the specification provides a non-transitory machine-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to collectively: monitor requests for content issued by user computers to a computer network; categorize the requests for content into categories of content; group users of the user computers into cohorts; and for a particular category of content: generate a periodic forecast of future requests for content by a cohort over a period using time-series clustering; compare newrequests for content issued by the cohort within the period to the forecast; detect an anomaly when the new requests deviate from the forecast; and output an indication of the anomaly.
[0009] An aspect of the specification provides a non-transitory machine-readable medium, wherein the instructions are to apply an error, trend, seasonality (ETS) model to generate the periodic forecast of future requests for content.
[0010] An aspect of the specification provides a non-transitory machine-readable medium, wherein the period is selected to encapsulate cyclic user behavior.
[0011] An aspect of the specification provides a non-transitory machine-readable medium, wherein the period is weekly.
[0012] An aspect of the specification provides a non-transitory machine-readable medium, wherein the instructions are to perform a normal analysis detect the anomaly.
[0013] An aspect of the specification provides a non-transitory machine-readable medium, wherein the instructions are to categorize the requests for content by applying a machine learning model.
[0014] An aspect of the specification provides a system including: one or more computer servers, each server including memory and one or more processors, the one or more computer servers connected to a computer network and configured to: monitor requests for content issued by user computers to the computer network; categorize the requests for content into categories of content; group users of the user computers into cohorts; and for a particular category of content: generate a periodic forecast of future requests for content by a cohort over a period using timeseries clustering; compare new requests for content issued by the cohort within the period to the forecast; detect an anomaly when the new requests deviate from the forecast using; and output an indication of the anomaly.
[0015] An aspect of the specification provides a system, wherein the one or more computer servers are configured to apply an error, trend, seasonality (ETS) model to generate the periodic forecast of future requests for content.
[0016] An aspect of the specification provides a system, wherein the period is selected to encapsulate cyclic user behavior.
[0017] An aspect of the specification provides a system, wherein the period is weekly.
[0018] An aspect of the specification provides a system, wherein the one or more computer servers are configured to perform a normal analysis detect the anomaly.
[0019] An aspect of the specification provides a system, wherein the one or more computer servers are configured to categorize the requests for content by applying a machine learning model.Detailed Description
[0020] Understanding user behavior is important to serve digital content that interests users and to avoid serving content that does not interest users, whether such content is news, learning materials, entertainment, healthcare resources, general information, advertising, etc.
[0021] This disclosure provides techniques for automatic forecasting of user behavior in the time-domain and determining when a user or group of users behaves in a way that does not conform to the forecast. Detection of anomalies can offer insights into user behavior. For example, group of people that habitually reads a certain type of news at certain times during the week may exhibit an anomaly that demonstrates a shift to another type of news. This shift may be useful for news providers to understand, so as to change their offerings to meet user interest. In another example, user behavior may change when a large purchase nears, in a user may conduct research into the purchase or an alternative purchase. Digital advertisers may find this information useful to tailor their advertisement to such users.
[0022] FIG. 1 shows an example system 100 for generating cohorts of user profiles. The system 100 may be implemented by one or more computer servers, each of which includes one or more processors, memory, and non-transitory machine -readable medium.
[0023] The system 100 operates on log files 102 that represent user requests for content from content providers available on the internet. The system 100 processes the log files 100 among other information to maintain a cohort database 104. The cohort database 104 maps user profilesto cohorts, which are groups users that demonstrate similar interests and behavior. The cohort database 104 is responsive to outside queries 106 with responses 108. A query 106 may indicate user identifier, and the response 108 may indicate a cohort to which the indicated user belongs. External media entities may use the query -response mechanism to tailor content and / or advertisements to users.
[0024] User data may be anonymized to comply with privacy requirements. For example, the core network 140 may provide a core-network user identifier that links a user to an account, as well as private information such as a name, an address, etc. The cohort database 104 may provide a cohort user identifier that links a user to a cohort. A mapping of core-network user identifier to cohort user identifier may link the two sets of data.
[0025] The system 100 includes a categorization engine 110, a user database 112, a Uniform Resource Locator (URL) database 114, a cohort engine 116, a cohort database 104, an audience segments Application Programming Interface (API) 118, and a profiling engine 120.
[0026] The categorization engine 110 receives and operates on log files 102 to categorize requests for content into categories, such as Interactive Advertising Bureau (IAB) categories. Log files 102 track user requests for content from content providers that offer content via the internet in the form of websites, web services, webapps, etc. The log files 102 may be generated by a computer within an Internet Service Provider’s (ISP) core network 140, where such computer has access to all requests and responses made by users of the ISP.
[0027] Logs files 102 include logs, each of which is a string of information describing a single communication event that occurred within the network. An example log schema is as follows:
[0028] Date Time User-name Group-name Denied-categories All-categories Denied-Flag URI Client-IP Destination-IP Interceptor-IP Logger-ID
[0029] Each individual log thus indicates a user (“Client-IP") has accessed a website or other resource (“URI” or Uniform Resource Identifier) at a particular time (“Date” and “Time”). Note that URL and URI are related and are used interchangeably herein, given that it is well known how to translate between the two.
[0030] An example log is as follows:
[0031] 2022-06-1602:01:26 tral60518k default 45,57,301,2079 45,57,301,2079 0 http: / / olive01.cdnl8.cloudtech.live / live / cnn / l_526283_74274218_43771.ts 84.235.107.71 178.218.196.218 84.235.107.71 mens 1. production
[0032] Typical requests and responses to obtain visit a webpage may result in tens if not hundreds of such logs.
[0033] Log files 102 thus describe the monitoring of user requests for content issued via user computers 142 to a computer network, such as the internet.
[0034] The categorization engine 110 uses machine learning to categorize the content indicated by logs contained in the log files 102. Natural language processing may be used to extract text from raw HyperText Markup Language (HTML) at a URL. For example, a logistic regression model as applied to a Term Frequency - Inverse Document Frequency (TF-IDF) representation of a webpage’s text may be used.
[0035] The categorization engine 110 may scrape URLs at a rate and scale commensurate with the magnitude of the logs 102 received. There is no need to scrape every URL present within the log files 102 as all communications within the network core are present within the log files and many individual logs may represent a single human request for content. Unconstrained scraping may result in hundreds of connections being opened, which would be unnecessary. As such, a set of filters is applied to each URL to reduce noise and obtain a summary of human-oriented requests.
[0036] The categorization engine 110 sends output to the user database 112 and the URL database 114. The user database 112 maps users to URLs and stores user-centric information found within the log files 102, i.e., indications of which users accessed which URLs at which times. The URL database 114 maps each URL to its category determined by the categorization engine 110. A link 122, such as URL or other identifier, associates the data contained in the two databases 112, 114 to form a transitive link between user and the categories of content that interest them.
[0037] The cohort engine 116 transforms the output of the categorization engine 110, as stored in the user database 112 and the URL database 114, into the cohorts which are useful provide insight into user behavior, which may be of interest to content delivery / generation companies, advertisers, etc.
[0038] The behavior of each user is divided into time-based representations across various categories. That is, time is tracked by the user database 112 and category is tracked by the URL database 114. The cohort engine 116 may perform a rate-of-engagement computation on timebased category-specific user behavior.
[0039] Rate of engagement refers to movement between peaks and troughs of interest within some category across time. See FIG. 2, for example. Various computations may be used to calculate rate of engagement as there are many possible interpretations of how the times at which a user triggers a given category translate into the user’s interest in that category. For instance, one might view a high frequency of category triggers as implying a high interest. In the examples discussed in detail herein, this approach, in which high frequency of requests indicates high interest, is used to determine rate of engagement: more category triggers within a predefined time window imply a higher interest in that category at that time. A conceptual example of this interest-calculation is shown by FIG. 2. In other examples, the methodology may take an opposite approach in that a long period visiting a single webpage indicates high interest in its contents. That is, a low frequency of requests while on the same page or same website indicates high interest.
[0040] The calculated rate of engagement is used to create and maintain cohorts in the cohort database 104. In various examples, the cohort engine 116 uses unsupervised machine learning to cluster users with similar rates of engagement. Time-series clustering may be used given the temporal nature of rate of engagement. As the data may have high dimensionality, spherical clustering may be used as opposed to algorithms based on Euclidean metrics. This may help circumventing problems caused by high dimensionality.
[0041] FIG. 3 shows an example of five prototypical cohorts as generated through 41 real users across roughly 10 minutes of web activity. In each case, the horizontal axis represents time (seconds) and the vertical axis represents the computation of each user’s engagement with the“automotive” IAB category. The peaks and troughs present with each user’s calculated engagement roughly match at an intra-cohort level, albeit the similarity shows limits given the small number of users. Th same or similar approach may be used for vast collections of users over large timeframes.
[0042] Once an appropriate level of intra-cohort similarity is achieved, the cohort engine 116 performs time-series calculations, which may be supplemented by expert knowledge, to forecast, model, and interpret macro-level behavior of each cohort of the cohort database 104. The cohort engine 116 is configured to generate a periodic forecast of future requests for content by users belonging to a cohort over a period, such as a week. Any suitable period may be selected with a week being a useful example. A suitable period will encapsulate cyclic user behavior that allows anomalies to be detected. Other suitable examples include a day and a year. The cohort engine 116 may use an error, trend, seasonality (ETS) model to generate the periodic forecast of future requests for content. An example of a time-series forecast is shown in FIG. 4, which has a horizontal axis of time and a vertical axis as interest.
[0043] Each resultant annotated cohort describes forecasted behavior that may be useful to entities that deliver content, create and / or disseminate advertisements, and similar. Such entities may be better able to predict what content at what times will be requested by their users.
[0044] The cohort engine 116 may be configured to perform time-series clustering as discussed in the accompanying document entitled “Time Series,” which forms an integral part of this disclosure.
[0045] The profiling engine 120 is configured to interpret the cohorts of the cohort database 104 in terms of marketability and describe each cohort to content delivery entities, such as advertisers.
[0046] The profiling engine 120 may be configured with hard-coded expert knowledge in the form of rules, statistical time-series processes, machine learning, or a combination of such to recover a profile for each cohort. The profiling engine 120 may be configured with purely interpretational functionality, in that the profiling engine 120 does not alter the existing cohorts of the cohort database 104 in any way. Thus, the profiling engine 120 may be configured toannotate cohorts with use-case specific information. As such, the profiling engine 120 merely appends its determinations to the cohort database 104. Annotated cohort data may be exposed to content delivery entities, such as advertisers, through the audience segments API 118.
[0047] The profiling engine 120 may be configured and controlled by a human profiler 124.
[0048] The profiling engine 120 includes an anomaly detector 130. The anomaly detector 130 detect anomalies in actual interest compared to expected cohort behavior, which may be useful for content or advertisement creation or delivery. If a cohort, subgroup of a cohort, or individual user behaves in a way that does not conform to the forecast, then such behavior may useful to highlight, so as to better deliver and / or create content and / or advertisements.
[0049] The anomaly detector 130 is configured to operate on the periodic forecasts generated by the cohort engine 116. As discussed above, for a given cohort, a forecast relates to future requests for content by users belonging to the cohort operating their user computers over a predefined period, such as a week (e.g., Monday to Sunday). FIG. 4 shows an example timeseries forecast.
[0050] For a given cohort, the anomaly detector 130 is configured to compare new requests for content issued by the users within the cohort over a period to the forecast. The anomaly detector 130 is further configured to detect an anomaly when the new requests deviate from the forecast and output an indication of the anomaly. Example actual interest is shown in FIG. 5, in which an anomaly is visible.
[0051] Referring back to FIG. 1, an anomaly may be outputted as a response 108 to a query 106 of the cohort database 104. The response 108 may indicate a user identifier of each user that exhibits the anomaly. The response 108 may also indicate the nature of the anomaly, such as increased interest in the category or decreased interest in the category. A scale may also be provided, such as a percentage compared to expected, e.g., 40% increased interest over expected. The anomaly may be useful in the short term. Content delivery entities may use this information to tailor the amount and types of content to publish. Advertisers may use this information to target advertisements to certain users and / or adjust advertisement prices.
[0052] As shown in FIG. 6, the anomaly detector 130 may be configured to perform a normal analysis detect the anomaly.
[0053] In view of the above, it should be apparent that that techniques discussed herein allow for a more accurate understanding of user interaction with content delivery via computer networks.This may allow increase accuracy in tailoring content to users, so as to save network resources in unnecessarily delivering content that is of low or no interest.
Claims
Claims1. A non-transitory machine-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to collectively: monitor requests for content issued by user computers to a computer network; categorize the requests for content into categories of content; group users of the user computers into cohorts; and for a particular category of content: generate a periodic forecast of future requests for content by a cohort over a period using time-series clustering; compare new requests for content issued by the cohort within the period to the forecast; detect an anomaly when the new requests deviate from the forecast; and output an indication of the anomaly.
2. The non-transitory machine -readable medium of claim 1 , wherein the instructions are to apply an error, trend, seasonality (ETS) model to generate the periodic forecast of future requests for content.
3. The non-transitory machine -readable medium of claim 1, wherein the period is selected to encapsulate cyclic user behavior.
4. The non-transitory machine -readable medium of claim 1 , wherein the period is weekly.
5. The non-transitory machine -readable medium of claim 1, wherein the instructions are to perform a normal analysis detect the anomaly.
6. The non-transitory machine -readable medium of claim 1 , wherein the instructions are to categorize the requests for content by applying a machine learning model.
7. A system comprising: one or more computer servers, each server including memory and one or more processors, the one or more computer servers connected to a computer network and configured to: monitor requests for content issued by user computers to the computer network; categorize the requests for content into categories of content; group users of the user computers into cohorts; and for a particular category of content: generate a periodic forecast of future requests for content by a cohort over a period using time-series clustering; compare new requests for content issued by the cohort within the period to the forecast; detect an anomaly when the new requests deviate from the forecast using; and output an indication of the anomaly.
8. The system of claim 7, wherein the one or more computer servers are configured to apply an error, trend, seasonality (ETS) model to generate the periodic forecast of future requests for content.
9. The system of claim 7, wherein the period is selected to encapsulate cyclic user behavior.
10. The system of claim 7, wherein the period is weekly.
11. The system of claim 7, wherein the one or more computer servers are configured to perform a normal analysis detect the anomaly.
12. The system of claim 7, wherein the one or more computer servers are configured to categorize the requests for content by applying a machine learning model.
Citation Information
Patent Citations
Network anomaly detection
US20170099311A1
Deep recurrent neural network for cloud server profiling and anomaly detection through DNS queries
US20190141067A1
Anomaly detection in data protection operations
US20210073097A1