Using machine learning models to detect the leakage of sensitive browsing information.

The sensitivity detection system uses machine learning models to identify and mitigate sensitive user information leakage, enhancing accuracy and efficiency in preventing unintended disclosure and resource waste.

JP2026509398APending Publication Date: 2026-03-19MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-05
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing systems fail to accurately detect and prevent the leakage of sensitive user information during browsing activities, leading to unintended disclosure of personal or sensitive information, and subsequent delivery of discriminatory or predatory digital content.

Method used

A sensitivity detection system utilizing machine learning models, such as decision tree classification models, to identify and mitigate the leakage of sensitive user information by generating synthetic training data, training models to detect sensitivity classifications, and performing mitigation actions.

Benefits of technology

Accurately detects and prevents the leakage of sensitive user information, improving computational efficiency and reducing wasteful resource usage by modifying user profiles and digital content delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509398000001_ABST
    Figure 2026509398000001_ABST
Patent Text Reader

Abstract

This disclosure relates to a sensitivity detection system that accurately and efficiently determines when information based on a user's browsing activity unintentionally reveals personal or other sensitive information about the user. For example, the sensitivity detection system generates and utilizes machine learning models for sensitivity detection to accurately detect when sensitive user information is leaked from a set of user information, such as a user profile. In addition, if it determines that sensitive user information has been revealed, the sensitivity detection system often takes mitigation actions to prevent and / or reduce the undesirable disclosure of sensitive user information.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Background

[0001] In recent years, significant advancements in hardware and software have occurred in the field of computing devices, particularly in the management of user information and digital content. For example, individuals using computing devices are increasingly being provided with digital content, including requested content and uncommitted content, when browsing online. In any case, since people who guard sensitive information do not want personal or sensitive information to be leaked or inadvertently shared, individuals expect their privacy to be maintained. Unfortunately, many existing systems for managing user information and providing digital content do not have accurate or flexible safeguards to determine when or how they are leaking sensitive user information. As a result, many existing systems generate and share user information that inadvertently leaks or otherwise reveals personal or sensitive information about individuals.

[0002] Brief Description of the Drawings

[0002] The detailed description provides one or more implementations with additional specificity and detail using the accompanying drawings, as briefly described below.

Brief Description of the Drawings

[0003] [Figure 1]

[0003] An exemplary overview for implementing a sensitivity detection system for detecting and reducing the leakage of sensitive user information according to one or more implementations is shown. [Figure 2]

[0004] An exemplary computing system environment in which a sensitivity detection system according to one or more implementations is implemented is shown. [Figure 3]

[0005] This document provides an example process flow for generating user profiles using a profile generation model, with one or more implementations. [Figure 4A]

[0006] This document illustrates an exemplary process flow for generating synthetic training data to train a sensitivity detection model using one or more implementations. [Figure 4B]

[0006] An exemplary process flow for generating synthetic training data for training a sensitivity detection model is shown in one or more implementation forms. [Figure 4C]

[0006] An exemplary process flow for generating synthetic training data for training a sensitivity detection model is shown in one or more implementation forms. [Figure 5A]

[0007] An exemplary block diagram is shown for training and utilizing a sensitivity detection model in one or more implementation forms. [Figure 5B]

[0007] An exemplary block diagram is shown for training and utilizing a sensitivity detection model in one or more implementation forms. [Figure 6A]

[0008] This document provides an example process flow for performing mitigation actions based on the detection of sensitive user information leaks, using one or more implementations. [Figure 6B]

[0008] An exemplary process flow for performing mitigation actions based on the detection of leakage of sensitive user information is shown in one or more implementation forms. [Figure 6C]

[0008] An exemplary process flow for performing mitigation actions based on the detection of leakage of sensitive user information is shown in one or more implementation forms. [Figure 7]

[0009] This describes an exemplary set of actions for determining a ranked list of related monitoring incident tickets corresponding to a fault ticket, using one or more implementations. [Figure 8]

[0010] This shows exemplary components included in a computer system. [Modes for carrying out the invention]

[0004] Detailed explanation

[0011] This disclosure describes a sensitivity detection system that accurately and efficiently determines when information based on a user's browsing activity unintentionally reveals personal or other sensitive information about the user. For example, the sensitivity detection system generates and utilizes machine learning models for sensitivity detection to accurately detect when sensitive user information is leaked from a set of user information, such as a user profile. In addition, if it determines that sensitive user information has been revealed, the sensitivity detection system often takes mitigation actions to prevent and / or reduce the undesirable disclosure of sensitive user information.

[0005]

[0012] As background information, when users browse the web and perform other actions on online services, their user information is often recorded and subsequently shared. For example, web cookies, trackers, and browser fingerprints store user information and share it with websites and web services. In addition, this user information is often used to generate user profiles, which include a set of descriptive labels that indicate the user's characteristics and attributes. Furthermore, websites and web services use user profiles to deliver personalized digital content to users.

[0006]

[0013] In some cases, digital content provided to users enhances the user experience by delivering content that users desire. However, in other cases, digital content provided to users is discriminatory or predatory. For example, users are presented with digital content such as advertisements that target them based on race, gender, medical conditions, or other sensitive issues that the user wishes to keep private. In these cases, benign information from the user profile unintentionally reveals sensitive information about the user.

[0007]

[0014] Accordingly, this specification describes a sensitivity detection system that utilizes machine learning models to accurately detect when user information, such as user profiles, unintentionally reveals personal or other sensitive user information. In addition, the sensitivity detection system provides a range of mitigation actions and other tools to prevent inappropriate digital content from being delivered to users, from system-level solutions to user-level solutions.

[0008]

[0015] For illustrative purposes, in various implementations, a sensitivity detection system identifies a user profile generated by a profile generation model based on the user's browsing activity. In addition, the sensitivity detection system provides the user profile to a sensitivity detection machine learning model trained to classify users into sensitivity classifications based on their user profiles in order to determine whether the user profile should be classified as a specific sensitive topic. Furthermore, if the user profile is found to leak sensitive information, the sensitivity detection system provides the profile generation model with mitigation actions, such as instructing the model to decouple the given user profile from a specific sensitivity classification.

[0009]

[0016] As stated above, the implementations of this disclosure solve one or more of the problems described above and other problems in the art. The system, computer-readable medium, and method utilize a sensitivity detection system to determine when a user's user profile contains descriptive labels that inadvertently reveal personal or sensitive information about the user. In fact, the sensitivity detection system trains and utilizes machine learning models, such as decision tree classification models, to determine when a user profile inadvertently associates the user with a sensitive topic. In some implementations, the sensitivity detection system also indicates one or more labels or combinations of labels from the user profile that influenced the sensitivity classification result.

[0010]

[0017] As described herein, sensitivity detection systems offer several technical advantages in terms of computational accuracy and efficiency compared to existing computing systems. In fact, sensitivity detection systems offer several practical applications that provide benefits and solve problems related to detecting data sources (e.g., reputable data aggregators) that inadvertently reveal sensitive user information, and mitigating future information leaks.

[0011]

[0018] For illustrative purposes, a sensitivity detection system accurately determines when a user profile indicates or suggests sensitive user information. For example, by training and utilizing a sensitivity detection machine learning model, a sensitivity detection system accurately determines when a user profile is leaking sensitive information about a given user. In various implementations, sensitivity detection systems improve the accuracy of sensitivity detection machine learning models by intelligently creating data to train and tune the machine learning models. As further provided below, since sensitive-based user data is largely unavailable, a sensitivity detection system synthesizes usage data by generating algorithms that accurately mimic the user's real-world behavior without unnecessarily exposing the user's privacy.

[0012]

[0019] In addition, by utilizing sensitivity detection machine learning models, the sensitivity detection system can determine when and / or how the profile generation model generates user profiles that leak sensitive information. For example, in various implementations, the sensitivity detection system utilizes sensitivity detection machine learning models to identify combinations, characteristics, and / or sequences of descriptive labels contained in a user's profile that may unintentionally reveal sensitive information about the user. In response, the sensitivity detection system provides the profile generation model with instructions to decouple a given combination of descriptive labels from specific browser activities.

[0013]

[0020] Furthermore, sensitivity detection systems improve computational efficiency by rectifying and / or preventing the leakage of sensitive user information. For example, if a user profile leaks sensitive information, a digital content provider will provide the user with unnecessary and unwanted digital content. This unnecessary and unwanted digital content wastes computing resources and network bandwidth across multiple computing devices. Therefore, if a user profile is determined to potentially leak sensitive information, the sensitivity detection system employs several actions to mitigate the user profile's leakage of sensitive information and prevent the waste of computing resources. In fact, as will be further described below, these mitigation actions range from changes to the profile generation model to instructions and modifications on the user client device.

[0014]

[0021] For illustrative purposes, in one example, a sensitivity detection system facilitates data aggregators in generating better user profiles that prevent sensitive user information from being revealed. In another example, a sensitivity detection system facilitates digital content providers in refraining from providing digital content based on leaked sensitive user information. As a further example, a sensitivity detection system provides measures to offset user actions that would otherwise affect a user's user profile in relation to revealing sensitive user information. In addition, a sensitivity detection system helps users modify their browsing activity to better protect themselves from malicious actors.

[0015]

[0022] As described above, the present disclosure uses various terms to describe the features and advantages of one or more of the described implementations. For the purposes of the description, the present disclosure describes a sensitivity detection system in the context of a network. In the present disclosure, a "network" is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. The network can include not only private networks but also public networks such as the Internet. When information is transferred or provided to a computer via a network or another communication connection (either hardwired, wireless, or a combination of hardwired or wireless), the computer appropriately regards that connection as a transmission medium. The transmission medium can be used to carry program code means necessary in the form of computer-executable instructions or data structures, and can include networks and / or data links that are accessible to general-purpose or special-purpose computers. The above combinations are also included within the scope of computer-readable media.

[0016]

[0023] As an example, the term "user identifier" refers to an identifier of a user associated with one or more client devices. For simplicity of explanation, the term "user" refers to the user's user identifier. For example, when a user performs an action that is captured, the computing device detects that action and associates it with the user's user identifier. In many implementations, the term "user's browsing activity" (or simply "browsing activity") refers to the detected digital actions performed by the user with respect to websites and web services, and these actions are stored by the computing device in association with the user's user identifier. Similarly, the term "sensitive browsing activity" represents the user browsing a website associated with a sensitive topic or using a web service associated with a sensitive topic.

[0017]

[0024] In various implementations, the profile generation model generates a user profile from browsing activities associated with a user identifier. For example, the profile generation model generates a user profile by determining a set of descriptive labels for assignment to the user identifier from a larger set of possible descriptive labels, where the descriptive labels are based on the user's browsing activities. In some implementations, the user profile includes additional attributes, characteristics, and / or information about the user. In one or more implementations, the user profile is a user advertising profile generated by an advertising profile generation model.

[0018]

[0025] In the present disclosure, the terms "sensitive topic" and "sensitivity classification" refer to topics, labels, or issues that a user desires to protect personally and not have revealed. For example, a sensitive topic is considered sensitive due to the potential for negative consequences such as social disgrace, discrimination, or invasion of privacy. Examples of sensitive topics include medical conditions, political affiliations, race, sexual orientation, personal financial status, past traumas, emotionally impactful topics, and non-mainstream creeds.

[0019]

[0026] As an additional example, in this specification, the term “machine learning model” refers to a computer model or computer representation that can be tuned (e.g., trained) based on an input to approximate an unknown function. For example, machine learning models may include, but are not limited to, decision trees (e.g., gradient-boosted decision trees or decision tree classification models), transformer models, sequence-to-sequence models, neural networks (e.g., convolutional neural networks or deep learning models), regression-based models (e.g., quantile or linear), random forest models, clustering models, support vector learning models, Bayesian network models, principal component analysis models, or combinations thereof. In addition, machine learning models include deep learning models and / or shallow learning models.

[0020]

[0027] For example, a sensitivity detection system generates and utilizes a sensitivity detection machine learning model that determines one or more identifiable sensitivity classifications (e.g., sensitivity topics) from a vulnerable user profile. In various implementations, the sensitivity detection machine learning model is a decision tree model. In some implementations, the sensitivity detection machine learning model outputs a decision path for classifying a user profile into a specific sensitivity classification and / or a combination of descriptive labels within the user profile that resulted in that specific sensitivity classification.

[0021]

[0028] Additional details relating to exemplary implementations of the sensitivity detection system are described in relation to the following figures. For example, Figure 1 shows an exemplary overview of one or more implementations of a sensitivity detection system for detecting and initiating the mitigation of sensitive user information leaks. As illustrated, Figure 1 shows a series of actions 100 performed by the sensitivity detection system on a cloud computing system in many examples.

[0022]

[0029] For illustrative purposes, a series of actions 100 includes action 102, which generates training data by simulating the user's browsing habits. As mentioned above, to protect user privacy, the sensitivity detection system operates without the need to use actual user data. Instead, the sensitivity detection system simulates the user's browsing activity in a way that intelligently mimics the user's behavior. Additional details regarding the sensitivity detection system that generates training data are provided below in relation to Figures 4A-4C.

[0023]

[0030] As illustrated, a series of actions 100 includes action 104, which trains a sensitivity classification model to determine sensitivity classifications from a user profile. For example, in various implementations, the sensitivity detection machine learning model trains a sensitivity detection model, such as a sensitivity detection machine learning model and / or a sensitivity detection neural network, to point out sensitivity classifications that the user profile may unintentionally indicate. In various implementations, the sensitivity detection system uses the generated training data to train and refine the sensitivity detection model. Additional details regarding the training of the sensitivity detection machine learning model are provided below in relation to Figure 5A.

[0024]

[0031] As also illustrated, a series of actions 100 includes an action 106 that uses a profile generation model to generate a user profile (a user identifier associated with the user). For example, a sensitivity detection system and / or another system utilizes a profile generation model to convert a given user's browsing activity into a user profile. Additional details regarding user profile generation are provided below in relation to Figure 3.

[0025]

[0032] Figure 1 also shows that a series of actions 100 may include action 108 in which a sensitivity detection model determines that the user profile leaks sensitive information. For example, using a sensitivity detection model, the sensitivity detection system determines that a user profile generated for a given user results in one or more sensitivity classifications. In other words, the user profile contains one or more combinations of benign labels that unintentionally reveal personal or sensitive user information. Additional details regarding determining sensitivity classifications from user profiles using a sensitivity detection machine learning model are provided below in relation to Figure 5B.

[0026]

[0033] In addition, the sequence of actions 100 includes actions 110 that perform mitigation actions to reduce sensitivity leaks in user profiles. In particular, if the sensitivity detection system detects that a given user's user profile leaks sensitive information, it takes one or more mitigation actions to prevent future user profile leaks. As illustrated, in some implementations, the sensitivity detection system provides information to the profile generation model to patch the leak. In some examples, the sensitivity detection system notifies the user when the user profile reveals potentially sensitive information.

[0027]

[0034] In addition, the sensitivity detection system may modify the user's tracking settings on the client device. Furthermore, in various cases, the sensitivity detection system may conceal the user's browsing activity by injecting artificial counteracting browsing activity. Also, in some implementations, the sensitivity detection system may notify the digital content provider about the sensitivity leak to prevent the digital content provider from inadvertently discriminating based on the revealed sensitive information. Additional details regarding the sensitivity detection system that performs mitigation actions are provided below in relation to Figures 6A-6C.

[0028]

[0035] With a general overview of the sensitivity detection system in place, additional details regarding the components and elements of the sensitivity detection system are provided. For illustrative purposes, Figure 2 provides an exemplary diagram of the sensitivity detection system. In particular, Figure 2 shows an exemplary computing system environment in which the sensitivity detection system is implemented in one or more implementation forms. Figure 2 shows an exemplary arrangement and configuration within the computing system environment 200, but other arrangements and configurations are possible.

[0029]

[0036] As illustrated, Figure 2 includes a server device 202, a client device 230, a resource device 240, and a web server device 250, each connected by a network 260. Additional details regarding these and other computing devices are provided below in relation to Figure 8. In addition, Figure 8 also provides additional details regarding the network, such as the illustrated network 260.

[0030]

[0037] The computing system environment 200 includes a server device 202 having a user management system 204. In various implementations, the user management system 204 manages user information, including storing user information associated with user identifiers, providing communications and other digital content to users, protecting user privacy, and / or performing other functions. As shown in the figure, the user management system 204 includes a sensitivity detection system 206. In some implementations, the sensitivity detection system 206 is located outside the user management system 204.

[0031]

[0038] As described above, the sensitivity detection system 206 protects users by accurately detecting the leakage of personal and sensitive user information and protecting users from such leakage. In many implementations, the sensitivity detection system 206 utilizes a sensitivity detection model to detect when user profiles and other user information unintentionally reveal sensitive user information.

[0032]

[0039] As illustrated, the sensitivity detection system 206 includes various components and elements implemented in hardware and / or software. For example, the sensitivity detection system 206 includes a user data simulation manager 210 that generates synthetic user browser data (e.g., part of user browser data 222). In addition, the sensitivity detection system 206 includes a sensitivity detection manager 212 that trains and utilizes a sensitivity detection model 226 to detect sensitivity classifications from user profiles 224. The sensitivity detection system 206 also includes a sensitivity response manager 214 that provides one or more mitigation actions when sensitive user information is leaked. Furthermore, the sensitivity detection system 206 includes a storage manager 220 for storing data corresponding to the sensitivity detection system 206. The functionality of the components is further described below.

[0033]

[0040] In addition, the computing system environment 200 includes a client device 230 having a client application 232. In various implementations, the client device 230 is associated with a user having one or more user identifiers. In many implementations, the user (i.e., represented by the user identifier) ​​interacts with a server device 202 (e.g., a user management system 204 and / or a sensitivity detection system 206), a resource device 240, and / or a web server device 250 to access content and / or services. The computing system environment 200 may include any number of client devices. Also, as illustrated, the client device 230 includes a client application 232. For example, the client application 232 is a web browser application, a mobile application, or another type of application that accesses internet-based content to access and receive digital content. In some implementations, the client device 230 includes a plug-in associated with a sensitivity detection system 206 that communicates with the client application 232 to perform mitigation actions. In some implementations, a portion of the sensitivity detection system 206 is incorporated into a client application 232 to perform mitigation actions.

[0034]

[0041] As illustrated, the computing system environment 200 includes a resource device 240 containing a profile generation model 242. In one or more implementations, the sensitivity detection system 206 accesses user profiles 224 generated by the profile generation model 242, which may be operated by another system (e.g., a data aggregator system). In some implementations, the profile generation model 242 is part of the sensitivity detection system 206 and / or is internally connected to the sensitivity detection system 206.

[0035]

[0042] In addition, the computing system environment 200 includes a web server device 250 having web services 252. For example, the web services 252 include a website for the user to browse via a client application 232. In some implementations, information from the web services 252 is monitored and stored on the client device 230 in relation to a user identifier (e.g., browsing activity). In many implementations, the user consents to the tracking and / or storage of user data and has complete control over such user data.

[0036]

[0043] With the foundation of the sensitivity detection system 206 in place, we will now describe additional details regarding the various functions of the sensitivity detection system 206. As a brief roadmap, Figure 3 relates to the generation of user profiles using a profile generation model. Figures 4A-4C relate to the generation of synthetic training data that protects user privacy, Figures 5A-5B relate to the generation and use of a sensitivity detection machine learning model to determine when a user profile leaks sensitive user information, and Figures 6A-6C relate to the sensitivity detection system taking mitigation actions when a leak of sensitive user information is detected.

[0037]

[0044] As described above, Figure 3 relates to the generation of a user profile using a profile generation model. In particular, Figure 3 shows an exemplary process flow for generating a user profile using a profile generation model in one or more implementation forms. As illustrated, Figure 3 includes a series of actions 300 performed by the sensitivity detection system 206 and / or other systems.

[0038]

[0045] As illustrated, a series of actions 300 includes an action 302 that identifies user browser data based on the user's browsing activity. For example, when a user interacts with content and web services via a client device, the web services and / or client device monitor the user's actions and associate them with user identifiers. In other words, as a user browses the web, a set of actions and labels is accumulated based on each website they visit, the articles they read, the health websites they visit for information, the images they view, the videos they watch, and the products they buy. In various implementations, user browser activity is stored in the form of internet cookies, variables, browser or device digital fingerprints, and / or other trackers.

[0039]

[0046] Furthermore, as illustrated, the series of actions 300 includes an action 304 that generates a user profile using a profile generation model. For example, another system, such as a sensitivity detection system 206 or a data aggregator system, receives, identifies, or otherwise accesses the user's browsing activity 306 associated with a user identifier and provides it to the profile generation model 310. The profile generation model 310 generates a user profile 308 (e.g., an advertising user profile) based on the user's browsing activity 306. As already mentioned, the user profile 308 often includes a set of descriptive labels representing the corresponding user identifier.

[0040]

[0047] The profiling model 310 can generate user profiles using rules, features, weights, and / or parameters. For example, the profiling model 310 is a heuristic model that translates a user's browsing activities 306 into descriptive labels based on a set of rules. In another example, the profiling model 310 is a machine learning model trained to learn and encode latent features from a user's browsing activities 306 based on tuned weights and parameters, and then decode these latent features into descriptive labels.

[0041]

[0048] In various implementations, the profile generation model 310 generates a user profile 308 that includes a set of descriptive labels (e.g., a subset) selected from a larger set of potential descriptive labels. For example, some of the descriptive labels may include user information corresponding to interests and hobbies, income, car or home ownership, pet status, family relationships, shopping habits, etc. The profile generation model 310 can generate a variety of user profiles, ranging from a few descriptive labels to thousands. In various implementations, the labels may include a hierarchical structure (e.g., "Hobbies and Interests > Exercise > Running and Jogging" or "Hobbies and Interests > Games > Board Games").

[0042]

[0049] As illustrated, a series of actions 300 includes an action 312 that provides content to the user based on a user profile. For example, a sensitivity detection system, data aggregation system, digital content distribution system, or another system uses the user's browsing activity 306 to identify the user and provide digital content to the user via a client device. As illustrated, backpack supplies are provided to the user based on a user profile.

[0043]

[0050] As described above, existing computer systems may provide users with discriminatory or predatory content because a user's user profile may leak personal or sensitive information. For example, digital content may be provided to a user based on a medical condition or other sensitive issue that the user does not wish to disclose. For simplicity of explanation, this specification refers to a user who has a medical condition that they wish to keep private.

[0044]

[0051] As described above, Figures 4A-4C relate to the generation of synthetic training data that protects user privacy. Generally, Figures 4A-4C illustrate the generation and simulation of web browsing activity. For example, to determine whether information about visits to sensitive websites is being leaked through seemingly harmless labels in a user's profile, the sensitivity detection system 206 trains a sensitivity detection machine learning model to discover how and when the user profile is being leaked. In particular, Figures 4A-4C show exemplary process flows for generating synthetic training data to train a sensitivity detection model in one or more implementation forms.

[0045]

[0052] To generate training data, the sensitivity detection system 206 synthetically generates user browsing activity. As part of simulating this synthetic data, the sensitivity detection system 206 determines the various web resources to access in order to accurately mimic a real-world user. To achieve this, the sensitivity detection system 206 identifies both frequently visited and popular websites, as well as websites associated with sensitive topics. Part of this process is shown in Figure 4A.

[0046]

[0053] Figure 4A includes two data streams. The first data stream corresponds to frequently visited websites and includes popular websites 402. For example, popular websites 402 include top search and / or visited websites in a given region or network. For example, the sensitivity detection system 206 identifies popular websites 402 based on website traffic. From popular websites 402, the sensitivity detection system 206 selects some of the top sites 404. As will be discussed later, the top sites 404 can serve as controls when generating training data.

[0047]

[0054] In various implementations, to determine the list of sensitive sites 410, the sensitivity detection system 206 identifies web trends 406 of sensitive topics. For example, the sensitivity detection system 206 identifies popular search terms associated with each sensitive topic (e.g., sensitive groups and / or sensitive behaviors) included in the list of sensitive topics. In various implementations, the sensitivity detection system 206 accesses one or more web services that track trending terms associated with these sensitive topics. In some implementations, the sensitivity detection system 206 uses a natural language processing model or another type of topic grouping machine learning model to identify terms associated with each sensitive topic from a database or other resources.

[0048]

[0055] Next, the sensitivity detection system 206 uses those terms to perform web queries and retrieves or captures a list of resulting web resources (e.g., websites) for each sensitive topic (or part thereof), which are presented as web search results 408. In some examples, the sensitivity detection system 206 removes duplicate entries and / or identifies a threshold number of search results for each sensitive topic. In addition, in various implementations, the sensitivity detection system 206 ranks the list of websites from the web search results 408 based on one or more metrics (e.g., traffic rank, search score, statistically unlikely phrase, sensitivity score, and / or other metrics) to generate a list of sensitive sites 410 for each sensitive topic.

[0049]

[0056] Once the top site 404 and sensitive site 410 are identified, the sensitivity detection system 206 can proceed to generate synthetic browsing activity. For illustrative purposes, Figure 4B shows a series of actions 420 for the sensitivity detection system 206 to generate synthetic browsing activity. As shown in action 422, the sensitivity detection system 206 generates synthetic users. In various implementations, as part of action 422 for generating synthetic users, the sensitivity detection system 206 generates a list of user identifiers for a corresponding number of synthetic users.

[0050]

[0057] In some examples, the sensitivity detection system 206 assigns parameters to each user, such as the type of browser application, geographical location, and / or demographic information. In addition, in many examples, the sensitivity detection system 206 also assigns browsing activity to the synthetic user, such as the length of the browsing session, the time of day or week during which browsing occurs, the number of days or months during which browsing occurs, the amount of mouse movement during browsing, and / or the duration of browsing at each site.

[0051]

[0058] In addition, the sequence of actions 420 includes action 424, which determines a sensitivity percentage of 0-100% for each synthetic user and each sensitivity topic. In one or more implementations, the sensitivity detection system 206 divides users into two initial groups, including a control or baseline group that does not visit any of the sensitive sites 410, and a non-baseline group that visits at least some of the sensitive sites 410. Alternatively, in some implementations, the sensitivity detection system 206 assigns each user to one of zero, one, or more sensitive topics. For example, the sensitivity detection system 206 assigns users to sensitive topics according to recent statistical data that match real-world ratios.

[0052]

[0059] As shown in the figure, the sensitivity detection system 206 assigns a 10% sensitivity percentage to sensitive topic A, a 25% sensitivity percentage to sensitive topic B, and a 0% sensitivity percentage to sensitive topic C to synthetic user 1. In this example, the remaining 65% is assigned to non-sensitive topics. Using different methods, the sensitivity detection system 206 assigns a sensitivity percentage between 0% and 100% to each sensitivity topic assigned to the synthetic user.

[0053]

[0060] For each synthetic user assigned to one or more sensitive topics, the sensitivity detection system 206 can determine a sensitivity percentage. For example, as shown in relation to action 424, the sensitivity detection system 206 assigns a value between 0 and 100% to each of the three illustrated sensitive topics. In one or more implementations, the sensitivity percentage or amount is assigned randomly. In some implementations, the sensitivity detection system 206 uses a non-random or uniform distribution when assigning sensitivity levels to sensitive topics.

[0054]

[0061] As illustrated, a series of actions 420 includes an action 426 that simulates a user's browsing activity between top sites 404 and sensitive sites 410 based on their sensitivity percentage. For example, for each synthetic user, the sensitivity detection system 206 visits top sites 404 and sensitive sites 410 according to their assigned sensitivity percentage. For illustrative purposes, in the case of synthetic user 1, the sensitivity detection system 206 visits a website corresponding to sensitive topic A with a 10% probability, a website corresponding to sensitive topic B with a 25% probability (e.g., X% = 10% + 25% = 35%), and a website corresponding to a non-sensitive topic with a 65% probability (e.g., 100% - X% or 100% - 35%).

[0055]

[0062] In the second method, each sensitivity topic assigned to a synthetic user is given an independent sensitivity percentage ranging from 0% to 100% (e.g., X%), the sensitivity detection system 206 selects several websites (e.g., 20 sites or 5,000) and visits both websites corresponding to a given sensitivity topic and websites corresponding to a non-sensitive topic, according to the assigned sensitivity percentage (e.g., X% and 100%-X%).

[0056]

[0063] When visiting a website on a non-sensitive topic, in various implementations, the sensitivity detection system 206 browses websites from the top sites 404. For example, the sensitivity detection system 206 selects one or more websites to visit and / or browse from the top sites 404. In various implementations, the sensitivity detection system 206 randomly selects websites to visit. In some implementations, the sensitivity detection system 206 weights its selections based on website rankings. For example, the sensitivity detection system 206 browses higher-ranked websites more frequently than lower-ranked websites.

[0057]

[0064] Similarly, when visiting a website corresponding to a given sensitive topic, the sensitivity detection system 206 uses the sensitive site 410 corresponding to the given sensitive topic to select the website to visit. In addition, the sensitivity detection system 206 can alternate between websites corresponding to non-sensitive topics and websites corresponding to one or more sensitive topics.

[0058]

[0065] As illustrated, action 426 includes simulating the user's browsing activity between the top site 404 and the sensitive site 410. Thus, in various implementations, the sensitivity detection system 206 mimics the real-world browsing behavior and browsing activity of each synthetic user when visiting selected sites. For example, the sensitivity detection system 206 utilizes instructions that simulate the browsing habits of an actual user when visiting a website, such as mouse and scroll movements, dwell time, selection of links within a website, and / or other browsing behaviors described above. In this way, each synthetic user generates browsing activity, so the data collected for the synthetic users is the same as, or very similar to, that of an actual user's browsing activity.

[0059]

[0066] For further explanation, Figure 4C shows an example of a virtual environment 430 in which the sensitivity detection system 206 operates a synthetic user. For example, the host device includes a virtual machine that implements the virtual environment 430. As shown in the figure, the virtual environment 430 includes the user data simulation manager 210 described above.

[0060]

[0067] In various implementations, the user data simulation manager 210 receives a site list 432 that includes both the top site 404 and the sensitive site 410. In addition, the user data simulation manager 210 includes a command file 434 that contains various instructions for creating and managing synthetic users.

[0061]

[0068] In addition, the virtual environment 430 represents two simulated browsers (simulated browser A (436a) and simulated browser B (436b)). In various implementations, the sensitivity detection system 206 assigns a unique synthetic user to each browser, with simulated browser A (436a) being assigned the first synthetic user and simulated browser B (436b) being assigned the second synthetic user. In various implementations, the sensitivity detection system 206 generates an independent simulated browser for each synthetic user, thereby enabling each synthetic user to maintain and save its own browser activity as it visits the assigned website over time. This method ensures that stored browser activity, such as internet cookies collected over time, matches that of a real-world user.

[0062]

[0069] As illustrated, the virtual environment 430 includes a local network 438 that enables the synthetic user to access Internet websites and web services, such as the target website 442. In addition, in various implementations, the synthetic user can be assigned to one or more proxies and / or virtual private networks (VPNs). For illustrative purposes, Figure 4C includes two VPN proxies (VPN Proxy A (440a) and VPN Proxy B (440b)). For example, each proxy allows the synthetic user to appear to be in a different geographical and / or network location in order to more accurately emulate a real-world user with a different network address (e.g., an Internet Protocol (IP) address).

[0063]

[0070] To illustrate with a non-limiting example, Sensitivity Detection System 206 generates 30,000 synthetic users who visit a mix of the top 500 popular sites and sensitive websites (based on 60 sensitive topics). Sensitivity Detection System 206 assigned each synthetic user a set of 500 sites to visit, divided into browsing sessions of varying lengths over a period of at least two months. In addition, Sensitivity Detection System 206 distributed the synthetic users across approximately 1,200 VPN proxies. Furthermore, Sensitivity Detection System 206 allowed cookies and other trackers to accumulate during browsing and maintained them between sessions. In some examples, Sensitivity Detection System 206 waited for a threshold time (e.g., one week or eight days) after the synthetic user's browsing activity was complete to allow browser activity to stabilize before generating the user's user profile.

[0064]

[0071] In this way, the sensitivity detection system 206 generates training data by simulating browser activity and generating user profiles from the simulated data. Furthermore, because the sensitivity detection system 206 knows which sensitive topics are associated with which synthetic users, it also generates ground truth data (e.g., ground truth sensitivity classification) corresponding to each user profile (e.g., training user profile) as part of the training data (e.g., how strongly each user profile correlates with one or more given sensitive topics).

[0065]

[0072] Figures 5A and 5B provide additional details for generating and utilizing sensitivity detection machine learning models. For example, Figure 5A corresponds to the action of training a sensitivity detection machine learning model. Figure 5B corresponds to the action of using the trained sensitivity detection machine learning model to determine sensitivity classifications from user profiles.

[0066]

[0073] As shown in the figure, Figure 5A includes training data 502, a sensitivity detection machine learning model 510, a user sensitivity classification 512, and a loss model 520. Also as shown in the figure, the training data 502 includes training user profiles 504 and the ground truth sensitivity classification 506 described above. For example, the training user profiles 504 include a set of descriptive benign labels that describe the corresponding users (e.g., synthetic users).

[0067]

[0074] In various implementations, the sensitivity detection machine learning model 510 is a decision tree machine learning model. For example, the sensitivity detection machine learning model 510 is a multiclass decision tree classifier that predicts sensitivity classification using a set of descriptive labels from a user profile as features. In some examples, each node of the decision tree machine learning model represents a binary decision about whether a given label exists in the user profile. In one or more implementations, the sensitivity detection system 206 trains a decision tree machine learning model with a depth of 90 and a minimum (or average) leaf size of 5. In addition, in some cases, the sensitivity detection system 206 utilizes a heterogeneous distribution metric as a partitioning algorithm when generating the sensitivity detection machine learning model 510.

[0068]

[0075] As illustrated, training data 502 is provided to the sensitivity detection machine learning model 510. In particular, training user profiles 504 are provided to the sensitivity detection machine learning model 510, which is trained to generate user sensitivity classifications 512 using descriptive labels as features. In fact, the sensitivity detection system 206 trains the sensitivity detection machine learning model 510 to accurately re-identify when a synthetic user is associated with one or more sensitive topics. In various implementations, the sensitivity detection system 206 trains the sensitivity detection machine learning model 510 to classify a given user profile of a simulated user with a specific sensitivity classification when the synthetic user has browsed a threshold number of websites associated with a sensitivity classification.

[0069]

[0076] During training, the loss model 520 determines the amount of loss or error by comparing the user sensitivity classification 512 with the corresponding sensitivity classification from the ground truth sensitivity classification 506. The loss model 520 may use one or more loss functions to determine the amount of loss, which is then fed back to the sensitivity detection machine learning model 510 as feedback 522 to adjust the model's weights, parameters, layers, and / or nodes. In this way, the sensitivity detection system 206 trains the sensitivity detection machine learning model 510 by backpropagation in an end-to-end manner until the model converges or meets another training criterion.

[0070]

[0077] Furthermore, by training a decision tree model (i.e., a decision tree machine learning model), the sensitivity detection system 206 can track one or more decision paths of descriptive labels taken to arrive at a particular sensitivity description. In addition, the sensitivity detection system 206 can identify which descriptive labels best represent a particular sensitive topic (e.g., regardless of the decision path). In this way, the trained decision tree model can accurately detect when a user profile leaks sensitive information, indicate which sensitive topics are being leaked, and determine which descriptive labels in the user profile encouraged, caused, or triggered the leakage of sensitive user information.

[0071]

[0078] In one case, a trained sensitivity detection machine learning model was found to achieve a re-identification accuracy of 77.4%, compared to a control / reference classifier using random assignment, which would achieve a re-identification accuracy of 2.08%. Furthermore, the trained sensitivity detection machine learning model can re-identify sensitive topics for 63% of users with 99% confidence, based on the average of five descriptive labels. As mentioned above, user profiles typically have hundreds to thousands of descriptive labels.

[0072]

[0079] In some cases, a decision tree model is generated, but in other cases, the sensitivity detection system 206 generates a different type of machine learning model. For example, the sensitivity detection system 206 uses the training data to generate a convolutional neural network. In fact, the sensitivity detection system 206 uses the training data to oversee the training of various types of machine learning models and / or neural networks. In this way, the sensitivity detection system 206 generates a sensitivity detection machine learning model 510 that can very accurately identify how a user's seemingly harmless interests relate to belonging to a sensitive group or having an interest in a sensitive topic.

[0073]

[0080] As shown in the figure, Figure 5B includes a given user profile 524, a trained sensitivity detection machine learning model 510', and a given user sensitivity classification 532. As shown in the figure, the trained sensitivity detection machine learning model 510' includes a sensitivity decision tree classification model 530.

[0074]

[0081] In various implementations, a given user profile 524 corresponds to an actual user, not a synthetic user. The sensitivity detection system 206 receives the given user profile 524 directly or directly from the browser activity of a given user. For example, the sensitivity detection system 206 receives the given user profile 524 from a profile generation model. In addition, the sensitivity detection system 206 uses a sensitivity decision tree classification model 530 to generate a given user sensitivity classification 532 for a given user. In fact, the given user sensitivity classification 532 may reveal one or more sensitivity classifications of a given user that the given user wishes to keep private.

[0075]

[0082] In some implementations, the sensitivity detection system 206 uses a trained sensitivity detection machine learning model 510' to classify a given user profile 524 as sensitive or non-sensitive to one or more sensitive topics. Specifically, the sensitivity detection system 206 provides a descriptive label of the given user profile 524 to a sensitivity decision tree classification model 530, which then classifies the given user profile 524 as non-sensitive or belonging to one or more sensitive topics based on the decision paths within the sensitivity decision tree classification model 530.

[0076]

[0083] Once the sensitivity classification of a given user (or group of users) is determined, the sensitivity detection system 206 may perform one or more mitigation steps to prevent future leaks of sensitive user information and / or to prevent users from being inappropriately targeted by digital content providers. As described above, Figures 6A–6C provide additional details regarding the sensitivity detection system performing mitigation actions.

[0077]

[0084] For illustrative purposes, Figure 6A includes a series of actions 600 performed by the sensitivity detection system 206. As illustrated, action 602 includes the sensitivity detection system 206 identifying a given user's sensitivity classification. For example, as previously stated, the sensitivity detection system 206 uses a sensitivity detection neural network to determine that a given user's user profile reveals sensitive user information that the given user wishes to keep private.

[0078]

[0085] As illustrated, a series of actions 600 includes an action 604 that generates and provides instructions based on identified sensitivity classifications. For example, the sensitivity detection system 206 determines whether to send one or more instructions, where to send the instructions, and what to include in the instructions. For example, the sensitivity detection system 206 decides to send a first set of information to a first target recipient and a different second set of information to a second target recipient. In some implementations, the recipient is a backend device that is instructed to make system-wide modifications. In one or more implementations, the recipient is a frontend device, such as a client device, that is instructed to make modifications that affect individual users.

[0079]

[0086] For illustrative purposes, in various implementations, the sensitivity detection system 206 generates instructions to provide to backend devices. As illustrated, a series of actions 600 includes action 606, which provides instructions to facilitate backend modifications. In fact, in various implementations, the sensitivity detection system 206 provides audits of backend devices and services with respect to sensitive topics that are prone to leakage. The audits enable backend devices and services to obtain accurate feedback that allows them to eliminate undesirable influences on the profiling process, which is otherwise a black box where the internal structure is difficult or impossible to discover. For example, a device receiving instructions can use them to modify the algorithms and / or rules that generate one or more descriptive topics to decouple them from sensitive topics. In this way, the system and models can ensure that they generate user profiles that do not inadvertently reveal sensitive topics about a given user.

[0080]

[0087] Furthermore, in many cases, the instructions provide real-time feedback on how to better generate user profiles that prevent the leakage of sensitive user information. For example, the sensitivity detection system 206 generates instructions that indicate information to cause a device implementing a profile generation model to modify one or more features to decouple one or more given user profiles (or subsets of descriptive labels) from a particular sensitivity classification. For instance, the profile generation model continues to modify the features, weights, associations, and / or connections of one or more descriptive labels with respect to a given user's browser activity until the profile generation model no longer generates a user profile for a given user that leaks particular sensitive information.

[0081]

[0088] In various implementations, instructions to the profile generation model cause the profile generation model to generate different sets of descriptive labels for the user profile. For example, the profile generation model generates a first set of descriptive labels for a given user profile that the sensitivity detection system 206 has determined to be leaking sensitive user information. Upon receiving instructions from the sensitivity detection system 206, the profile generation model generates a different second set of descriptive labels for the given user profile by changing how it determines descriptive labels and using the same browser activity as before. The sensitivity detection system 206 can then verify that the updated user profile for a given user does not leak sensitive user information.

[0082]

[0089] In some implementations, the instruction results in the removal of one or more descriptive labels. In fact, in some implementations, the instruction includes a specific sensitive topic that has been leaked, refers to a given user profile, and / or lists a set of descriptive labels that likely led to the leaked sensitive topic (e.g., one or more decision paths for classifying a given user profile into a specific sensitivity category).

[0083]

[0090] As illustrated, the sequence of actions 600 includes an action 608 that provides instructions to facilitate front-end modification. For example, the sensitivity detection system 206 may provide various messages that result in front-end changes on the client device. For example, the sensitivity detection system 206 may send instructions that automatically cause the client device to take action and / or allow the user to know the potential consequences of a particular action, as further provided below in Figures 6B and 6C.

[0084]

[0091] For illustrative purposes, Figure 6B shows the graphical user interface on a client device for the website 610 concerning medical disease X. The website 610 includes various links to further information about medical disease X, such as treatment links 612 related to medical disease X.

[0085]

[0092] A user may have a medical condition X, but does not want to disclose this personal information to others. However, based on the user's browser activity, the user's user profile may unintentionally, or nearly unintentionally, reveal to a third party that the user has a medical condition X. Therefore, in one or more implementations, the sensitivity detection system 206 utilizes a sensitivity detection machine learning model to proactively predict how the user's potential browsing activity may lead to the leakage of user information.

[0086]

[0093] For illustrative purposes, in response to detecting that the user is about to select a treatment link 612, the sensitivity detection system 206 determines how the additional action of visiting the treatment link 612 will affect the leakage of sensitive information regarding medical condition X. As illustrated, the sensitivity detection system 206 determines that visiting the treatment link 612 (e.g., an online resource) increases the likelihood by 15% that the user has medical condition X. In response, the sensitivity detection system 206 provides the user with a visual indication 614 of the treatment link 612 before the user selects it. In this way, the sensitivity detection system 206 performs mitigation actions (which will be further described below) that enable the user not to increase the likelihood of their sensitive user information being leaked.

[0087]

[0094] In one or more implementations, the sensitivity detection system 206 provides the user with a report on the user's current status regarding sensitive topics. For example, the sensitivity detection system 206 provides the user with a report showing the probability that one or more sensitive topics are leaked. Furthermore, the sensitivity detection system 206 can show trends and other statistics on how the probability has changed over time.

[0088]

[0095] As described above in several implementations, the sensitivity detection system 206 performs one or more mitigation actions to prevent and reduce the future leakage of sensitive topics. The sensitivity detection system 206 may perform these actions automatically or based on user selection. For illustrative purposes, Figure 6C shows a prompt 620 that changes the user's response to the possibility of sensitive user information being leaked, along with the option to perform mitigation actions.

[0089]

[0096] The sensitivity detection system 206 can perform one or more mitigation actions to prevent and reduce the future leakage of sensitive topics, which can be performed automatically or based on user selection. As shown in prompt 620 in Figure 6C, the sensitivity detection system 206 may perform action 622 to modify the tracking settings for selected browser activity. For example, the sensitivity detection system 206 can enable private or untracked browsing (e.g., untracked mode) for all of the user's browsing activity or browser activity corresponding to one or more sensitive topics, either directly or by communicating with a client application on the client device. For example, if a user visits website 610 and / or selects a therapeutic link 612, the sensitivity detection system 206 can enable untracked browsing (or provide a message to the client application to enable it) to prevent sensitive browsing activity data from being added to the user's browser activity used to generate the user profile of the user. In some cases, the sensitivity detection system 206 can also remove sensitive browser activity data from the user's browser activity (for example, removing some browser activity stored directly on the client device or via a client application). This helps protect the user's sensitive information and prevent unintended data leaks.

[0090]

[0097] As shown in the box in the lower right, the sensitivity detection system 206 may perform act 624 to conceal the user's user browser activity by injecting artificial browsing activity. For example, the sensitivity detection system 206 may generate, load, or otherwise acquire browsing activity that is the opposite of a sensitive topic or a generally irrelevant topic. For illustrative purposes, the sensitivity detection system 206 may conceal medical condition X by supplementing the user's browser activity with visits and interactions on websites related to exercise, vacation, news, or other opposite and / or irrelevant topics. As another example, the sensitivity detection system 206 may have the user revisit non-sensitive websites that they have previously visited, generate further browser activity from those non-sensitive websites, and / or inject that further browser activity (e.g., weighted by recency, frequency, and / or preference). In various implementations, the sensitivity detection system 206 supplements the user's browsing activity with default or generic browsing activity. In one or more implementations, the sensitivity detection system 206 provides the supplemented and / or artificial browsing activity to a client application on a client device for the client application to save it as if the user had generated the supplemented browsing activity.

[0091]

[0098] In some examples, the sensitivity detection system 206 times-stamps visits to non-sensitive websites around the same time that the user visits a website related to medical condition X. In various implementations, the sensitivity detection system 206 includes a significant amount of non-sensitive browser activity data (e.g., 3, 5, 10, or 50 times) to better conceal sensitive browser activity. In fact, the sensitivity detection system 206 may perform various actions to hide the user's sensitive browsing activity.

[0092]

[0099] Referring now to Figure 7, this figure shows an exemplary flowchart of a series of actions 700 for utilizing the sensitivity detection system 206 in one or more implementations. In particular, Figure 7 shows an exemplary series of actions for determining a ranked list of relevant monitoring incident tickets corresponding to a fault ticket in one or more implementations.

[0093]

[0100] Figure 7 illustrates actions in one or more implementations, but alternative implementations may omit any of the illustrated actions, add to any of the illustrated actions, rearrange any of the illustrated actions, and / or modify any of the illustrated actions. Furthermore, the actions in Figure 7 may be performed as part of a method (e.g., a computer implementation method). Alternatively, a non-temporary computer-readable medium may, when executed by a processing system with a processor, contain instructions that cause a computing device to perform the actions in Figure 7. In further implementations, a system (e.g., a processing system with a processor) may perform the actions in Figure 7.

[0094]

[0101] In one or more implementations, the system includes a given user profile for a given user identifier generated by a profile generation model that generates a user profile for a user identifier including one or more descriptive labels for each user ID; a sensitivity detection machine learning model trained to classify users into sensitivity classifications based on the user profile; at least one processor in a server device; and / or computer memory containing instructions that, when executed by at least one processor in a server device, cause the system to perform one or more operations or actions.

[0095]

[0102] As illustrated, a series of actions 700 includes an action 710 that identifies a user profile from a profile generation model. For example, in an exemplary implementation, action 710 includes identifying a given user profile generated by the profile generation model based on browsing activity associated with a given user identifier. In various implementations, action 710 includes generating a given user profile for a given user identifier by the profile generation model based on browsing activity.

[0096]

[0103] In one or more implementations, act 710 includes identifying a given user profile by providing a profile generation model with browsing activities associated with a given user identifier and / or receiving a given user profile that includes a first descriptive label subset for a given user profile from a set of descriptive labels. In some implementations, modifying one or more features of the profile generation model causes the profile generation model to use a second set of features to determine a second descriptive label subset for a given user profile from a set of descriptive labels.

[0097]

[0104] As further illustrated, a series of actions 700 includes an action 720 that provides a user profile to a sensitivity classification model. For example, in an exemplary implementation, action 720 includes providing a given user profile to a sensitivity detection machine learning model trained to classify users into sensitivity classifications based on the user profile. In some implementations, action 720 includes utilizing a decision tree model as the sensitivity detection machine learning model that provides a decision path for classifying a given user profile into a specific sensitivity classification. In some implementations, the decision path represents a combination of descriptive labels identified for a given user profile that resulted in a specific sensitivity classification, and / or the indication represents that the combination of descriptive labels resulted in the given user profile being classified into a specific sensitivity classification.

[0098]

[0105] As further illustrated, a series of actions 700 includes an action 730 that uses a sensitivity classification model to determine that a user profile belongs to a particular sensitivity classification. For example, in an exemplary implementation, action 730 includes determining that a given user profile is classified as a particular sensitivity classification (e.g., a particular sensitivity topic) by a sensitivity detection machine learning model. In various implementations, action 730 includes providing a given user profile to a sensitivity detection machine learning model in order to determine that a given user profile is classified as a particular sensitivity classification.

[0099]

[0106] As further illustrated, a series of actions 700 includes an action 740 that indicates to the profile generation model that a given user profile has a particular sensitivity classification. For example, in an exemplary implementation, action 740 includes providing the profile generation model with an instruction that, based on the particular sensitivity classification, a given user profile is classified as a particular sensitivity classification. In some implementations, action 740 includes providing a visual instruction to a client device associated with a given user that a given user profile has been determined to be associated with a particular sensitivity classification.

[0100]

[0107] In various implementations, action 740 causes the profile generation model to modify one or more features of the profile generation model to decouple a given user profile from a specific sensitivity classification. In one or more implementations, action 740 includes causing the profile generation model to determine different descriptive labels for a given user profile that have been previously determined, thereby causing the profile generation model to decouple a given user profile from a specific sensitivity classification.

[0101]

[0108] In some implementations, the set of actions 700 includes additional actions. For example, in certain implementations, the set of actions 700 includes actions that supplement the browsing activity of a given user identifier with artificial browsing activity that is not associated with a particular sensitivity classification of that user identifier, based on that particular sensitivity classification of the user identifier.

[0102]

[0109] In various implementations, a set of actions 700 includes generating a training dataset that simulates the browsing activity of users associated with multiple sensitivity classifications. In some implementations, a set of actions 700 includes generating an additional training dataset that simulates the additional browsing activity of control users not associated with multiple sensitivity classifications. In one or more implementations, generating a training dataset that simulates the browsing activity of users associated with multiple sensitivity classifications includes generating simulated users over several months who visit or browse a random number of websites associated with one or more sensitivity classifications, in addition to visiting and browsing additional websites not associated with one or more sensitivity classifications. In various implementations, a set of actions 700 includes generating a sensitivity detection machine learning model by tuning the sensitivity detection machine learning model to classify the user profiles of simulated users who visited more than a threshold number of websites associated with a particular sensitivity classification into one or more sensitivity classifications.

[0103]

[0110] In some implementations, a series of actions 700 includes actions that identify a potential interaction between a given user and an online resource, determine how the addition of the online resource will change the user's given user profile, and / or provide the user with additional instructions that performing the potential interaction with the online resource will change the classification status of the sensitivity classification, if the potential interaction will change the classification status of the sensitivity classification.

[0104]

[0111] In addition, the networks described herein may represent a network or combination of networks (such as the Internet, a corporate intranet, a virtual private network (VPN), a local area network (LAN), a wireless local area network (WLAN), a cellular network, a wide area network (WAN), a metropolitan area network (MAN), or a combination of two or more of the above networks) on which one or more computing devices can access the sensitivity detection system 206. In fact, the networks described herein may include one or more networks that use one or more communication platforms or technologies for transmitting data. For example, the network may include the Internet or other data links that enable the transport of electronic data between each client device and components of the cloud computing system (e.g., server devices and / or virtual machines on them).

[0105]

[0112] Furthermore, upon reaching various computer system components, program code in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to a non-transient computer-readable storage medium (device), or vice versa. For example, computer-executable instructions or data structures received via a network or data link can be buffered in random access memory (RAM) within a network interface module (NIC), and then finally transferred to the computer system RAM and / or a non-volatile computer storage medium (device) within the computer system. Therefore, it should be understood that non-transient computer-readable storage medium (device) can be included in computer system components that also utilize (or primarily utilize) the transmission medium.

[0106]

[0113] Computer executable instructions include, for example, instructions and data that, when executed by a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a specific function or set of functions. In some implementations, computer executable instructions are executed by a general-purpose computer to transform the general-purpose computer into a special-purpose computer that implements the elements of this disclosure. Computer executable instructions may be, for example, binary, intermediate format instructions such as assembly language, or source code. While the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or acts described above. Rather, the described features and acts are disclosed as exemplary forms that realize the claims.

[0107]

[0114] Figure 8 shows certain components that may be included in computer system 800. Computer system 800 can be used to implement various computing devices, components, and systems described herein. In this specification, “computing device” means an electronic component that performs a set of actions based on a programmed set of instructions. Computing devices include groups such as electronic components, client devices, and server devices.

[0108]

[0115] In various implementations, computer system 800 represents one or more of the client devices, server devices, or other computing devices described above. For example, computer system 800 may refer to a network, a cloud computing system, or various types of network devices capable of accessing data on another system. For example, a client device may refer to a mobile device such as a mobile phone, smartphone, personal digital assistant (PDA), tablet, laptop, or wearable computing device (e.g., a headset or smartwatch). A client device may also refer to a non-mobile device such as a desktop computer, a server node (e.g., from another cloud computing system), or another non-portable device.

[0109]

[0116] The computer system 800 includes a processing system including a processor 801. The processor 801 may be a general-purpose single-chip or multi-chip microprocessor (e.g., a Novel Reduced Instruction Set Computer (RISC) machine (ARM)), a special-purpose microprocessor (e.g., a Digital Signal Processor (DSP)), a microcontroller, a programmable gate array, etc. The processor 801 is sometimes referred to as a central processing unit (CPU). The illustrated processor 801 is just a single processor in the computer system 800 of Figure 8, but alternative configurations may use a combination of processors (e.g., ARM and DSP).

[0110]

[0117] The computer system 800 also includes a memory 803 that electronically communicates with the processor 801. The memory 803 may be any electronic component capable of storing electronic information. For example, the memory 803 may be embodied as random access memory (RAM), read-only memory (ROM), magnetic disk storage medium, optical storage medium, flash memory device in RAM, onboard memory included with the processor, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, etc. (including combinations thereof).

[0111]

[0118] Instruction 805 and data 807 may be stored in memory 803. Instruction 805 may be executable by processor 801 to implement some or all of the functionality disclosed herein. Execution of instruction 805 may involve the use of data 807 stored in memory 803. Any of the various examples of modules and components described herein may be implemented, in part or in whole, as instruction 805, stored in memory 803 and executed by processor 801. Any of the various examples of data described herein may be in data 807, stored in memory 803 and used by processor 801 during the execution of instruction 805.

[0112]

[0119] The computer system 800 may also include one or more communication interfaces 809 for communicating with other electronic devices. One or more communication interfaces 809 may be based on wired communication technology, wireless communication technology, or both. Some examples of one or more communication interfaces 809 include Universal Serial Bus (USB), Ethernet adapters, wireless adapters operating according to the IEEE (Institute of Electrical and Electronics Engineers) 802.11 wireless communication protocol, Bluetooth® wireless communication adapters, and infrared (IR) communication ports.

[0113]

[0120] The computer system 800 may also include one or more input devices 811 and one or more output devices 813. Some examples of one or more input devices 811 include a keyboard, mouse, microphone, remote control device, button, joystick, trackball, touchpad, and light pen. Some examples of one or more output devices 813 include a speaker and a printer. A particular type of output device commonly included in the computer system 800 is a display device 815. The display device 815 used in the implementations disclosed herein may utilize any suitable image projection technology such as a liquid crystal display (LCD), light-emitting diode (LED), gas plasma, or electroluminescence. A display controller 817 may also be provided for converting data 807 stored in memory 803 into text, graphics, and / or moving images to be displayed on the display device 815 (as appropriate).

[0114]

[0121] The various components of the computer system 800 can be connected by one or more buses, which may include a power bus, a control signal bus, a status signal bus, a data bus, and so on. For clarity, the various buses are illustrated in Figure 8 as a bus system 819.

[0115]

[0122] Those skilled in the art will understand that this disclosure can be implemented in network computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablets, pagers, routers, switches, and the like. This disclosure can also be implemented in distributed system environments where local and remote computer systems linked over a network (by hardwired data links, wireless data links, or a combination of hardwired and wireless data links) work together to perform tasks. In a distributed system environment, program modules may reside on both local and remote memory storage devices.

[0116]

[0123] The technologies described herein may be implemented in hardware, software, firmware, or any combination thereof, unless otherwise specifically stated as being implemented in a particular manner. Any feature described as a module or component may be implemented together in an integrated logical device, or separately as individual but interoperable logical devices. When implemented in software, these technologies can be at least partially realized by a non-transient processor-readable storage medium containing instructions that, when executed by at least one processor, perform one or more of the methods described herein. Instructions may be organized into routines, programs, objects, components, data structures, etc., which can perform specific tasks and / or implement specific data types, and can be combined or distributed as desired in various implementations.

[0117]

[0124] The computer-readable medium may be any available medium accessible by a general-purpose or special-purpose computer system. The computer-readable medium that stores computer-executable instructions is a non-transient computer-readable storage medium (device). The computer-readable medium that carries computer-executable instructions is a transmission medium. Therefore, as an example, implementations of this disclosure may include at least two distinctly different types of computer-readable medium: a non-transient computer-readable storage medium (device) and a transmission medium.

[0118]

[0125] In this specification, non-temporary computer-readable storage media (devices) may include RAM, ROM, EEPROM, CD-ROM, solid-state drives (SSDs) (e.g., RAM-based), flash memory, phase-change memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other media that can be used to store desired program code means in the form of computer-executable instructions or data structures and can be accessed by a general-purpose computer or a special-purpose computer.

[0119]

[0126] The steps and / or actions of the methods described herein may be interchanged without departing from the claims. In other words, unless a particular order of steps or actions is required for the proper operation of the described method, the order and / or use of any particular steps and / or actions may be changed without departing from the claims.

[0120]

[0127] The term "determining" encompasses a wide variety of actions, and therefore, "determining" may include calculation, computing, processing, derivation, investigation, lookup (e.g., lookup in a table, data repository, or other data structure), confirmation, etc. It may also include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), etc. Furthermore, "determining" may include resolving, selecting, choosing, establishing, etc.

[0121]

[0128] The terms “comprising,” “including,” and “having” are intended to be comprehensive and mean that additional elements may exist beyond those listed. Furthermore, it should be understood that any reference in this disclosure to “one implementation” or “multiple implementations” is not intended to exclude the existence of further implementations that similarly incorporate the described features. For example, any element or feature described in relation to one implementation described herein can be combined with any element or feature of any other implementation described herein, provided they are compatible.

[0122]

[0129] This disclosure may be embodied in other specific forms without departing from its spirit or characteristics. The described implementations are intended to be illustrative and not restrictive. The scope of this disclosure is indicated not by the foregoing description but by the appended claims. Any changes that fall within the meaning and scope of the claims are included therein.

Claims

1. Identifying a given user profile of a given user identifier generated by a profile generation model based on browsing activity, The method involves providing the given user profile to a sensitivity detection machine learning model trained to classify users into sensitivity categories based on their user profile, The sensitivity detection machine learning model determines that the given user profile is classified as a specific sensitivity classification, Based on the specific sensitivity classification of the given user identifier, the profile generation model is instructed to modify one or more features that separate the given user profile from the specific sensitivity classification. Computer implementation methods, including those mentioned above.

2. Identifying the given user profile is Providing the browsing activity associated with the given user identifier to the profile generation model, Receiving a given user profile which includes a first descriptive label subset for the given user profile from a set of descriptive labels, The computer implementation method according to claim 1, including the method described in claim 1.

3. The computer implementation method according to claim 1 or 2, wherein modifying one or more features of the profile generation model causes the profile generation model to use a second set of features to determine a second subset of descriptive labels for the given user profile from the set of descriptive labels.

4. The computer implementation method according to any one of claims 1 to 3, wherein the sensitivity detection machine learning model is a decision tree model that provides a decision path for classifying a given user profile into the specific sensitivity classification.

5. The decision path indicates a combination of descriptive labels identified for the given user profile that resulted in the specific sensitivity classification, and The computer implementation method according to claim 4, wherein the instruction indicates that the combination of descriptive labels results in the given user profile being classified into the specific sensitivity classification.

6. A computer implementation method according to any one of claims 1 to 5, further comprising supplementing the browsing activity of a given user identifier with artificial browsing activity not associated with the given sensitivity classification, based on the specific sensitivity classification of the given user identifier.

7. A computer implementation method according to any one of claims 1 to 6, further comprising generating a training dataset that simulates the browsing activities of a user associated with a plurality of sensitivity classifications.

8. The computer implementation method according to claim 7, further comprising generating an additional training dataset to simulate additional browsing activities of control users not associated with the plurality of sensitivity classifications.

9. The computer implementation method according to claim 7, wherein generating the training dataset for simulating the browsing activity of the user associated with the plurality of sensitivity classifications includes generating a simulated user over several months who visits an additional number of websites associated with the plurality of sensitivity classifications, in addition to visiting additional websites not associated with the plurality of sensitivity classifications.

10. The computer implementation method of claim 9, further comprising generating the sensitivity detection machine learning model by adjusting the sensitivity detection machine learning model to classify the user profiles of simulated users who have visited websites exceeding a threshold amount associated with the particular sensitivity classification into the one or more sensitivity classifications.

11. It is a system, A given user profile for a given user identifier, generated by a profile generation model that generates a user profile for the user identifier, which includes one or more descriptive labels for each of the user identifiers, A sensitivity detection machine learning model trained to classify users into sensitivity categories based on the user profile, A server device comprising at least one processor, When executed by the at least one processor in the server device, In order to determine that the given user profile is classified as a specific sensitivity classification, the given user profile is provided to the sensitivity detection machine learning model (226, 510), Based on the specific sensitivity classification of the given user identifier, the profile generation model is instructed to separate the given user profile from the specific sensitivity classification. Computer memory containing instructions that cause the system to perform an operation including, A system that includes these features.

12. The system according to claim 11, further comprising instructions, when executed by the at least one processor, causing the system to perform an operation that includes supplementing the browsing activity of a given user identifier with artificial browsing activity not associated with the particular sensitivity classification of the given user identifier, based on the particular sensitivity classification of the given user identifier.

13. The system according to claim 11, wherein the instruction causes the profile generation model to determine different descriptive labels for the previously determined given user profile, thereby causing the profile generation model to decouple the given user profile from the particular sensitivity classification.

14. The system according to claim 11, further comprising an instruction that, when executed by the at least one processor, causes the system to perform an operation which includes providing a client device associated with a given user with a visual indication that it has been determined that the given user profile is associated with the particular sensitivity classification.

15. When executed by the aforementioned at least one processor, Identifying potential interactions between the given user and online resources, The addition of the online resources will determine how the user's given user profile changes, If the aforementioned potential interaction changes the classification status of the sensitivity classification, the user is given additional instructions that performing the aforementioned potential interaction with the online resource will change the classification status of the sensitivity classification. The system according to claim 14, further comprising an instruction causing the system to perform an operation including the above.

16. The system according to claim 11, further comprising instructions, wherein the sensitivity detection machine learning model is a decision tree model that provides a decision path for classifying a given user profile into the specific sensitivity classification.

17. Identifying a given user profile generated by a profile generation model based on browsing activity associated with a given user identifier, The given user profile is provided to a sensitivity detection machine learning model (226, 510) that has been trained to classify users into sensitivity categories based on the user profile, The sensitivity detection machine learning model determines that the given user profile is classified as a specific sensitivity classification, Based on the aforementioned specific sensitivity classification, the profile generation model is instructed to classify the given user profile as the aforementioned specific sensitivity classification. Computer implementation methods, including those mentioned above.

18. The computer implementation method according to claim 17, wherein providing the instructions to the profile generation model causes the profile generation model to modify one or more features that separate the given user profile from the specific sensitivity classification.

19. To generate a training dataset that simulates the user's browsing activity associated with multiple sensitivity classifications, To generate additional training data to simulate additional browsing activities of control users not associated with the aforementioned multiple sensitivity classifications, The computer implementation method according to claim 17, further comprising:

20. The computer implementation method according to claim 19, wherein generating the training dataset that simulates the browsing activity of the user associated with the plurality of sensitivity classifications includes generating a simulated user over several months who browses a random amount of websites associated with one or more sensitivity classifications, in addition to browsing additional websites not associated with one or more sensitivity classifications.