Machine learning models for preventing sensitive data from being exposed online

By integrating machine learning models in content editing tools, real-time detection and evaluation of potential privacy leaks, the problem of unintentional disclosure of sensitive data on online platforms is solved, and automated privacy protection and editing suggestions are achieved.

CN114462616BActive Publication Date: 2025-08-08ADOBE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110864458.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-11-09
Filing Date
2021-07-29
Publication Date
2025-08-08
Estimated Expiration
2041-07-29

AI Technical Summary

Technical Problem

Existing content editing tools have the risk of unintentional disclosure of sensitive data, especially on online platforms, leading to the rapid spread and exposure of privacy issues.

Method used

Using machine learning models combined with content editing tools to detect and identify potential privacy leaks in real time, generating modification suggestions to reduce the disclosure of sensitive data by calculating privacy scores between entities.

Benefits of technology

It effectively reduces the unintentional disclosure of sensitive data on online platforms, provides real-time feedback and automated editing suggestions, reduces users' reliance on subjective judgments, and improves the security of content editing tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114462616B_ABST
    Figure CN114462616B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to machine learning models for preventing the online disclosure of sensitive data. Systems and methods use machine learning models in conjunction with content editing tools to prevent or mitigate the unintentional disclosure and dissemination of sensitive data. By applying a trained machine learning model to a set of unstructured text data received via an input field of an interface, entities associated with private information can be identified. A privacy score is calculated for the text data by identifying connections between entities, the connections between entities contributing to the privacy score based on the cumulative privacy risk, the privacy score indicating the potential exposure of the private information. The interface is updated to include an indicator that distinguishes a target portion of a set of unstructured text data within the input field from other portions of the set of unstructured text data within the input field, wherein modifications to the target portion change the potential exposure of the private information indicated by the privacy score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to using artificial intelligence to prevent the unintentional disclosure of sensitive data. More specifically, but not by way of limitation, the present disclosure relates to techniques for using machine learning models with content editing tools to prevent or mitigate the unintentional disclosure and dissemination of sensitive data in real time. Background Art

[0002] Artificial intelligence technologies for processing text are useful in various content editing tools. For example, when a user enters content into an online search, a machine learning model is used to predict the next word. As another example, machine learning is used in online word processing software to suggest changes to improve the readability of text content.

[0003] However, these types of content editing tools often present the risk that sensitive information, such as personally identifiable information, may be unintentionally disclosed. For example, a user may enter seemingly innocuous information in an online forum, such as stating that the user is a "software engineer from Florida," which can be used in conjunction with other online content to identify the user. In some cases, the online nature of certain content editing tools presents a unique risk of allowing this sensitive data to be rapidly disseminated, sometimes irrevocably, once it is unintentionally disclosed. As the amount of information that individuals post to the Internet increases rapidly, the privacy concerns arising from the exposure of personally identifiable information also increase rapidly. Seemingly innocuous data elements, when aggregated, can provide a complete view of someone that they never intended to publish or were aware of that could be obtained through their interaction with the Internet. Summary of the Invention

[0004] Certain embodiments relate to techniques for using machine learning models to flag potential privacy leaks in real time.

[0005] In some aspects, a computer-implemented method includes: detecting, by a content retrieval subsystem, entry of a set of unstructured text data into an input field of a graphical interface; identifying, in response to detecting the entry and utilizing a natural language processing subsystem, a plurality of entities associated with private information by applying at least a trained machine learning model to the set of unstructured text data in the input field; calculating, by a scoring subsystem, a privacy score for the text data by identifying connections between entities, the connections between entities contributing to the privacy score based on a cumulative privacy risk, the privacy score indicating potential exposure of the private information by the set of unstructured text data; and updating, by a reporting subsystem, the graphical interface to include an indicator that distinguishes a target portion of the set of unstructured text data within the input field from other portions of the set of unstructured text data within the input field, wherein modifications to the target portion change the potential exposure of the private information indicated by the privacy score.

[0006] In some aspects, the method also includes: detecting, by the content retrieval subsystem, modifications to a set of unstructured text data entered into an input field of the graphical interface; identifying, in response to detecting the modifications and utilizing the natural language processing subsystem, modified multiple entities associated with the private information by applying at least a trained machine learning model to the modified text data in the input field; calculating, by the scoring subsystem, a modified privacy score for the text data based on the modified entities; and updating, by the reporting subsystem, the graphical interface based on the modified privacy score.

[0007] In some aspects, the method further includes: receiving, by the content retrieval subsystem, an image or video associated with the unstructured text data; and processing, by the media processing subsystem, the image or video to identify metadata, wherein at least a subset of the identified metadata is further input to the machine learning model to identify entities.

[0008] In some aspects, the set of unstructured text data is a first set of unstructured text data and the plurality of entities is a first plurality of entities, and the method further includes: prior to receiving the first set of unstructured text data: detecting, by the content retrieval subsystem, entry of a second set of unstructured text data entered into an input field; and in response to detecting the entry and utilizing the natural language processing subsystem, identifying a second plurality of entities associated with the private information by applying at least the trained machine learning model to the second set of unstructured text data in the input field, wherein the scoring subsystem calculates a privacy score based on connections between the first plurality of entities and the second plurality of entities.

[0009] In some aspects, the updated graphical interface further displays an indication of the privacy score. In some aspects, the machine learning model includes a neural network, and the method further includes training the neural network by: retrieving, by the training subsystem, first training data for a first entity type associated with the privacy risk from a first database; retrieving, by the training subsystem, second training data for a second entity type associated with the privacy risk from a second database; and training, by the training subsystem, the neural network using the first training data and the second training data to identify the first entity type and the second entity type.

[0010] In some aspects, the method further includes: determining, by the natural language processing subsystem, an entity type for the identified entity; and assigning, by the scoring subsystem, weights to links between entities in the graph model based on the determined entity type, wherein the privacy score is based on the weights.

[0011] In some aspects, a computing system includes: a content retrieval subsystem configured to detect entry of unstructured text data into an input field of a graphical interface; a natural language processing subsystem configured to identify multiple entities associated with private information by applying at least a trained machine learning model to the unstructured text data; a scoring subsystem configured to calculate a privacy score for the text data by applying a graph model to multiple entities to identify connections between the entities, the connections between the entities contributing to the privacy score based on a cumulative privacy risk, the privacy score indicating potential exposure of the private information by the unstructured text data; and a reporting subsystem configured to update the graphical interface to include an indicator that distinguishes a target portion of the unstructured text data within the input field from other portions of the unstructured text data within the input field, the target portion causing the potential exposure of the private information indicated by the privacy score.

[0012] In some aspects, a non-transitory computer-readable medium having instructions stored thereon, the instructions being executable by a processing device to perform operations, the operations comprising: detecting entry of a set of unstructured text data into an input field of a graphical interface; steps for calculating a privacy score for the text data, the privacy score indicating potential exposure of private information by the set of unstructured text data; and updating an indicator based on the privacy score, the indicator distinguishing a target portion of the set of unstructured text data within the input field from other portions of the set of unstructured text data within the input field.

[0013] These illustrative embodiments are mentioned not to limit or define the present disclosure, but to provide examples to aid its understanding.Additional embodiments are discussed in the detailed description, and further description is provided there. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The features, embodiments, and advantages of the present disclosure may be better understood when the following detailed description is read with reference to the accompanying drawings.

[0015] Figure 1 An example of a computing environment in which a content editing tool uses a machine learning model to indicate content modifications for addressing potential privacy leaks in real time is depicted, according to certain embodiments of the present disclosure.

[0016] Figure 2 Depicted is one example of a process for updating an interface of a content editing tool in real time to indicate potential edits that would reduce exposure of private information, in accordance with certain embodiments of the present disclosure.

[0017] Figures 3A-3D Illustrated is the use of certain embodiments according to the present disclosure Figure 2 An example of a sequence of graphical interfaces generated by the process depicted in FIG.

[0018] Figure 4 Described is a method for training a Figure 2 An example of a process using a machine learning model.

[0019] Figure 5 One example of a computing system that performs certain operations described herein is depicted in accordance with certain embodiments of the present disclosure.

[0020] Figure 6 One example of a cloud computing environment is depicted for performing certain operations described herein, according to certain embodiments of the present disclosure. DETAILED DESCRIPTION

[0021] The present disclosure includes systems and methods for using machine learning models with content editing tools to prevent or mitigate the unintentional disclosure and dissemination of sensitive data in real time. As described above, online services and other content editing tools present the risk of inadvertently disclosing sensitive data, which can spread rapidly via the Internet or other data networks. Certain embodiments described herein address this risk by using machine learning models to detect potentially problematic content during the editing phase and indicate potential modifications to the content that would reduce the disclosure of sensitive data. For example, such embodiments analyze unstructured text data to identify words or phrases associated with private information. A privacy score is generated based on the connections between these words or phrases, and based on the privacy score, information is displayed that can encourage users to modify the text data to reduce the exposure of private information.

[0022] The following non-limiting example is provided to introduce certain embodiments. In this example, a privacy monitoring system communicates with a web server that provides data for presenting a graphical interface (e.g., a graphical user interface (GUI)) on a user device. The graphical interface includes a text field configured to receive text data. The privacy monitoring system retrieves the text data as the user enters the text data, identifies elements of the text data, and identifies relationships between the various elements of the text data that pose privacy risks. For example, the privacy monitoring system detects the entry of a set of unstructured text data into an input field of the graphical interface. The graphical interface is used to edit and publicly publish information, such as product reviews, social media posts, classified ads, etc. A content retrieval subsystem monitors the entry of information into the input field and, upon detecting the entry of information, initiates processing of the text to identify privacy issues. Privacy issues may arise from information that exposes sensitive data, such as personally identifiable information (PII), which can be used alone or in combination with other publicly accessible data to identify an individual. Examples of such sensitive data include a person's address, city, bus stop, medical issues, etc.

[0023] Continuing with this example, the privacy monitoring system processes text data to identify entities associated with private information. To do this, the privacy monitoring system applies a machine learning model to the text data. The machine learning model is a named entity recognizer that is trained to identify specific categories of entities associated with potential privacy issues (such as location information, medical information, etc.). The privacy monitoring system generates a graph model of entities, identifying the connections between entities and how the entities are related to each other, which is used to generate a privacy score that indicates the potential exposure of private information by a set of unstructured text data. The connections between entities contribute to the privacy score based on the cumulative privacy risk. This can be achieved via the graph model by weighting the links between different entities based on the cumulative risk of the different entities - for example, leaking two pieces of location information such as a city and a street will result in more privacy leakage than leaking one piece of health information and one piece of location information (for example, the person has asthma and lives in Dallas). Therefore, the weights for links between similar entity types in the graph model can be heavier than the weights for links between different entity types.

[0024] In this example, the privacy monitoring system calculates a privacy score based on the entities in the graph and the weighted links between the entities. The privacy score can be used to recommend or otherwise indicate edits that will reduce the risk of sensitive data being disclosed. The privacy monitoring system compares the privacy score to one or more thresholds to identify whether the text should be modified and identifies the recommended modification (e.g., removing the name of a street from a comment). Such information generated by the privacy monitoring system is output to the privacy monitoring system for reporting via a graphical interface. To facilitate editing of text, the privacy monitoring system updates the graphical interface to include an indicator that distinguishes a target portion of a set of unstructured text data (e.g., one or more entities) within an input field from other portions of the set of unstructured text data within the input field. When a modification to the target portion is detected, the privacy monitoring system can repeat the analysis to identify an updated privacy score and modify or remove the suggestion. Thus, the system can identify privacy issues in real time by retrieving and processing text as the user enters it, to instantly generate and provide suggestions that can be used to help produce text content (e.g., online posts) with reduced exposure of private information or other sensitive data.

[0025] As described herein, certain embodiments provide improvements to computing environments by addressing issues specific to online content editing tools. These improvements include providing real-time feedback in editing tools that alerts users to the potential disclosure of sensitive data before it is published on the Internet. Online computing environments pose unique risks to this type of sensitive data exposure because the Internet or other data networks allow for near-instant transmission and publication to a large number of recipients, while the utility provided by online content editing tools (e.g., publication via the click of a single button) increases the risk that such publication and transmission may occur unexpectedly. Furthermore, the variety of information available via the Internet limits a user's ability to accurately determine whether any given data fragment posted on an online forum can be combined with other publicly available data to identify the user. Because these issues are specific to computing environments, the embodiments described herein utilize machine learning models and other automated models that are uniquely suited to mitigating the risk of unintentional dissemination of user data via the Internet or other data networks. For example, computing systems sometimes automatically apply various rules of a specific type (e.g., various functions captured in one or more models) to text entered into a user interface in real time. These rules can more effectively detect the potential disclosure of sensitive data, not least because the system is trained using information from a large corpus to identify and quantify different levels of sensitive private information in text (both individually and in relation to previous posts), rather than relying on the subjective judgment of the user posting the content.

[0026] Additionally or alternatively, certain embodiments provide improvements to existing software tools for safely creating online content. For example, existing software tools require users of editing tools executed on computers to subjectively determine the level of risk associated with entering certain data into the online editing tools. Relying on these subjective determinations may reduce the effectiveness of the editing tools used to create online content. The embodiments described herein can support an automated process for creating online content that avoids such reliance on subjective, manual determinations made by users. For example, the combination of machine learning models with structural features of user interfaces (e.g., suggestions for reducing exposure risks or other indicators of potential editing) improves the functionality of online editing tools. These features can reduce the manual, subjective work involved in preventing the disclosure of sensitive data in existing content editing tools.

[0027] As used herein, the term "private information" refers to information that can be used to identify an individual or sensitive information about that individual. For example, private information can include information that directly identifies an individual, such as name, address, or social security information, as well as information that indirectly identifies an individual, such as race, age, and region of residence. Certain categories of information about an individual are also private, such as medical conditions and employment information.

[0028] As used herein, the term "entity" is used to refer to a word or phrase that corresponds to a defined category or type of information. An entity can be a proper noun (e.g., "124 Main Street"). An entity can also be a phrase that represents a selected category of information (e.g., "backache," "pineapple," "seven grandchildren"). Entities may fall into categories or types such as places, things, people, medical conditions, etc. Some entities are associated with private information, such as location information, medical information, and employment information.

[0029] As used herein, the term "privacy risk" refers to the level of potential exposure of private information. The more private information there is, and the more sensitive the private information is, the higher the privacy risk. Privacy risk can be determined for a single exposure (e.g., a single online post) or cumulatively (e.g., across multiple online posts).

[0030] An example of an operational environment for real-time privacy leak prediction

[0031] Figure 1 An example of a computing environment 100 is depicted in which a content editing tool uses a machine learning model to indicate content modifications in real time to address potential privacy breaches. Figure 1 In the depicted example, user device 102 publishes information via web server 109. Privacy monitoring system 110 evaluates the information to identify privacy issues using content retrieval subsystem 112, natural language processing (NLP) subsystem 114, media processing subsystem 116, and reporting subsystem 120. The subsystem includes one or more trained machine learning models trained using training data 126A-126N using training subsystem 122.

[0032] The various subsystems of the privacy monitoring system 110 can be implemented on the same computing system or on different, independently operating computing systems. For example, the training subsystem 122 can be a separate entity from the NLP subsystem 114, the media processing subsystem 116, and the scoring subsystem 118, or the same entity. A different, independently operating web server 109 can communicate with the privacy monitoring system 110, or the privacy monitoring system 110 can be part of the same online service as the web service. Although it is possible to use Figure 1 system, but other embodiments may involve building the privacy monitoring system 110 into a software application executing on the client device 102, for example, as a plug-in to some word processing software.

[0033] Some embodiments of computing environment 100 include user device 102. Examples of user devices include, but are not limited to, personal computers, tablet computers, desktop computers, processing units, any combination of these devices, or any other suitable device with one or more processors. A user of user device 102 interacts with graphical interface 104 by exchanging data with web server 109 and privacy monitoring system 110 via a data network.

[0034] The user device is communicatively coupled to the web server 109 and the privacy monitoring system 110 via a data network. Examples of data networks include, but are not limited to, the Internet, a local area network ("LAN"), a wireless LAN, a wired LAN, a wide area network, and the like.

[0035] Graphical interface 104 is an interface, such as a GUI, capable of displaying and receiving information. Graphical interface 104 includes content editing tools for receiving and modifying content (e.g., content to be published online). Graphical interface 104 includes a text field 105 for receiving text data 106. For example, text field 105 is an interface element configured to receive typed text data 106 from a user of user device 102. Alternatively or additionally, in some embodiments, text field 105 is configured to receive text data identified by the system by processing voice user input (e.g., using speech-to-text processing technology).

[0036] In some embodiments, the graphical interface 104 also includes an upload element 107, through which the user can upload additional information, such as images or videos. In response to the user selecting the upload element, the graphical interface 104 transitions to a view showing available files to upload, prompting the user to take a photo, etc.

[0037] Graphical interface 104 is further configured to display a privacy alert 108 in response to a signal from privacy monitoring system 110 (directly or via web server 109). For example, privacy alert 108 includes information characterizing the risk associated with a portion of text data 106 (e.g., a privacy risk score, a different color marker, a warning, etc.). In some embodiments, privacy alert 108 indicates the portion of text data 106 associated with the potential exposure of private information (e.g., highlighted, printed in a different color, a bubble with explanatory text, etc.). An example of graphical interface 104 including text field 105, upload element 107, and privacy alert 108 is shown in FIG. Figures 3A-3D is shown in the figure.

[0038] In some embodiments, web server 109 is associated with an entity such as a social network, an online merchant, or a variety of websites that allow users to post information. Web server 109 includes functionality for serving the website (which may include content editing tools) and accepting input from user device 102 and / or privacy monitoring system 110 for modifying the website. In some embodiments, web server 109 is a separate entity or computing device from privacy monitoring system 110. Alternatively, in some embodiments, web server 109 is a component of privacy monitoring system 110.

[0039] The privacy monitoring system 110 monitors updated information received from the user device 102 via the graphical interface 104 and analyzes the information for privacy risks. In some embodiments, an indication of the privacy risk is then presented by updating the graphical interface 104. The privacy monitoring system 110 includes a content retrieval subsystem 112, a natural language processing (NLP) subsystem 114, a media processing subsystem 116, a scoring subsystem 118, a reporting subsystem 120, and a training subsystem 122. In some embodiments, the privacy monitoring system also includes or is communicatively coupled to one or more data storage units (124A, 124B, ..., 124N) for storing training data (training data A 126A, training data B 126B, ..., training data N 126N).

[0040] The content retrieval subsystem 112 includes hardware and / or software configured to retrieve content that a user is entering into the graphical interface 104. The content retrieval subsystem 112 is configured to retrieve unstructured text data 106 as it is entered into the text field 105 of the graphical interface 104. In some implementations, the content retrieval subsystem 112 is also configured to retrieve media such as images and videos uploaded via the upload element 107.

[0041] The NLP subsystem 114 includes hardware and / or software configured to perform natural language processing to identify entities (e.g., certain words or phrases) associated with privacy risks. In some embodiments, the NLP subsystem 114 applies a machine learning model that is trained to identify entities associated with privacy risks, such as health-related words or phrases, street names, city names, etc. Examples of phrases that may be associated with privacy risks include:

[0042] - "For our upstairs bathroom" - means homes with more than 1 floor

[0043] - "Texas Summer" - helps triangulate the user's location

[0044] - "Get a privacy-focused screen reader at a nearby coffee shop" - helps triangulate the user's location

[0045] - 'Bought for my son's asthma' - Leaked health status

[0046] The media processing subsystem 116 includes hardware and / or software configured to analyze media files to identify entities. The media processing subsystem 116 is configured to process images or videos to identify metadata and / or text within the images themselves. In some aspects, entities are identified by analyzing media files to identify metadata (e.g., including location information). Alternatively or additionally, the media processing subsystem 116 identifies entities by analyzing images (e.g., identifying words on a sign in a photograph).

[0047] Scoring subsystem 118 includes hardware and / or software configured to generate a privacy score based on entities identified by NLP subsystem 114 and / or identified by media processing subsystem 116. For example, scoring subsystem 118 generates a graph of the identified entities. By taking into account the weights assigned to the links between the entities, scoring subsystem 118 generates a privacy score that represents the overall information exposure of the entity as a whole. In some aspects, the scoring subsystem further identifies recommended actions, specific words that should be removed or modified, etc., as described herein.

[0048] The reporting subsystem 120 includes hardware and / or software configured to generate and transmit an alert to a user, which may include a privacy score and other information generated by the scoring subsystem 118. The reporting subsystem 120 causes the privacy alert 108 to be displayed to the graphical interface 104. The privacy alert 108 includes a graphical display, such as text, a highlighted portion of text, etc. Alternatively or additionally, in some embodiments, the privacy alert 108 includes an audio alert, such as a beep or voice output.

[0049] The training subsystem 122 includes hardware and / or software configured to train one or more machine learning models, such as used by the NLP subsystem 114, the media processing subsystem 116, and / or the scoring subsystem 118. An example training process is described below with respect to Figure 4 Be described.

[0050] Data storage units 124A, 124B...124N may be implemented as one or more databases or one or more data servers. Data storage units 124A, 124B...124N include training data 126A, 126B...126N, which are used by training subsystem 122 and other engines of privacy monitoring system 110, as described in further detail herein.

[0051] Example of operations for real-time privacy leak prediction

[0052] Figure 2 An example of a process 200 for updating an interface of a content editing tool in real time to indicate potential edits that would reduce exposure of private information is depicted. In this example, a privacy monitoring system 110 detects input to a graphical interface 104 via a content retrieval subsystem 112. The input is processed in a pipeline that includes an NLP subsystem 114, a scoring subsystem 118, and in some cases, a media processing subsystem 116. If a portion of the input poses a risk of exposure of private information above an acceptable threshold, a reporting subsystem 120 modifies the graphical interface 104 to include a privacy alert 108, which can enable the user to modify the information entered. Alternatively or additionally, in some other embodiments, the privacy monitoring system can be executed as part of a software application executed on a client device, where the software application can execute one or more of blocks 202-206, 212, and 214. In some embodiments, one or more processing devices implement the process by executing appropriate program code. Figure 2 For illustrative purposes, process 200 is described with reference to certain examples depicted in the figures. However, other implementations are possible.

[0053] At block 202, the content retrieval subsystem receives a set of unstructured text data entered into an input field of a graphical interface. When a user enters text data into the graphical interface, the content retrieval subsystem detects and identifies the entered text data. As the user types text via the graphical interface, the content retrieval subsystem retrieves the unstructured text data, for example, as a stream or in blocks. The content retrieval subsystem may retrieve the set of unstructured text data directly from the user device or via an intermediate web server.

[0054] The processing device executes program code of the content retrieval subsystem 112 to implement block 202. For example, program code for the content retrieval subsystem 112 stored in a non-transitory computer-readable medium is executed by one or more processing devices.

[0055] One or more operations in blocks 204-210 implement steps for calculating a privacy score for the text data, the privacy score indicating the potential exposure of private information by a set of unstructured text data. In some embodiments, at block 204, the content retrieval subsystem receives an image or video associated with the unstructured text data. For example, the content retrieval subsystem identifies the image or video in response to detecting that a user interacts with an "upload" button and selects a media file to be stored on the user's device. Alternatively or additionally, the user captures the image or video when submitting it via a graphical interface.

[0056] At block 206, the media processing subsystem processes the image or video file to identify metadata. In some embodiments, the media processing subsystem extracts metadata from the received media file (e.g., JPEG, MP4, etc.). Alternatively or additionally, the media processing subsystem analyzes the image or video data itself to identify words. For example, an image may include the name of a street, building, or bus stop. The media processing subsystem performs optical character recognition on the picture or video still to identify any words therein. Both the metadata and the identified words can be treated by the privacy monitoring system as additional textual data for privacy analysis.

[0057] At block 208, the NLP subsystem processes the text data using the trained machine learning model to identify a plurality of entities associated with the private information. Examples of the types of entities associated with privacy risks include names, streets, and local landmarks such as schools, museums, and bus stops. Other examples of entities associated with privacy risks include information about health conditions, information about family status, and information about employment status. In some embodiments, at least a subset of the metadata identified at block 206 is further input into the machine learning model to identify entities.

[0058] In some embodiments, the NLP subsystem processes the data in response to detecting the entry of text data at block 202. In some implementations, at block 206, the NLP subsystem further processes the information identified from the media file. The NLP subsystem identifies multiple entities associated with the private information by applying at least a trained machine learning model to a set of unstructured text data in an input field. Alternatively or additionally, at block 206, the NLP subsystem applies the trained machine learning model to the identified image metadata and / or words identified from the image.

[0059] In some aspects, the trained machine learning model is a named entity recognizer that has been trained to identify certain words or categories of words that are associated with privacy risks. The named entity recognizer processes text data to identify entities within the text data and then labels the text data with information about the identified entities. The machine learning model uses information such as the following about Figure 4 In some embodiments, the machine learning model is a neural network, such as a recurrent neural network (RNN), a convolutional neural network (CNN), or a deep neural network. In some embodiments, the machine learning model is an ensemble model (e.g., including a neural network and another type of model, such as a rule-based model).

[0060] At block 210, the scoring subsystem calculates a privacy score for the text data by identifying connections between entities. In some embodiments, the scoring subsystem generates a graph model of entities (also referred to as a graph), the graph model including connections between entities. The nodes of the graph are entities, which may include entities identified from the text data at block 202 and entities identified from image metadata or the image itself at block 206. Connections between entities contribute to the privacy score based on the cumulative privacy risk. For example, connections between different entities may be weighted differently to account for the increased risk of exposing certain entities together. As a specific example, street names and city names together pose a relatively large cumulative privacy risk because they can be used together to identify a location, while the combination of medications and street names poses a smaller cumulative privacy risk because the entities are less related. The scoring subsystem may then generate a privacy score based on the number of links and the weights of these links. Thus, in some embodiments, the scoring subsystem determines entity types (e.g., medical condition, street, age, etc.). Using the determined entity types, the scoring subsystem assigns weights to the links between entities in the graph model, where the privacy score varies based on the weights. The privacy score indicates the potential exposure of private information by a set of unstructured text data.

[0061] In some aspects, the scoring subsystem determines a sensitivity level for each identified entity. In some aspects, entities are weighted or labeled with different sensitivity categories. For example, depending on the entity type, certain entities are assigned a higher weight than other entities. As a specific example, more specific entities are weighted more heavily than more general entities (e.g., the name of the street where the user lives is weighted more heavily than the name of the continent where the user lives). In some embodiments, a machine learning model is trained to recognize these sensitivity levels (e.g., using the assigned labels). For example, entities related to medical, health, and financial information are labeled with the highest sensitivity level. Then, another group of entities (example: entities related to demographics and geolocation) can be labeled with a medium sensitivity level.

[0062] In some aspects, the scoring subsystem generates a personalized graph for the user based on one or more text entries. In some embodiments, the scoring subsystem generates a graph that includes information derived from multiple text entries (e.g., multiple comments, multiple social media posts, etc.). As an example, the text received at box 202 is a product review detected by the system in real time. The privacy monitoring system is coupled to other sites (such as social media) to identify other posts posted by the user in other contexts. This information can be used together to generate a graph. Alternatively or additionally, the scoring subsystem uses the current text entry to generate a graph. The graph includes nodes in the form of identified entities and connections between nodes that are weighted according to the relationships between the entities. In some embodiments, weights are assigned according to rules. Alternatively, machine learning is used to calculate appropriate weights. Based on the connections and their weights, the scoring subsystem generates a score that indicates the overall exposure of sensitive information.

[0063] For example, when a user enters a review, the scoring subsystem creates a personalized graph of extracted entities ranked by sensitivity level, which generates a score for the user's review. When the user returns to the system and begins submitting another review, their sensitive entity graph is enhanced (so that entities from the previous review are linked to the new review). In this way, a review is scored based on the information it reveals both alone and in combination with information revealed by previous reviews.

[0064] Thus, in some aspects, prior to receiving the first set of unstructured text data (e.g., in a previous post by the user), the content retrieval subsystem detects the entry of a second set of unstructured text data into an input field. In response to detecting the entry and utilizing the natural language processing subsystem, the content retrieval subsystem identifies a second plurality of entities associated with the private information by applying at least the trained machine learning model to the second set of unstructured text data in the input field. The second plurality of entities may represent the same or different entities that the user entered in the previous post. For example, the user enters text including the entities "Main Street," "Georgia," and "Neurosurgeon" in a product review on September 6. Subsequently, on October 25, the user enters another review including the entities "Georgia," "fifth floor," and "restaurant next to my apartment building." The scoring subsystem updates the graph for the user and calculates a privacy score based on the connections between the first plurality of entities and the second plurality of entities.

[0065] In some aspects, the weights assigned to links between entities degrade over time. For example, links between entities in the same post are weighted more heavily, and the weights degrade over time. As a specific example, an entity in the current post may have a link weight of 0.7 with another entity in the current post, a link weight of 0.5 with another entity in a post from the previous day, and a link weight of 0.1 with a post from two months ago.

[0066] In some aspects, the scoring subsystem generates a privacy score based on the weighted links between entities and the sensitivity levels of the entities themselves. For example, the scoring subsystem uses the generated graph to identify nodes and links between nodes and uses the corresponding weights to calculate the privacy score. As a specific example, the privacy score can be calculated using the following function:

[0067]

[0068] Where P is the privacy score, W ei is the weight of the i-th entity, W lj is the jth link weight. In some embodiments, the scoring subsystem continuously updates the score as additional text is detected. For example, as the user continues to type additional text, the privacy score is updated to reflect the additional detected entities.

[0069] In some aspects, the privacy score is further used by the scoring subsystem to identify a privacy risk level (e.g., a security rating). For example, the scoring subsystem compares the calculated privacy score to one or more thresholds. If the privacy score is below the threshold, the privacy risk level is "low"; if the privacy score is below a second threshold, the privacy risk level is "medium"; and if the privacy score is equal to or greater than the second threshold, the privacy risk level is "high."

[0070] The processing device executes the program code of the scoring subsystem 118 to implement block 210. In one example, the program code for the scoring subsystem 118 (which is stored in a non-transitory computer-readable medium) is executed by one or more processing devices. Executing the scoring subsystem 118 causes the processing device to calculate a privacy score.

[0071] At box 212, the reporting subsystem updates the graphical interface to include an indicator that distinguishes a target portion of a set of unstructured text data within the input field from other portions of a set of unstructured text data within the input field. For example, the reporting subsystem updates the graphical interface by transmitting instructions to the user device (and / or the intermediate web server), causing the user device to display the updated graphical interface. For example, the reporting subsystem transmits instructions that cause the graphical interface to be modified to highlight an entity, display the entity in bold or other fonts, place a box around the entity, etc. Alternatively or additionally, the reporting subsystem causes display of an indication of the privacy risk level (e.g., security level), such as a color code and / or text. Alternatively or additionally, the reporting subsystem transmits a signal that causes the graphical interface to display text that explains the potential privacy risk posed by the marked text data. An example of a graphical interface view displaying an indicator that distinguishes the target portion of text and the privacy risk level is shown in FIG. Figures 3A-3D In some implementations, the reporting subsystem causes display of a word cloud describing all content collectively disclosed by the user across posts that can be used to identify the user.

[0072] In some embodiments, as Figure 3A-3C As shown in , as the user enters additional text data, the additional words are highlighted and the privacy level is modified to a higher risk level. Thus, as the user modifies the text, the privacy monitoring system dynamically repeats steps 202-212 to generate an updated privacy score and displays an updated or additional indicator that distinguishes the target portion of the text.

[0073] At block 214, the modification to the target portion changes the potential exposure of the private information as indicated by the privacy score. For example, a user interacts with the graphical interface to modify the target portion. The content retrieval subsystem detects the modification to a set of unstructured text data entered into an input field of the graphical interface. In response to detecting the modification, the natural language processing subsystem identifies modified entities associated with the private information by applying at least a trained machine learning model to the modified set of unstructured text data in the input field. The scoring subsystem calculates a modified privacy score for the text data based on the modified entities.

[0074] For example, at block 212, in response to the indication(s) displayed by the privacy monitoring system via the graphical interface, the user deletes or modifies a portion of the textual data. As a specific example, the user deletes a phrase that has been highlighted as a potential privacy risk. As a result, the scoring subsystem recalculates the privacy score, now with fewer entities and links, resulting in the privacy score indicating a lower risk level (e.g., a lower privacy score). An example of such a situation is Figure 3C and Figure 3D is shown in the figure.

[0075] In some embodiments, the privacy monitoring system provides a content editing tool that includes an element for users to provide feedback to control the sensitivity of the privacy score. Figures 3A-3D As shown in , the graphical interface includes a slider (e.g., 312) that a user can use to control the privacy sensitivity of the model. If the privacy sensitivity is higher, the system is more likely to generate a privacy alert. For example, if the privacy sensitivity level increases, the model used to generate the privacy score is modified to identify more entities and / or weight entities and links between entities more heavily. For lower privacy sensitivity levels, certain entities are not identified as risky and / or are weighted less heavily. In some aspects, the privacy monitoring system re-executes the operations at blocks 202-210 in response to detecting a change to such a privacy sensitivity modification element, which can result in a modified privacy score.

[0076] Based on the updated privacy score, the reporting subsystem updates the graphical interface. For example, the reporting subsystem updates the graphical interface to include fewer indicators that distinguish the target portion of the text data. Alternatively or additionally, the reporting subsystem updates the graphical interface to indicate the new privacy score or privacy risk level.

[0077] Example GUI with privacy alerts

[0078] Figures 3A-3D Depicted are examples of graphical interface views 300-370 according to certain embodiments of the present disclosure. In some aspects, the graphical interface 104 includes an online content editing tool having an edit mode in which a user can create a post (e.g., a product review, a comment, etc.). The online tool also includes a "publish" mode in which the comment is available to other users (and the original user may not edit it). When text is entered via the graphical interface 104, the above description of the Figure 2 Analysis of the text of the description is triggered.The resulting privacy score is used to display an indication of the privacy risk via the graphical interface 104 as shown in the graphical interface views 300-370.

[0079] Figure 3A An example of a graphical interface view 300 is shown. The graphical interface view 300 includes a text entry field 302 in which a user has entered text 304. The graphical interface view 300 also includes a photo upload element 308 (labeled "Add Photo") and a video upload element 306 (labeled "Add Video"). As the user enters text 304 into the text entry field 302, the privacy monitoring system generates a privacy score in real time, as described above with respect to the privacy score. Figure 2 As described. Figure 3AIn the example shown, a privacy score is used by a privacy monitoring system to identify a privacy risk level. In this case, there is a phrase highlighted as a potential privacy risk 310 - "My back hurts." The privacy monitoring system causes the text to be highlighted to show that the user may wish to remove or modify user content. Since there is only one risky phrase in the text 304, the privacy risk level 314 is relatively low. This is indicated by displaying a "smart meter" in green, with the text "Mostly safe review content." In some embodiments, the graphical interface view 300 also includes a slider 312 for accepting user feedback to control the sensitivity of the privacy score. Via the slider 312, the user can modify the privacy sensitivity level used by the privacy monitoring system to generate the privacy score and determine whether to display an alert. The slider 312 can start with some default privacy sensitivity level (e.g., medium), which can be adjusted via user input.

[0080] Figure 3B An example of an updated graphical interface view 330 is shown. The graphical interface view 330 includes a text entry field 332 in which the user has entered text 334. The graphical interface view 330 also includes a photo upload element 338 (labeled "Add Photo") and a video upload element 336 (labeled "Add Video"). As the user enters text 334 into the text entry field 332, the privacy monitoring system updates the privacy score. As the user continues to enter text, the system updates the privacy score in real time, as described above with respect to Figure 2 As described. Figure 3B In the example shown, text 334 includes four phrases that have been highlighted as potential privacy risks 340 - "my back hurts," "wife and grandchildren," "Florida," and "software engineer." As more phrases of potential privacy risks are added, the privacy risk level 344 has increased to a medium level. This is indicated by displaying a "smart meter" in orange, with the text "Something that might not be appropriate to disclose." The graphical interface view 330 also includes a slider 342 for accepting user feedback to control the sensitivity of the privacy score. In this case, the privacy sensitivity selected is high, which will result in more words being highlighted and a higher privacy risk level 344 than if the privacy sensitivity were medium or low, in which case certain phrases could be used without triggering a privacy warning.

[0081] Figure 3CAnother example of an updated graphical interface view 350 is shown. The graphical interface view 350 includes a text entry field 352 in which the user has entered text 354. The graphical interface view 350 also includes a photo upload element 358 (labeled "Add Photo") and a video upload element 356 (labeled "Add Video"). As the user enters text 354 into the text entry field 352, the privacy monitoring system updates the privacy score in real time, as described above with respect to the Figure 2 As described. Figure 3C In the example shown, five phrases are highlighted as potential privacy risks 360: "My back hurts," "Wife and grandchildren," "Florida," "Software engineer," and "Coffee shop down the street." With the addition of another phrase with potential privacy risk, the privacy risk level 364 has increased to a relatively high level. This is indicated by the "smart meter" being displayed in red with the text "Several pieces of content that should not be disclosed." Graphical interface view 350 also includes a slider 362 for accepting user feedback to control the sensitivity of the privacy score. Via slider 362, the user can modify the privacy sensitivity level used by the privacy monitoring system to generate the privacy score and determine whether to display an alert.

[0082] Figure 3D 370. Graphical interface view 370 includes a text entry field 372 in which a user has entered text 374. Graphical interface view 370 also includes a photo upload element 378 (labeled "Add Photo") and a video upload element 376 (labeled "Add Video").

[0083] exist Figure 3D In the example shown, in response to Figure 3C , the user has removed the text (including "software engineer"). As a result, the privacy monitoring system has recalculated the privacy score based on the updated text 374, resulting in a reduced privacy risk level 384, which is displayed in the graphical interface view 370. Figure 3D In the example shown, four phrases were highlighted as potential privacy risks 380 - "my back hurts," "wife and grandchildren," "Florida," and "cafe down the street." With the potential privacy risk phrases removed, the privacy risk level 384 has been reduced back to a medium level. This is indicated by the "smart meter" being displayed in orange with the text "Something that might not be appropriate to disclose." The graphical interface view 370 also includes a slider 382 for accepting user feedback to control the sensitivity of the privacy score. Via slider 382, the user can modify the privacy sensitivity level used by the privacy monitoring system to generate the privacy score and determine whether to display an alert.

[0084] Examples of operations for training machine learning models

[0085] Figure 4 Described according to some embodiments for training as in Figure 2 In this example, the training subsystem 122 of the privacy monitoring system 110 retrieves training data from multiple databases (e.g., data storage unit 124A, data storage unit 124B, etc.). The training subsystem 122 trains the machine learning model to identify different types of entities associated with privacy risks, and the machine learning model can be used in Figure 2 In some embodiments, one or more processing devices implement the Figure 4 For illustrative purposes, process 400 is described with reference to certain examples depicted in the figures. However, other implementations are possible.

[0086] At block 402, the training subsystem retrieves first training data for a first entity type associated with a privacy risk from a first database. For example, the data storage unit 124A may store a list of email addresses. Other examples of entity types that may be retrieved from a particular database include health conditions (e.g., from a health advice website), a person's name, a country's name, a street name, an address, and the like.

[0087] At block 404, the training subsystem receives second training data for a second entity type associated with a privacy risk from a second database. The training subsystem can receive the second training data in a substantially similar manner as the first training data was received at block 402. However, in some cases, the second training data is associated with a different entity type and is from a different database (e.g., the first training data is a list of medical conditions from a medical website and the second training data is a list of email addresses from an online directory).

[0088] At block 406, the training subsystem associates the first training data and the second training data with labels for the first entity type and the second entity type. In some embodiments, the training subsystem labels the first training data according to a named entity type for the overall dataset (e.g., "email address," "employer," "nearby landmark," etc.). In some cases, the training subsystem labels the second training data according to another named entity type for the respective dataset.

[0089] In some aspects, the training subsystem identifies a dataset that has been grouped by a particular entity type (such as name, email address, street, medical condition, etc.). In some embodiments, the training subsystem automatically associates each element in the dataset with a label that identifies the data element as being of the corresponding type. In this way, the labels are already associated with the entity types in the dataset, and there is no need to individually analyze and annotate each entity (a time-consuming process that is typically used to generate training data).

[0090] In some aspects, a curated set of entities are labeled with different sensitivity levels. Entities related to medical, health, and financial information are labeled with the highest sensitivity level. Then, another set of entities (for example, entities related to demographics and geolocation) can be labeled with a medium sensitivity level. This entity labeling can be done at a coarse high, medium, low, or more refined level.

[0091] At box 408, the training subsystem uses the first training data and the second training data to train a machine learning model (e.g., a neural network) to identify the first entity type and the second entity type. In some embodiments, the machine learning model is trained using backpropagation. For example, the machine learning model receives training data as input and outputs a predicted result. The result is compared with the label assigned to the training data. In some embodiments, the comparison is performed by determining a gradient based on the input and the predicted result (e.g., minimizing a loss function by calculating and minimizing a loss value representing the error between the predicted result and the actual label value). The calculated gradient is then used to update the parameters of the machine learning model.

[0092] Alternatively or additionally, the training subsystem trains a model to recognize formats associated with private information. For example, the model is trained to recognize @.com and @.org as email addresses, and to recognize St. and Ave. as street names.

[0093] In some aspects, a machine learning model is trained on prepared datasets of text with varying degrees of sensitivity. For example, a prepared dataset of text related to personal financial information, medical, and health-related information may be categorized as the highest sensitivity level. This sensitive text dataset is then used to train a model to detect entities highlighted in this prepared set. A prepared set of named entities reflecting varying degrees of sensitivity is used, either alone or in combination with other entities, to train a model to detect their use and to score the sensitivity of user-provided comments.

[0094] The processing device executes the program code of the training subsystem 122 to implement blocks 402-408. For example, the program code for the training subsystem 122 stored in a non-transitory computer-readable medium is executed by one or more processing devices. Executing the code of the training subsystem 122 causes the processing device to access the training data 126A-126N from the same non-transitory computer-readable medium or a different non-transitory computer-readable medium. In some embodiments, accessing the training data involves transmitting appropriate signals between the local non-transitory computer-readable medium and the processing device via a data bus. In additional or alternative embodiments, accessing the training data involves transmitting appropriate signals between a computing system including the non-transitory computer-readable medium and a computing system including the processing device via a data network.

[0095] Example of a computational system for real-time privacy leak prediction

[0096] Any suitable computing system or group of computing systems may be used to perform the operations described herein. For example, Figure 5 An example of a computing system 500 executing the scoring subsystem 118 is depicted. In some embodiments, the computing system 500 also executes Figure 1 In some other embodiments, there are similar Figure 5 A separate computing system from those devices depicted in (eg, processor, memory, etc.) executes one or more of subsystems 112 - 122 .

[0097] The depicted example of computing system 500 includes a processor 502 communicatively coupled to one or more memory devices 504. Processor 502 executes computer-executable program code stored in memory device 504, accesses information stored in memory device 504, or both. Examples of processor 502 include a microprocessor, an application specific integrated circuit ("ASIC"), a field programmable gate array ("FPGA"), or any other suitable processing device. Processor 502 may include any number of processing devices, including a single processing device.

[0098] Memory device 504 includes any suitable non-transitory computer-readable medium for storing data, program code, or both. Computer-readable media may include any electronic, optical, magnetic, or other storage device capable of providing computer-readable instructions or other program code to the processor. Non-limiting examples of computer-readable media include disks, memory chips, ROM, RAM, ASICs, optical storage, tape or other magnetic storage, or any other medium from which a processing device can read instructions. Instructions may include processor-specific instructions generated by a compiler or interpreter from code written in any suitable computer programming language, including, for example, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, and ActionScript.

[0099] The computing system 500 may also include a number of external or internal devices, such as input or output devices. For example, the computing system 500 is shown as having one or more input / output ("I / O") interfaces 508. The I / O interfaces 508 may receive input from input devices or provide output to output devices. One or more buses 506 are also included in the computing system 500. The buses 506 communicatively couple one or more components of a corresponding one of the computing systems in the computing system 500.

[0100] The computing system 500 executes program code that configures the processor 502 to perform one or more of the operations described herein. For example, the program code includes the content retrieval subsystem 112, the NLP subsystem 114, or other suitable applications that perform one or more of the operations described herein. The program code may reside in the memory device 504 or any suitable computer-readable medium and may be executed by the processor 502 or any other suitable processor. In some embodiments, both the content retrieval subsystem 112 and the NLP subsystem 114 are stored in the memory device 504, such as Figure 5 In additional or alternative embodiments, one or more of the content retrieval subsystem 112 and the NLP subsystem 114 are stored in different memory devices of different computing systems. In additional or alternative embodiments, the program code is stored in one or more other memory devices accessible via a data network.

[0101] The computing system 500 may access one or more of the training data A 126A, training data B 126B, and training data N 126N in any suitable manner. In some embodiments, some or all of one or more of these data sets, models, and functions are stored in the memory device 504, such as at Figure 5For example, computing system 500 executing training subsystem 122 may access training data A 126A stored by an external system.

[0102] In additional or alternative embodiments, one or more of these data sets, models, and functions are stored in the same memory device (e.g., one of the memory devices 504). For example, common computing systems such as Figure 1 The privacy monitoring system 110 depicted in FIG may host the content retrieval subsystem 112 and the scoring subsystem 118 as well as the training data 126A. In additional or alternative embodiments, one or more of the programs, data sets, models, and functions described herein are stored in one or more other memory devices accessible via a data network.

[0103] The computing system 500 also includes a network interface device 510. The network interface device 510 includes any device or group of devices suitable for establishing a wired or wireless data connection to one or more data networks. Non-limiting examples of the network interface device 510 include an Ethernet network adapter, a modem, etc. The computing system 500 can use the network interface device 510 to communicate with one or more other computing devices (e.g., computers executing a computer) via a data network. Figure 1 104 depicted in the graphical interface 104) communicates with the computing device).

[0104] In some embodiments, the functionality provided by computing device 500 may be provided via cloud-based services provided by cloud infrastructure 600 provided by a cloud service provider. Figure 6 An example of a cloud infrastructure 600 is depicted that provides one or more services, including a service that provides virtual object functionality as described in the present disclosure. Such a service can be subscribed to and used by multiple user subscribers using user devices 610A, 610B, and 610C across a network 608. The service can be provided under a software as a service (SaaS) model. One or more users can subscribe to such a service.

[0105] exist Figure 6 In the depicted embodiment, cloud infrastructure 600 includes one or more server computers 602 configured to perform processing for providing one or more services provided by a cloud service provider. One or more of server computers 602 may implement Figure 16. The content retrieval subsystem 112, NLP subsystem 114, media processing subsystem 116, scoring subsystem 118, reporting subsystem 120, and / or training subsystem 122 are depicted in FIG. Subsystems 112-122 can be implemented using software alone (e.g., code, programs, or instructions executable by one or more processors provided by cloud infrastructure 600), in hardware, or a combination thereof. For example, one or more of server computers 602 can execute software to implement the services and functionality provided by subsystems 112-122, wherein the software, when executed by one or more processors of server computer 602, causes the services and functionality to be provided.

[0106] The code, program, or instructions may be stored on any suitable non-transitory computer-readable medium, such as any electronic, optical, magnetic, or other storage device capable of providing computer-readable instructions or other program code to a processor. Non-limiting examples of computer-readable media include disks, memory chips, ROM, RAM, ASICs, optical storage, tapes, or other magnetic storage, or any other medium from which a processing device can read instructions. The instructions may include processor-specific instructions generated by a compiler or interpreter from code written in any suitable computer programming language, including, for example, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, and ActionScript. In various examples, the server computer 602 may include volatile memory, non-volatile memory, or a combination thereof.

[0107] exist Figure 6 In the depicted embodiment, cloud infrastructure 600 also includes a network interface device 606 that enables communications to and from cloud infrastructure 600. In certain embodiments, network interface device 606 includes any device or group of devices suitable for establishing a wired or wireless data connection to network 608. Non-limiting examples of network interface device 606 include an Ethernet network adapter, a modem, and the like. Cloud infrastructure 600 can communicate with user device 610A, user device 610B, and user device 610C via network 608 using network interface device 606.

[0108] Graphical interfaces (e.g. Figure 1The graphical interface 104 depicted in FIG. 1 can be displayed on each of user devices A 610A, user device B 610B, and user device C 610C. The user of user device 610A can interact with the displayed graphical interface, for example, to enter text data and upload media files. In response, processing for identifying and displaying privacy alerts can be performed by server computer 602. In response to these alerts, the user can again interact with the graphical interface to edit the text data to resolve any privacy concerns.

[0109] General considerations

[0110] Numerous specific details are set forth herein to provide a thorough understanding of the claimed subject matter. However, those skilled in the art will appreciate that the claimed subject matter may be practiced without these specific details. In other instances, methods, devices, or systems known to those skilled in the art have not been described in detail in order to avoid obscuring the claimed subject matter.

[0111] Unless otherwise specifically noted, it should be understood that throughout this specification, discussions utilizing terms such as "process," "compute," "calculate," "determine," and "identify" refer to actions or processes of a computing device, such as one or more computers or one or more similar electronic computing devices, that manipulate or transform data represented as physical electronic or magnetic quantities within a memory, registers, or other information storage device, transmission device, or display device of a computing platform.

[0112] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide a result conditioned on one or more inputs. Suitable computing devices include multi-purpose microprocessor-based computer systems that access stored software that programs or configures the computing system from a general-purpose computing device to a dedicated computing device that implements one or more embodiments of the present subject matter. Any suitable programming, scripting, or other type of language or combination of languages may be used to implement the teachings contained herein in software to be used to program or configure a computing device.

[0113] Embodiments of the methods disclosed herein can be performed in the operation of such a computing device. The order of the blocks presented in the above examples can be changed - for example, the blocks can be reordered, combined and / or decomposed into sub-blocks. Certain blocks or processes can be executed in parallel.

[0114] The use of "suitable for" or "configured to" herein is intended to be open and inclusive language that does not exclude devices that are adapted or configured to perform additional tasks or steps. Furthermore, the use of "based on" is intended to be open and inclusive, as a process, step, calculation, or other action that is "based on" one or more stated conditions or values may actually be based on additional conditions or values beyond those stated conditions or values. The headings, lists, and numbering included herein are for ease of explanation only and are not meant to be limiting.

[0115] Although the subject matter has been described in detail with respect to specific embodiments thereof, it should be understood that those skilled in the art, after obtaining an understanding of the foregoing, may readily make changes, modifications, and equivalents to such embodiments. Therefore, it should be understood that the present disclosure is presented for purposes of illustration and not limitation, and does not exclude the inclusion of such modifications, variations, and / or additions to the subject matter as would be apparent to those skilled in the art.

Claims

1. A computer-implemented method comprising: detecting, by the content retrieval subsystem, entry of a set of unstructured text data into an input field of the graphical interface; In response to detecting the entry and utilizing a natural language processing subsystem, identifying a plurality of entities associated with the private information by applying at least a trained machine learning model to the set of unstructured text data in the input field; determining, by the natural language processing subsystem, an entity type for the identified entity; as well as Based on the determined entity types, the scoring subsystem assigns weights to links between entities in the graph model; calculating, by the scoring subsystem, a privacy score for the textual data by identifying connections between the entities, the connections between the entities contributing to the privacy score according to a cumulative privacy risk, wherein the privacy score is based on the weights, the privacy score indicating potential exposure of the private information by the set of unstructured textual data; as well as The graphical interface is updated in real time by the reporting subsystem to include an indicator that distinguishes a target portion of the set of unstructured text data within the input field from other portions of the set of unstructured text data within the input field, wherein modification of the target portion changes the potential exposure of the private information indicated by the privacy score.

2. The method according to claim 1, further comprising: detecting, by the content retrieval subsystem, a modification to the set of unstructured textual data entered into the input field of the graphical interface; In response to detecting the modification and utilizing the natural language processing subsystem, identifying a modified plurality of entities associated with private information by applying at least the trained machine learning model to the modified text data in the input field; calculating, by the scoring subsystem, a modified privacy score for the text data based on the modified entity; as well as The graphical interface is updated, by a reporting subsystem, based on the modified privacy score.

3. The method according to claim 1, further comprising: receiving, by the content retrieval subsystem, an image or video associated with the unstructured text data; as well as Processing the image or the video by a media processing subsystem to identify metadata, At least a subset of the identified metadata is further input into the machine learning model to identify the entity.

4. The method of claim 1 , wherein the set of unstructured text data is a first set of unstructured text data and the plurality of entities is a first plurality of entities, the method further comprising: Before receiving the first set of unstructured text data: detecting, by the content retrieval subsystem, entry of a second set of unstructured text data into the input field; as well as In response to detecting the entry and utilizing the natural language processing subsystem, identifying a second plurality of entities associated with the private information by applying at least the trained machine learning model to the second set of unstructured text data in the input field, The scoring subsystem calculates the privacy score based on connections between the first plurality of entities and the second plurality of entities. The method of claim 1 , wherein the updated graphical interface further displays an indication of the privacy score.

6. The method of claim 1 , wherein the machine learning model comprises a neural network, the method further comprising training the neural network by: Retrieving, by the training subsystem, first training data for a first entity type associated with a privacy risk from a first database; Retrieving, by the training subsystem, second training data for a second entity type associated with a privacy risk from a second database; as well as The neural network is trained by the training subsystem using the first training data and the second training data to identify the first entity type and the second entity type.

7. A computing system comprising: The content retrieval subsystem is configured to: detect entry of unstructured text data into an input field of the graphical interface; a natural language processing subsystem configured to: identify a plurality of entities associated with the private information and determine entity types for the identified entities by applying at least a trained machine learning model to the unstructured text data; a scoring subsystem configured to: assign weights to links between entities in a graph model based on the determined entity types, and calculate a privacy score for the text data by applying the graph model to the plurality of entities to identify connections between the entities, the connections between the entities contributing to the privacy score according to a cumulative privacy risk, wherein the privacy score is based on the weights, the privacy score indicating potential exposure of the private information by the unstructured text data; as well as A reporting subsystem is configured to: update the graphical interface in real time to include an indicator that distinguishes a target portion of the unstructured text data within the input field from other portions of the unstructured text data within the input field, the target portion causing the potential exposure of the private information indicated by the privacy score.

8. The computing system of claim 7, wherein: The content retrieval subsystem is further configured to: detect a modification to text data entered into the input field of the graphical interface; The natural language processing subsystem is further configured to: in response to detecting the modification, identify a modified plurality of entities associated with the private information by applying at least the trained machine learning model to the modified text data in the input field; The scoring subsystem is further configured to: calculate a modified privacy score for the text data based on the modified entity; and The reporting subsystem is further configured to update the graphical interface based on the modified privacy score.

9. The computing system according to claim 7, The content retrieval subsystem is further configured to: receive an image or video associated with the unstructured text data; Also included is a media processing subsystem configured to process the image or the video to identify metadata, Wherein at least a subset of the identified metadata is further used to identify the entity.

10. The computing system of claim 7, wherein: The text data is a first set of unstructured text data and the plurality of entities is a first plurality of entities, The content retrieval subsystem is further configured to: receive a second set of unstructured text data before receiving the first set of unstructured text data; The natural language processing subsystem is further configured to: process the second set of unstructured text data using the trained machine learning model to identify a second plurality of entities associated with the private information; and The privacy score is calculated based on connections between the first plurality of entities and the second plurality of entities.

11. The computing system of claim 7, wherein the updated graphical interface further displays an indication of the privacy score.

12. The computing system of claim 7, wherein: The machine learning model comprises a neural network; and The computing system further includes a training subsystem configured to train the neural network by: retrieving first training data for a first entity type associated with a privacy risk from a first database; retrieving second training data for a second entity type associated with a privacy risk from a second database; as well as The neural network is trained using the first training data and the second training data to identify the first entity type and the second entity type.

13. A non-transitory computer-readable medium having instructions stored thereon, the instructions being executable by a processing device to perform operations comprising: detecting entry of a set of unstructured text data into an input field of a graphical interface; a step for calculating a privacy score for the text data, the privacy score indicating potential exposure of private information by the set of unstructured text data, the step comprising determining entity types for the identified entities, and assigning weights to links between entities in a graph model based on the determined entity types, wherein the privacy score is based on the weights; as well as An indicator is updated in real time based on the privacy score, the indicator distinguishing a target portion of the set of unstructured text data within the input field from other portions of the set of unstructured text data within the input field.

14. The non-transitory computer-readable medium of claim 13, the operations further comprising: detecting a modification to the set of unstructured text data entered into the input field of the graphical interface; a step for calculating a modified privacy score for said textual data; as well as The graphical interface is updated based on the modified privacy score.

15. The non-transitory computer-readable medium of claim 13, the operations further comprising: receiving an image or video associated with the unstructured text data; as well as processing the image or the video to identify metadata, At least one subset of the identified metadata is further used to calculate the privacy score.

16. The non-transitory computer-readable medium of claim 13, wherein the set of unstructured text data is a first set of unstructured text data, the operations further comprising: Before receiving the first set of unstructured text data, detecting entry of a second set of unstructured text data into the input field; The privacy score is calculated based on the first set of unstructured text data and the second set of unstructured text data.

17. The non-transitory computer-readable medium of claim 13, wherein the updated input field further displays an indication of the privacy score.

18. The non-transitory computer-readable medium of claim 13, wherein the step of calculating the privacy score comprises using a neural network to identify entities that contribute to the privacy score, the operations further comprising training the neural network by: retrieving first training data for a first entity type associated with a privacy risk from a first database; retrieving second training data for a second entity type associated with a privacy risk from a second database; and The neural network is trained using the first training data and the second training data to identify the first entity type and the second entity type.

Citation Information

Patent Citations

  • Data privacy quantitative evaluation method based on privacy information detection

    CN110175327A

  • Sensitive data detection in communication data

    US20200336501A1