Detecting and removing predefined sensitive information types from electronic documents

By combining a context-aware filtering layer and a proxy server, efficient, accurate, and flexible data security control of the document hiding tool is achieved, solving the problems of insufficient context awareness and insufficient flexibility of proxy services in existing tools.

CN121079671APending Publication Date: 2025-12-05MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480028932.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-31
Filing Date
2024-05-20
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing document hiding tools lack context awareness, leading to inappropriate hiding and high computational resource consumption. Furthermore, traditional proxy services lack flexibility and cannot achieve fine-grained data security control.

Method used

It combines a context-aware filtering layer with existing sensitive information detection logic, uses machine learning and rule combinations to detect sensitive information, combines a proxy server to achieve selective concealment, and provides a graphical user interface to assist editing.

Benefits of technology

It improves the efficiency and accuracy of document concealment, reduces the consumption of computing resources, provides more flexible data security controls, and reduces the cost of manpower and computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121079671A_ABST
    Figure CN121079671A_ABST
Patent Text Reader

Abstract

Automatic and semi-automatic document masking techniques are disclosed herein. In some example embodiments, "context aware" masking is provided. Automated techniques are used to identify a set of potentially sensitive item (s) within a document. The potentially sensitive item (s) are filtered based on contextual information (such as entity identifiers (such as personnel identifiers, personnel group identifiers identifying a group of multiple persons, organization identifiers, etc.), resulting in a filtered set of hidden candidate (s). The filtered concealment candidate (s) may be concealed from the document, for example, automatically, or output (e.g., via a document editing graphical user interface) as a suggestion in a secondary concealment tool. Other example embodiments account for selective concealment when uploading and / or downloading a document via a proxy server to prevent, for example, intentional or unintentional publication of potentially sensitive information in a web browsing context.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to systems, methods and computer programs for detecting and removing predetermined types of sensitive information from electronic documents. BACKGROUND

[0002] The need to remove certain types of sensitive information from electronic documents arises in a variety of contexts. For example, publication of certain types of information, such as user credentials, bank details, etc., can present a security risk. As another example, privacy restrictions can require that certain types of identifying data be removed from a document before it is published. SUMMARY

[0003] Automated and semi-automated document redaction techniques are disclosed herein. In certain example embodiments, "context-aware" redaction is provided. Automated techniques are used to identify a set of potential sensitive item(s) within a document. The potential sensitive item(s) are filtered based on contextual information, such as entity identifiers (such as, for example, a person identifier, a person group identifier identifying a group of multiple persons, an organization identifier, etc.), resulting in a filtered set of redaction candidates. The filtered redaction candidates can be automatically redacted from the document, for example, or output as suggestions in an auxiliary redaction tool (e.g., via a document editing graphical user interface). Other example embodiments consider selective redaction when uploading and / or downloading documents via a proxy server, to prevent intentional or unintentional publication of potentially sensitive information in, for example, a web browsing context. BRIEF DESCRIPTION OF DRAWINGS

[0004] Illustrative embodiments will now be described, by way of example only, with reference to the following drawings in which:

[0005] Figure 1 A schematic block diagram illustrating a document redaction system is shown;

[0006] Figure 2 A schematic block diagram illustrating a document retrieval system incorporating a document redaction system is shown;

[0007] Figure 3 A schematic block diagram illustrating a proxy-based download redaction architecture is shown;

[0008] Figure 4 A schematic block diagram illustrating a proxy-based upload redaction architecture is shown;

[0009] Figure 5 A schematic block diagram illustrating a web page content request-response exchange between a client device and a proxy server is shown;

[0010] Figure 6 A schematic block diagram illustrating a proxy-based upload redaction architecture incorporating a proxy client is shown;

[0011] Figure 7 A flowchart showing a method of downloading a document from an upstream server to a client device is shown;

[0012] Figure 8 A flowchart showing a method of downloading an edited version of a document from an upstream server to a client device via a proxy server is shown;

[0013] Figure 9 A flowchart showing a method of uploading a document from a client device to an upstream server is shown;

[0014] Figure 10 A flowchart showing a method of uploading an edited version of a document from a client device to an upstream server via a proxy server is shown;

[0015] Figure 11 A flowchart showing a method for validating a redaction recommendation is shown;

[0016] Figure 12 A flowchart showing a method of filtering redaction candidates based on a redaction context is shown;

[0017] Figure 13 Various conditions that can be used in rule-based redaction automation are shown;

[0018] Figure 14 A flowchart showing an assisted redaction method is shown; and

[0019] Figure 15 A schematic block diagram of a computer system is shown. DETAILED DESCRIPTION

[0020] Improvements in data security are achieved herein through automatic or semi-automatic document redaction.

[0021] Many existing document redaction tools merely facilitate manual redaction of electronic documents. A user must manually identify (e.g., highlight) item(s) to be redacted within a document. Certain existing tools are generally capable of automatically identifying certain types of potential sensitive information in a document using some form of pattern recognition. However, such tools lack context awareness. In certain example embodiments of the present disclosure, potential sensitive items are automatically identified within a document, and then filtered based on contextual information, such as an entity (e.g., person or group, etc.) identifier. One use case is to automatically redact personal information from a document, or to automatically identify and output candidate redaction items that potentially contain personal information, but not personal information related to the identified person or group of persons. For example, a person identifier (or group of persons identifier) can be associated with a document request, or with an uploaded or downloaded document, and any identified personal item(s) determined to match the person identifier can be filtered out of a set of potential sensitive items that have been identified. Thus, in some cases, a first item and a second item can be identified within an electronic document as belonging to a predefined sensitive information category (e.g., a personal information category that is generally related to personal information or to a particular type (or multiple particular types) of personal information). However, the first item can be determined to match an entity identifier that provides context to the redaction process, triggering an exception (e.g., preventing the first item from being redacted from the document, or preventing the first item from being indicated as a redaction candidate). This context awareness reduces the likelihood of improper document redaction, which ultimately makes the process more efficient. If a document is improperly redacted, it is often not possible to retrieve the redacted information from the document (which is the purpose of redaction), meaning that the process must be repeated from scratch in that instance. In assisted redaction tools, a set of redaction candidates can be manually revised before the document is actually redacted. However, this would require additional human effort, and has a resulting cost in terms of computational resources needed to correct errors in the identification of redaction candidates. Improved redaction (whether automatic or semi-automatic) ultimately improves the speed and efficiency with which a computer system implementing a redaction method is able to achieve a desired redaction result.

[0022] Context-aware redaction can involve detecting, within an electronic document, a first item and a second item that belong to a predefined sensitive information category. Once detected, the first item can be matched against a contextual entity identifier, with the result that the first item is filtered out (meaning that it is not redacted or is not output as a redaction candidate). In this way, a lightweight context-aware filtering "layer" is applied on top of the sensitive information detection logic. This does not require any context awareness within the sensitive information detection logic, which simplifies the implementation of the sensitive information detection logic (e.g., the context-aware filtering layer can be applied on top of existing sensitive information detection logic without modifying the latter). The context-aware filtering layer can be implemented efficiently with relatively simple filtering logic (as compared to the sensitive information detection logic, which is potentially richer and can use more complex processing) using minimal computational resources on the computer device implementing the filtering. This in turn avoids the high cost (in terms of time and computational resources) that would be required to build a context-aware sensitive information detector. Decoupling sensitive information detection from context-aware filtering in this way also provides greater scalability, as the sensitive information detection logic can be more easily refined and / or extended to new types of sensitive information or new sensitive information categories, etc. (e.g., through retraining where machine learning techniques are used), which can not require any modification to the context-aware filtering layer, or only direct modifications (e.g., to incorporate a new type of entity identifier).

[0023] When implemented in an assisted (semi-automatic) redaction tool, the refinement of redaction candidates presented via a graphical user interface (GUI) provides an improved human-machine interaction, as less human effort is required to manually complete and edit the redaction candidates. Such embodiments provide an improved document editing GUI as compared to existing redaction tools, which require the user to manually identify redaction candidates, or in the case of redaction tools that can automatically identify redaction candidates but lack context awareness, to manually remove inappropriate redaction candidates according to context.

[0024] Certain embodiments implement selective cloaking of documents uploaded and / or downloaded via a proxy server. In some deployment scenarios, a proxy server is "invisible" between a client device and an upstream server. Existing proxy architectures tend to be based on an "all or nothing" approach whereby a download / upload is allowed or blocked according to a download / upload policy. However, in the present context, selective cloaking of documents by the proxy server provides a more granular control, e.g., an upload or download action can be allowed, but the uploaded or downloaded document can be selectively cloaked (e.g., by "masking" some portion(s) of the document) to prevent sharing of unauthorized information. This approach provides improved data security, but with greater flexibility than traditional proxy-based approaches. Existing proxy services can provide improved data security (e.g., by blocking uploads / downloads related to certain websites, etc.), but can be overly burdensome to end users, particularly if uploads / downloads are needlessly blocked. The present technology can implement a given level of data security with respect to sensitive information, but in a manner that is less detrimental to the overall end user experience.

[0025] Figure 1 A schematic block diagram of a cloaking system 100 is shown. The cloaking system 100 is shown as including a document search component 106, a sensitive item detector 107, a filter component 108, and a cloaking component 110. The components 106, 107, 108, 110 are functional components, which can be implemented, e.g., in the form of code executed on a processor (or processors) of the cloaking system 100 (not shown). Such code can be stored in a memory (or memories) coupled to the processor(s), and configured to cause the processor(s) to implement the described functionality when executed on the processor(s).

[0026] The cloaking system 100 applies a context-aware cloaking process to the electronic document 102 in the manner described below.

[0027] The document search component 106 is configured to receive the electronic document 102 and search it for any“sensitive terms” that it can contain. A sensitive term refers to a document portion that is determined to belong to a predefined sensitive information category, such as a personal information category. Sensitive information can for example include user biometrics, user credentials, names, birth dates, addresses, phone numbers, identification numbers (e.g. passports, identity cards, social security, etc.), bank account details, private company information, etc. Such information types can be sensitive because, for example, they pose a security risk in the hands of a malicious user, because of user privacy considerations, or due to confidentiality considerations. The sensitive information categories can be relatively broad (e.g. “person identifiers” can be a single category, encompassing a broad variety of sensitive information types) or specific (e.g. with multiple separate categories for different forms of person identifiers). An“entity” in this context can refer to a person, but can also refer to other types of entities, such as organizations (e.g. companies), devices, etc.

[0028] The sensitive term detector 107 is associated with a predefined sensitive information category. The document search component 106 uses the sensitive term detector 107 to identify any sensitive (or potentially sensitive) terms within the electronic document 102 that belong to its associated sensitive information category. The sensitive term detector 107 can for example be a machine learning (ML) component that has been trained on examples of sensitive terms within this predefined sensitive information category. In this case, the sensitive information category can be implicitly defined in the selection of examples used to train the sensitive term detector 107. Alternatively, the sensitive term detector 107 can be a rule-based component, in which case the sensitive information category can be explicitly defined in the encoded rules in the sensitive term detector 107. Alternatively, a combination of ML-based and rule-based sensitive term detection can be used. Pattern detection (based on ML and / or rules) can be used to detect such terms within the electronic document 102. In some embodiments, multiple sensitive term detectors can be provided, which are associated with different sensitive information categories (e.g. different types of personal information).

[0029] The document search component 106 outputs a redaction candidate set 109. The redaction candidate set 109 contains or references any sensitive term(s) that the document search component 106 has located within the electronic document 102. Such terms are referred to as“redaction candidates” because at this stage they have not been redacted from the electronic document 102. Instead, the filter component 108 applies context-aware filtering to the redaction candidate set 109 to selectively remove term(s) from the redaction candidate set 109 before the electronic document 102 is redacted.

[0030] The filter component 108 receives the redaction candidate set 109, and additionally receives the redaction context 104 that is relevant to the electronic document 102.

[0031] In this example, the obfuscation context 104 is shown to include an entity identifier (eID) associated with the electronic document 102. The eID provides relevant context to the obfuscation process. For example, the eID can be a personnel identifier associated with the electronic document 102, or a personnel identifier associated with a request for the electronic document to be obfuscated prior to being published. The following examples consider eIDs that belong to a sensitive information category associated with the sensitive item detector 107. Thus, if the eID (or a detectable variant of the eID) appears somewhere in the content of the electronic document 102, the eID can be detected by the sensitive item detector 107 when the sensitive item detector 107 is applied to the electronic document 102. As such, the obfuscation candidate set 109 can include sensitive items that include the eID or some variant of the eID.

[0032] However, in certain contexts, obfuscating the eID from the electronic document 102 can be inappropriate or undesirable. For example, the eID can be an identifier of a person who has submitted a request for a copy of any document maintained within a document storage system that includes their personal information. In this case, it would be inappropriate to obfuscate instance(s) of the eID from the electronic document 102. However, in certain contexts, it can be necessary or desirable to obfuscate identifiable information of any other person (or other entity), referred to as “third-party” information.

[0033] The filtering component 108 searches the obfuscation candidate set 109 for any items that match the eID, and removes from the obfuscation candidate set 109 any items that are determined to match the eID. Such items can be identified via hard (exact) matching or soft matching, or via a combination of hard matching and soft matching. In some cases, multiple eIDs can be received, such as a person’s name and phone number, and used to filter the obfuscation candidate set 109. For example, an eID (e.g., a name or username) can be received, and used to locate one or more additional eIDs associated with the received eID (e.g., a phone number, email address, birth date, etc. associated with the name or username). Such additional eID(s) can be located, for example, in a database(s) of user information. With multiple eIDs, the following description applies to each ID that forms part of the obfuscation context 104. Thus, the eID associated with a message can be included in the message, or not included in the message but associated with another identifier included in the message (e.g.).

[0034] In the depicted example, the document search component 106 identifies a first item 109A and a second item 109B, each of which is determined to belong to a sensitive information category associated with the item detector 107. Thus, the first item 109A and the second item 109B are included in the obfuscation candidate set 109.

[0035] The first item 109A does contain the eID (or some variant thereof) of the redacted context 104. The filtering component 108 matches the eID to the redacted candidate set 109 that includes the first item 109A, and in response removes the first item 109A from the redacted candidate set 109.

[0036] The second item 109B relates to a different entity, meaning that the filtering component is unable to match the second item 109B to the eID of the redacted context 104.

[0037] The filtering component outputs 108 a filtered item set 111 that contains or references any items of the redacted candidate set 109 that have not been removed. In this example, the redacted candidate set 109 is shown to include the second item 109B, but not the first item 109A that matches the eID of the redacted context 104.

[0038] The redaction component 110 receives the filtered item set 111 and uses the filtered item set 111 to generate an edited document 112 that is a redacted version of the electronic document 102. The edited document 112 is generated by removing at least one sensitive item from the electronic document 102 or modifying the item so that it is no longer sensitive. For example, the item or some portion (or portions) of the item can be removed and optionally replaced with other context such as an image (e.g., a black box) or placeholder text (e.g., predetermined character(s) or string(s) of characters or randomly generated text). Note that any redacted item is not simply blurred visually, but is actually removed or modified so that the original item is no longer derivable from the edited document 112.

[0039] In some embodiments, the context-aware redaction process is fully automatic. In this case, the redaction component 110 automatically redacts each item of the filtered item set 111 from the electronic document 112. In other embodiments, an option for manual review is provided (referred to herein as “assisted” redaction). In this case, the filtered item set 111 can be more preliminary to a final redaction via user input of the redaction system 100, and the final redaction is also initiated via user input. For example, the filtered item set 111 can be visually indicated on a graphical user interface (GUI) associated with the redaction system 100 (not shown), and the filtered item set 111 can be modified via input of the GUI.

[0040] A copy of the original (unredacted) document 102 is retained, allowing (among other things) for different redacted versions of the document to be generated in the future based on different redaction contexts.

[0041] Figure 2 A document is shown to containFigure 1 An example document retrieval system 200 of the anonymization system 100. The document retrieval component 232 of the document retrieval system 200 receives a document search request 231 from a client device 230 that includes or indicates an entity identifier (eID) (e.g., identifying a person, device, or organization).

[0042] In the context of Figure 2 In the context of

[0043] The document retrieval component 232 conducts a search of the document store 234 (e.g., a database or multiple databases) to retrieve from among any documents therein that are found to satisfy the document search request 231. For example, with a person ID that identifies a person, the document retrieval component 232 can search for any documents containing personal information about the identified person. One or more other criteria can be applied, e.g., to limit the scope of the search or to exclude certain types of documents. As noted above, the search can alternatively or additionally be based on an eID (or eIDs) that is not contained in the document search request 231 but is otherwise indicated by the document search request 231 (e.g., an eID stored elsewhere in association with some other eID contained in the message).

[0044] Assuming the document retrieval component 232 finds at least one document 202 that satisfies the document search request 231, in one implementation the retrieved document 202 is automatically passed to the anonymization system 100 along with the anonymization context 204 that includes the eID. In another implementation, this step is subject to manual review of any retrieved documents (e.g., to identify irrelevant documents or gaps in the search before the documents 202 are passed to the anonymization system 100 along with the anonymization context 204). If multiple documents are identified (and, if applicable, approved for release in the manual review), each document is passed to the anonymization system 100 for sequential or parallel processing.

[0045] Upon receiving the document 202, the redaction system 100 uses the redaction context 204 to identify and filter redaction candidates. Note that in this example, the eID is included in the redaction context 204. Thus, in this example, the eID is used both to locate the document 202 and to provide context for redaction thereto. One use case is a person’s request for a document containing their own personal information. The person making the request is identified by the person identifier included in or otherwise indicated in the document search request 231. The goal in this case can be to publish any such requested document (e.g., to the extent defined by one or more document publication criteria, e.g., based on legal requirements regarding personal data), and to leave the requesting user’s personal information in such document, but to redact personal data (and / or other type(s) of sensitive information that can be identified, e.g., confidential information) of any other person identified in the same personal information category, for example.

[0046] In one implementation, the redaction candidates are identified, filtered, and any redaction candidate(s) remaining after filtering are automatically redacted. In another implementation, the redaction system 100 outputs or indicates via a user interface any redaction candidate(s) remaining after filtering. In this case, the redaction system 100 can receive user input and modify the filtered set of redaction candidates (e.g., add, remove, and / or modify one or more redaction candidates) prior to final redaction. Either way, the result is at least one edited document 212, which is communicated to the client device 230 (e.g., with one or more messages containing the edited document 212, or by way of a link indicating a storage location where the edited document 212 is stored, and the client device 230 can retrieve the edited document 212 from the storage location, for example).

[0047] Another deployment scenario is considered below, which involves a client device operating “behind” a proxy server. The proxy server implements a proxy service, such as a web proxy service through which web content is proxied (the term web proxy server can be used in this context). For example, incoming / outgoing network traffic to / from the client device can be routed via the proxy server, and the proxy server can selectively filter or block traffic according to a policy (or set of multiple policies). An example is described below that considers a document redaction policy applied to downloaded and / or uploaded documents.

[0048] Figure 3A schematic block diagram of a proxy download scenario with context-aware redaction using the redaction system 100 is shown. The client device 330 sends a download request 331 that includes a destination address corresponding to an upstream server 334. The download request 331 is intercepted by a proxy server 332, and in response to the download request 331, the proxy server 332 sends a proxied download request 333 to the upstream server 334. The proxied download request 333 includes a modified source address corresponding to the proxy server 332. For example, the download request 331 can include a source address corresponding to the client device (e.g., the IP address or other network address of the client device 330 in one or more source fields of the download request 331) that is replaced in the proxied download request 333 with the IP address (or other network address) of the proxy server 332. The modified source address causes the upstream server 334 to send a response to the proxy server 332 rather than the client device 330.

[0049] The response includes a document 302, and the proxy server 332 initiates selective redaction of the document 302 based on a download redaction policy 303. In this case, the redaction system 100 can be implemented as part of the proxy server 332, or as a separate (e.g., external) service that is accessible to the proxy server 332. The proxy server 332 derives a redaction context 304 from the download request 331, e.g., to extract (or otherwise obtain based on) an eID associated with the document 302 from the download request 331. For example, the eID can identify an entity that has initiated the download of the document 302. For example, the eID can be a user identifier or device identifier that is included in or otherwise indicated by the download request 331 and / or that is associated with the client device 330 (e.g., at the client device itself, or in a backend system that holds user / device details).

[0050] The proxy server 332 passes the document 302 to the redaction system 100 along with the redaction context 304. The redaction system 100 uses the redaction context 304 to selectively redact the document 302, resulting in an edited document 312. For example, the redaction system 100 can be configured to redact personal information from the document, with the exception of personal information associated with the user identifier in the redaction context 304 (which may, for example, identify a user of the client device 330; meaning that the user’s information is not redacted, but other personal information is redacted).

[0051] Note that where the eID identifies an entity that has initiated the download, the redaction of the document 302 is tailored to the entity that attempted to download the document 302.

[0052] The proxy server 332 sends the redacted document 312 to the client device 330 in response to the download request 331, rather than sending the (unredacted) document 302 received from the upstream server 334 to the client device 330 in response to the original download request 331.

[0053] Figure 4 A schematic block diagram of a proxy upload scenario with context-aware redaction using the redaction system 100 is shown. In this case, the proxy server 432 receives an upload request 431 from a client device 430. The upload request 431 comprises a document 402 to be uploaded to an upstream server 434. For example, the upload request 431 can be an HTTP POST request comprising the document 402 to be uploaded. The proxy server derives a redaction context 404 from the upload request 431 (e.g. by extracting (or otherwise obtaining based on) an eID associated with the document 402 from the upload request 431. For example, the eID can identify an entity that has initiated the upload of the document 402. For example, the eID can be a user or device identifier contained in or otherwise indicated by the upload request 431 and / or associated with the client device 430 (e.g. at the client device itself, or in a backend system holding details of the user / device).

[0054] The proxy server 432 passes the document 402 from the upload request 431 to the redaction system 100 along with the redaction context 404 derived from the upload request 431. The redaction system 100 can be implemented locally at the proxy server 432, or as a separate (e.g. external) service accessible to the proxy server 432. The redaction system 100 uses the redaction context 404 to selectively redact the document 402 based on the upload redaction policy 403, resulting in a redacted document 412. The proxy server 432 sends a proxied upload request 433 comprising or otherwise indicating the redacted document 412 to the upstream server 434, meaning that the redacted document 412 rather than the (unredacted) document 402 is uploaded to the upstream server 434. The upstream server 434 can for example store the redacted document 412 in a network (e.g. cloud) storage location.

[0055] This approach can for example be used to allow a given user to share their own personal information via document uploads (to the extent permitted by the upload redaction policy 403), but to prevent them from intentionally or unintentionally sharing personal information and / or other types of sensitive information (e.g. confidential information) about other people.

[0056] Note that where the eID identifies an entity that has initiated the upload, the redaction of the document 402 is tailored to that entity that attempted to upload the document 402.

[0057] In some implementations, a proxy client executing on the client device 430 detects the upload event and signals the upload event to the proxy server 432, causing the proxy server 432 to apply selective obfuscation to the document 402.

[0058] Figure 5 An illustrative overview of the proxy client injection scenario is provided. Figure 4 The client device 430 sends a content request 500 (such as an HTTP request) intended for the upstream server 434, e.g., requesting web page content indicated in the content request 500. The content request includes a resource identifier (e.g., a uniform resource locator (URL) or uniform resource identifier (URI)) identifying the requested web page content 505. The proxy server 432 intercepts the content request 500 and replaces the request 500 with a proxied content request 502 (e.g., replacing the first source address of the client device 430 with a second source address of the proxy server 432). In response to the proxied content request 502, the upstream server 434 returns a response 504 including the requested web page content 505 to the proxy server 432. The proxy server 432 receives the response 504 and injects a proxy client 507 in the response 504, resulting in a modified response 506 including the modified web page content, which in turn includes the requested web page content 505 and the proxy client 507. The proxy server 432 sends the modified response 506 to the client device 430 in response to the content request 500. The proxy client 507 is in the form of executable proxy client code (such as JavaScript code) adapted to be executed on the client device 430. The proxy client 507 is executed on the client device 430 when rendering the requested web page content 505. The requested web page content 505 may, for example, include a web page with an upload field or other document upload functionality for sending an upload request 431. Figure 4

[0059] Figure 6 The proxy client 507 running on the client device 430 is shown. In this example, the proxy client 507 is executed on the client device 430 when the requested web page content 505 is rendered. Figure 4 ​The upload tag 600 is inserted in the form of marker data included with the uploaded document 402 in the upload request 431. The proxy client 507 is configured to detect initiation of a document upload function in the requested web page content 505 at the client device 430 and in response insert the upload tag 600. The upload tag 600 signals to the proxy server 432 that the upload request 431 contains an uploaded document. The proxy server 432 detects the upload tag 600 in the upload request 431 and in response initiates selective obfuscation of the uploaded document 402 based on the upload obfuscation policy 403 in the manner described above.

[0060] Note that the term server is used in a broad sense to include not only a single server device but also a group of multiple server devices used to implement an application or deliver a service to a client device. For example, an upload server can include multiple server devices (sharing a network address or having different network addresses), and in some cases a first server device that receives a proxied content request can be different from a second server device that receives a proxied upload request. As another example, a proxy server can be implemented as a single proxy server device or as multiple proxy server devices.

[0061] Figure 7 A flowchart showing a method of downloading a document from an upstream server 730 to a user device 710 without using a proxy server is shown in context. At step 701, a web page is served to the browser 720. The web page contains a link to a document (e.g., docx, pdf, pptx, etc.). At step 702, user input selecting the link to the document is received, causing the browser 720 to send a request to the upstream server 730 at step 703 to retrieve the content of the document. The upstream server 730 receives the request at step 704 and responds with the content of the document at step 705. At step 706, the browser triggers a download action on the document content and saves it as a file to the local file system at step 707. At step 708, the user can then open the document using a desktop application separate from the browser.

[0062] Figure 8 A flowchart showing a method of downloading a document from an upstream server 860 to a user device 830 by using a proxy server 850 equipped with document obfuscation capabilities is shown. For example, the obfuscation system can run on the proxy server 850 or on a separate server in communication with the proxy server 850.

[0063] At step 801, a webpage of the web browser 840 contains a link to a document (e.g., docx, pdf, pptx, etc.). At step 802, the user selects the link to the document, causing the browser 840 to send a content request (e.g., HTTP request) to retrieve the content of the document at step 803. The proxy service 850 intercepts the request at step 804 and verifies that the response is a navigation request that can ultimately become a browser download action at step 805. The upstream server 860 receives the request at step 806 and responds with the content of the document at step 807. The proxy service 850 intercepts the response and detects that the response content type represents a document at step 809.

[0064] The administrator user 820 can log into the security and compliance portal of the proxy server at step 821 to configure session policies regarding downloads at step 822 to redact text and / or other content in documents based on specific keywords.

[0065] At step 810, the proxy service 850 finds a matching session policy from the session policies configured by the administrator at step 820 to redact text from the document. The proxy service 850 then parses the document content at step 811 (e.g., using a utility parsing method), finds text regions and / or other items that match the policy filters at step 812, and redacts the text (e.g., replaces the text with black rectangles at step 813). The modified document is reconstructed at step 814, and the modified document content is returned at step 815.

[0066] The browser 840 triggers a download action with the document content at step 816, and saves the document content as a file to the file system at step 817. At step 818, the user opens the document using a desktop application (Microsoft Word, Adobe Acrobat, Microsoft PowerPoint, etc.). At step 819, the user cannot view the redacted text and cannot extract any confidential content.

[0067] Figure 9 A flowchart illustrating a method of uploading a document from a user device 910 to an upstream server 930 without using a proxy server is shown. At step 901, a webpage of the browser 920 contains an input of type file. At step 902, the user clicks the input and selects a file from the local machine at step 910, and submits the upload form at step 903. The browser 920 sends an HTTP POST request with the content of the file at step 904. At step 905, the upstream server 930 receives the file for processing.

[0068] Figure 10A flowchart showing a method of uploading a document from a user device 1030 to an upstream server 1060 by using a proxy server 1050 is shown. At step 1001, a web page of the browser 1040 contains an input of type file. At step 1002, the user clicks on the input and selects a file from the local machine 1030 and submits the upload form at step 1003. At step 1004, the file is uploaded to the browser 1040.

[0069] At step 1006, the proxy client component 1005 detects the action of uploading a file into the browser 1040. At 1007 the browser 1040 sends an HTTP POST request with the content of the file.

[0070] The proxy client component 1005 adds at step 1008 an invisible input element for tagging the HTTP POST request, in this example corresponding to the upload tag 600 of Figure 6 For example, a "hidden" type of input element can be used. These types of elements allow web page developers to include data that cannot be seen or modified by the user when submitting the form. For example, the ID of the content that is currently being sorted or edited or a unique security token. Hidden inputs can also be used to store and submit security tokens or secrets for security purposes.

[0071] The proxy service 1050 intercepts at step 1009 the request from the browser 1040 and verifies at step 1010 that the request contains the input parameter added by the proxy client component 1005. The proxy service 1050 extracts at step 1011 the content of the document based on the hint in the input added by the proxy.

[0072] At step 1021, the administrator 1020 can log into the security and compliance portal of the proxy server 1050 to configure at step 1022 a session policy about the upload to redact text in the document based on specific keywords.

[0073] After extracting the content of the document at step 1011, the proxy server 1050 finds at step 1012 a matching session policy from the session policies configured by the administrator 1020 to redact text in the document. The proxy server 1050 then parses the document content at step 1013 (e.g., using a utility parsing method), finds the text region that matches the filter 1014 of the policy, and replaces the text with a black rectangle at 1015. The document is reconstructed with the modification at 1016 and the content of the request is updated at step 1017. The upstream server 1060 receives at step 1018 the modified (with redacted text) file for processing.

[0074] Figure 11A flowchart depicting a process for checking a recommendation for redaction of an item to be redacted at step 1100 using tags, policies, and zone redaction is shown. If the item under review is found to contain a tag or policy indicating a hold at step 1101, the item is flagged as a record requiring redaction / exception review at step 1102. If the item under review is found to contain a tag or policy indicating that the item is sensitive at step 1103, the item is flagged as confidential for redaction / exception review at step 1104. If the item is found to contain an email subject at step 1105, the subject is used as a zone redaction candidate at step 1106.

[0075] The method also allows for customization of redaction recommendations based on defined redaction context. Entity identifiers are used to represent items that should not be part of the redaction process. A document request can indicate entity identifiers to exclude from the redaction process.

[0076] Figure 12 A flowchart of a selective redaction process is shown. At step 1200, a document is searched for sensitive items, resulting in a set of redaction candidates. When a sensitive item is detected (step 1201), a check is performed to determine if the sensitive item matches an entity identifier that is excluded from the redaction process. If a match is found at step 1202, the sensitive item is filtered from the set of redaction candidates at step 1203. This default can be changed. For example, by an administrator. If no match is found at step 1202, the sensitive item is re-added as a redaction candidate in the set of redaction candidates at step 1204. The identification and filtering of redaction candidates can be performed in separate stages (e.g., the document can be searched to establish a full set of redaction candidates, which are then filtered), or they can be interleaved (e.g., each time a redaction candidate is found, it can be checked against one or more entity identifiers that apply to the redaction process, and if a match is found, it is filtered out at that point). For example, the candidate redaction items (which have not been filtered out based on contextual input) can be indicated by visual markers within the document itself (e.g., by automatically highlighting each candidate item within the document). The visual marker(s) can be modified or removed based on user input, and / or additional marker(s) can be added to add candidate redaction item(s) to the set of redaction candidates, prior to final redaction.

[0077] In some embodiments, the method allows custom item(s) or string(s) to be added to the search. In the case where a custom item is found, the number of instances of that custom item can be obtained. Custom items representing additional search terms are processed in a similar manner to recommendations. The properties of custom items allow them to be distinguished from recommendations. These custom item(s) or string(s) can be saved so that they can be viewed for any given item and modified at any time when the open is requested for review. Similarly, custom item(s) or string(s) added to the list can be removed, which will automatically undo any redaction or redaction actions that have been performed based on these custom items or strings.

[0078] In some embodiments, the method allows both recommendations and custom items to be visually identified during the item review process. Visual redactions are created for recommended items during the review experience without any substantive changes to the item. If needed, the method allows the visual redactions to be turned off during the review process. Visual redactions can be refreshed by turning off the visual redaction option and then turning it back on. This is useful when the item is rescanned on demand for redaction requests.

[0079] In some embodiments, the method provides a detailed view of custom items and recommendations for items being searched. A list of all recommendations can be provided for a single item or multiple items. These recommendations can be grouped or filtered based on various factors such as: classification type(s), confidence level of system recommendation, value, frequency of occurrence within content, location. As can be seen, each individual recommendation of an item can be displayed separately from the document with the surrounding document content (e.g., a predetermined number of characters before and after the detected result). During the review process, any recommendation within an item can be jumped to without having to review each recommendation in order.

[0080] In some embodiments, the method allows certain actions to be applied to custom items and recommendations, such as applying redaction, modifying redaction annotations, or removing any applied redaction. The recommendations can be acted on using the visual redactions described above. As can be seen, the action taken is immediately reflected within the review of the item(s). The specified action can be taken on a single instance of a recommendation or on multiple / all instances of a recommendation. For all recommendations that fall into a particular sensitive information type (e.g., all credit card numbers), it can be possible to take actions in bulk to redact, annotate redaction, or remove redaction. The method allows actions to be taken in bulk to redact, annotate redaction, or remove redaction for all recommendations based on various factors such as classification type, confidence level of recommendation, and value of recommendation. The privacy administrator is able to document the reasons for editing these redactions.

[0081] In some embodiments, the method allows for transparency of anonymization on demand during the review process without the need to remove the anonymization. The recommended action(s) can be updated at any time while the scheme is in a state that allows review and modification. For example, anonymization can be removed, performed, changed, or annotated. The method provides the ability to obtain how much anonymization has been performed in a single item or multiple or all items, and the ability to understand the difference in types of anonymization (custom search and custom anonymization, recommended anonymization, manual area anonymization). It can be discovered how many anonymizations were recommended and how many were taken. Anonymization can be statistically categorized by multiple pivots such as personal data type, value, location, frequency of occurrence, confidence.

[0082] In some embodiments, the method provides automatic customized recommendations for the anonymization process based on rules and / or policies and / or saved settings, suppress recommendations based on category(s), value(s), custom item(s), recommended confidence(s), add recommendations when found based on manually added custom item(s), and add recommendations based on machine learning patterns of anonymization behavior. Default automation of recommending anonymization can be configured based on various factors such as category type(s), confidence level(s), and value(s). For example, the automatic anonymization process can be programmed to “always anonymize,” “always anonymize + annotate,” or add a specific character count to anonymize before and / or after the recommended value.

[0083] Figure 13 It is shown how rules can be used at step 1300 to take automatic anonymization actions based on confidence level(s) (step 1301), personal data type(s) (step 1302), public area or custom value such as email header (e.g., from / to / CC / BCC), custom item(s) (step 1303) or string(s), file type(s) (step 1304), location (step 1305) (e.g., mailbox, site path, etc.), and by instance count (step 1306).

[0084] Figure 14A flowchart showing an example assisted redaction process is shown. At step 1400, the privacy administrator performs a document retrieval request associated with the (multiple) eIDs and receives a list of returned items in response at step 1401. When the privacy administrator selects an item from the list to review the item at step 1403, they can see recommended redactions at step 1404 where items such as personal data that do not match the associated eID (or any associated eIDs) have been found recommended for redaction in their view. The privacy administrator receives insight at step 1405 as to what data types are recommended in the selected item, what values are recommended, instance counts for those values, and confidence levels for those recommendations. The privacy administrator can select to redact all instances of a particular item type at step 1406, such as a particular personal data type (e.g., all social security numbers), all instances of a particular value (e.g., all occurrences of the value "John Doe"), all recommendations with a particular confidence threshold (e.g., all items detected by the system with "high" confidence or above an X% confidence). The privacy administrator can also see particular instances within context and select whether they want to redact only particular instances. The privacy administrator can navigate to any recommendation to view it within a preview pane for additional review or to make a broader redaction that overlays contextual information around the recommendation. The privacy administrator can also perform a manual addition of a value at step 1407 to search and redact the value across all data collected. Insights related to the value are then added.

[0085] The privacy administrator can see insights about the auxiliary anonymization activity for document access requests. The privacy administrator can see how many anonymizations for requests were made, and can see categorized statistics of the types of data actually anonymized and aggregate counts for each type. The privacy administrator can see insights into the confidence scores for anonymizations. This gives the privacy administrator a good understanding of what work has been done by the automated anonymization and where they can want to focus additional review. The privacy administrator can interact with any of these insights, which will bring a filtered list of the relevant items into their view (e.g., the privacy administrator can select the lowest confidence level insight to review those items in detail). At step 1408, the privacy administrator can also view all anonymized values, sorted to show the most frequently anonymized values first. This allows the privacy administrator to briefly review whether there are any anonymizations that are not in place. The privacy administrator can select to remove anonymization from any given value here, which will be performed in bulk on the review set. The privacy administrator can also select to choose values and view the files with those particular anonymizations for additional confirmation or modification. At any node in reviewing files with anonymizations, the privacy administrator can view what the value under the anonymization is, and can also select to remove that anonymization on demand. The privacy administrator can also select to remove anonymization at the file level, for multiple file selection, or for all files collected. When performing a de-anonymization activity, the privacy administrator can be prompted at step 1409 to add a comment that will be automatically saved in the file annotation.

[0086] When the privacy administrator performs an export at step 1410, the automated anonymization system can, for example, export the items to a format file that ensures that the copy of the data provided to the requesting entity (e.g., user or device) cannot be de-anonymized, and has the visible anonymizations placed by the administrator during review. The plaintext unedited copy of that information will not be included in the export package. At step 1411, the export file will be delivered to the data subject.

[0087] Figure 15 A non-limiting example of a computing system 1500, such as a computing device or system of connected computing devices, which can implement one or more of the above-described methods or processes, including filtering of data and implementation of the above-described structured knowledge base, is schematically shown. The computing system 1500 is shown in simplified form. The computing system 1500 includes a logic processor 1502, a volatile memory 1504, and a non-volatile storage device 1506. The computing system 1500 can optionally include a display subsystem 1508, input subsystem 1510, communication subsystem 1512, and / or other components not shown in FIG. 15. These components can be coupled by an interconnection element such as a bus 1514. Figure 15other components not shown. The logic processor 1502 includes one or more physical (hardware) processors configured to perform processing operations. For example, the logic processor 1502 can be configured to execute instructions as part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. The logic processor 1502 can include one or more hardware processors configured to execute software instructions based on a set of instruction architecture, such as a central processing unit (CPU), a graphics processing unit (GPU), or other form of accelerator processor. Additionally or alternatively, the logic processor 1502 can include hardware processors in the form of logic circuitry or firmware devices configured to execute hardwired logic or firmware instructions, programmable or non-programmable. The processor(s) of the logic processor 1502 can be single-core or multi-core, and the instructions executed thereon can be configured for sequential, parallel, and / or distributed processing. Individual components of the logic processor optionally can be distributed among two or more separate devices, which can be remotely located and / or configured for coordinated processing. Aspects of the logic processor 1502 can be virtualized and executed by remotely accessible, networked computing devices configured in a cloud-computing configuration. In such a case, multiple physical logic processors are functioning together as a cloud-computing platform to execute the instructions. The non-volatile storage device 1506 includes one or more physical devices configured to hold instructions executable by the logic processor 1502 to implement the methods and processes described herein. When such methods and processes are implemented, the state of the non-volatile storage device 1506 can be transformed— e.g., to hold different data. The non-volatile storage device 1506 can include removable and / or built-in physical devices. The non-volatile storage device 1506 can include optical memory (e.g., CD, DVD, HD-DVD, Blu-ray Disc, etc.), semiconductor memory (e.g., ROM, EPROM, EEPROM, FLASH memory, etc.), and / or magnetic memory (e.g., hard-disk drive), or other mass storage device technology. The non-volatile storage device 1506 can include nonvolatile, dynamic, static, read / write, read-only, sequential-access, location-addressable, file-addressable, and / or content- addressable devices. The volatile memory 1504 can include one or more physical devices configured to hold instructions executable by the logic processor 1502. The volatile memory 1504 is generally used to store temporary variables or other intermediate information used during the execution of instructions by the logic processor 1502. The logic processor 1502, the volatile memory 1504, and the non-volatile storage device 1506 can each be integrated into one or more hardware-logic components, such as a chipset, as opposed to existing in separate components. Such hardware-logic components can include one or more physical processors, ASICs, FPGAs, or other components configured to perform the methods and processes described herein.The terms "module," "program," and "engine" can be used to describe an aspect of computing system 1500 typically implemented in software by a processor using a portion of the volatile memory for performing specific functions in relation to transforming the processor to perform a function. Thus, a module, program, or engine can be instantiated via the execution of one or more "instructions" and / or "code" that, for example, is held by non-volatile storage device 1506 and that causes a logic processor 1502 to perform specific functions. Different "modules," "programs," and / or "engines" can be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Likewise, the same module, program, and / or engine can be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module," "program," and "engine" can encompass individual or groups of executable files, data files, libraries, drivers, scripts, database records, etc. When included therein, display subsystem 1508 can be used to present a visual representation of data held by non-volatile storage device 1506. This visual representation can take the form of a graphical user interface (GUI). As the herein described methods and processes change the data held by the non-volatile storage device, and thus transform the state of the non-volatile storage device, the same can also be said of the display subsystem 1508 as the state thereof is likewise transformed to visually represent the changes in the underlying data. Display subsystem 1508 can include one or more display devices utilizing virtually any type of technology. Such display devices can be combined with logic processor 1502, volatile memory 1504, and / or non-volatile storage device 1506 in a shared enclosure, or such display devices can be peripheral display devices. When included therein, input subsystem 1510 can comprise or interface with one or more user-input devices such as a keyboard, mouse, touch screen, or game controller. In some embodiments, the input subsystem can comprise or interface with selected natural user input (NUI) componentry. Such componentry can be integrated or peripheral, and the transduction and / or processing of input actions can be handled local or remote from a user's computer. Example NUI componentry can include a microphone for speech and / or voice recognition; an infrared, color, stereoscopic, and / or depth camera for machine vision and / or gesture recognition; a head tracker, eye tracker, accelerometer, and / or gyroscope for motion detection and / or intent recognition; as well as electric-field sensing componentry for assessing brain activity; and / or any other suitable sensor. When included therein, communication subsystem 1512 can be configured to communicatively couple various computing devices described herein with each other, and with other devices. Communication subsystem 1512 can include wired and / or wireless communication devices compatible with one or more different communication protocols. As non-limiting examples, the communication subsystem can be configured for communication via a wireless telephone network, or a wired or wireless local- or wide-area network.In some embodiments, the communication subsystem can allow the computing system 1500 to send messages and / or receive messages to and from other devices via a network, such as the Internet. As used herein, the term computer-readable media can include computer-storage media. Computer-storage media can include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, or program modules. Computer-storage media can include RAM, ROM, Electrically Erasable Read-Only Memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other article of manufacture which can be used to store information and which can be accessed by the computing device (e.g., the computing system 1500 or a component device thereof). Computer-storage media does not include a carrier wave or other propagated or modulated data signals. Communication media can be embodied by computer readable instructions, data structures, program modules, or other data by way of modulated data signals, such as carrier waves or other transport mechanisms, and includes any information delivery media. The term "modulated data signal" can describe a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media can include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.

[0088] In a first aspect disclosed herein, a computer-implemented method, comprising: obtaining an electronic document and an entity identifier associated with the electronic document, the entity identifier pertaining to a predefined sensitive information category; detecting, within the electronic document, a first item belonging to the predefined sensitive information category; detecting, within the electronic document, a second item belonging to the predefined sensitive information category; matching the first item to the entity identifier; based on the first item matching the entity identifier and based on detecting the second item, redacting the second item from the electronic document, producing an edited document comprising the first item; and outputting the edited document comprising the first item (108A).

[0089] In embodiments, the method can comprise receiving a document search request comprising the entity identifier, wherein the electronic document can be obtained from computer-readable storage via a document search based on the entity identifier.

[0090] The method of claim 1 can include receiving, at a proxy server, a download request associated with the entity identifier from a client device; in response to the download request: transmitting, from the proxy server, a proxied download request to an upstream server, in response to the proxied download request, receiving, at the proxy server, the electronic document from the upstream server, detecting the first item and the second item within the electronic document received from the upstream server, matching the first item to the entity identifier, redacting the second item, and transmitting, from the proxy server, the edited document to the client device.

[0091] The entity identifier can include a user identifier associated with the client device.

[0092] The proxy server can be a web proxy server and the download request can be received from a web browser executing on the client device.

[0093] The method can include receiving, at a proxy server, a message including the electronic document from a client device, the message being associated with the entity identifier; in response to the message including the electronic document: detecting the first item and the second item within the electronic document received from the client device, matching the first item to the entity identifier, redacting the second item, and transmitting, from the proxy server, a proxied message including the edited document to an upstream server.

[0094] The entity identifier can include a user identifier associated with the client device.

[0095] The proxy server can be a web proxy server and the message can be received from a web browser executing on the client device.

[0096] The method can include receiving, at the proxy server, a content request including a resource identifier from the client device; in response to the content request: retrieving, at the proxy server, web content associated with the resource identifier, generating modified web content including proxy client code based on the web content, causing the proxy client code to execute on the client device, and transmitting the modified web content to the client device; detecting, in the message including the electronic document, marker data inserted by the client proxy code executing on the client device; in response to detecting the marker data: detecting the first item and the second item within the electronic document received from the client device, matching the first item to the entity identifier, redacting the second item, and transmitting, from the proxy server, the proxied message including the edited document to the upstream server.

[0097] The method can comprise outputting, via a graphical user interface, an indication of the second item, wherein responsive to user input indicating the second item, the second item can be redacted from the electronic document.

[0098] The method can comprise displaying, via the graphical user interface, the electronic document, wherein the indication of the second item can comprise a visual marker marking the second item within the electronic document.

[0099] The method can comprise outputting, in association with the indication of the second item, an indication of the predefined class of sensitive information.

[0100] The entity identifier can be a person identifier or a group of persons identifier, and the predefined class of sensitive information can be a predefined class of personal information.

[0101] A second aspect herein provides a proxy server comprising: at least one memory configured to store computer-readable instructions; at least one processor coupled to the at least one memory and configured to execute the computer-readable instructions, the computer-readable instructions, when executed on the at least one processor, configured to cause the at least one processor to: generate proxied content based on a content request received from a client device; transmit, to an upstream server, a proxied content request; receive, in response to the proxied content request, a first response comprising requested web page content; transmit, to the client device, a second response comprising the requested web page content and executable proxy client code; receive, from the client device, an upload request comprising: a document, and an upload marker generated by the executable proxy client code when executed on the client device; identify the upload marker in the upload request; responsive to identifying the upload marker in the upload request, redact, from the document, an item determined to be a term of a predefined class of sensitive information, resulting in an edited document; generate a proxied upload request comprising the edited document; and transmit, to the upstream server, the proxied upload request.

[0102] A third aspect herein provides a proxy server comprising at least one memory configured to store computer-readable instructions; at least one processor coupled to the at least one memory and configured to execute the computer-readable instructions, the computer-readable instructions, when executed on the at least one processor, configured to cause the at least one processor to: receive, at the proxy server from a client device, a content request comprising a resource identifier; in response to the content request: retrieve web page content associated with the resource identifier, generate modified web page content comprising proxy client code based on the web page content, and transmit the modified web page content to the client device, causing the proxy client code to be executed on the client device; receive, at the proxy server from the client device, an upload request comprising a document; detect, in the upload request, marker data inserted by the client proxy code executed on the client device; in response to detecting the marker data in the upload request, cause items determined to belong to a predefined sensitive information category to be redacted from the document, resulting in an edited document; generate a proxied upload request comprising the edited document; and transmit the proxied upload request to an upstream server.

[0103] In embodiments, the computer-readable instructions can be configured to cause the at least one processor to: determine an entity identifier based on the upload request; and cause the items to be redacted from the document based on the entity identifier.

[0104] In response to determining that the items do not match the entity identifier, the items can be redacted from the document, for example.

[0105] A third aspect herein provides a computer-readable storage medium configured to store computer-readable instructions, the computer-readable instructions, when executed on at least one processor, configured to cause the at least one processor to perform operations comprising: receiving a message from a client device; determining an entity identifier associated with the message; obtaining a document associated with the message; and causing items i) determined to belong to a predefined sensitive information category and ii) determined to not match the entity identifier to be redacted from the document, resulting in an edited document.

[0106] In embodiments, the message can be a download request, and obtaining the document can comprise transmitting a proxied download request to an upstream server and receiving the document from the upstream server in response, in which case the operations can further comprise transmitting a response comprising the edited document to the client device.

[0107] Alternatively, the message can include the document, in which case the operations further include transmitting, to the upstream server, a proxied message including the edited document.

[0108] Alternatively, the document can be retrieved from a document store via a document search performed using the entity identifier.

[0109] The entity identifier can be a user identifier associated with the message or with the client device, and the predefined sensitive information category can be a predefined personal information category.

[0110] Further aspects provide a computer system comprising at least one processor configured to implement any of the above methods or functionalities, and computer readable instructions for programming the computer system to implement the above methods or functionalities.

[0111] It will be appreciated that the above examples have been disclosed by way of example only. Once the disclosure herein has been given, other modifications or uses can become apparent to those skilled in the art. The scope of the disclosure is not limited by the above examples, but only by the claims set forth below.

Claims

1. A computer-implemented method comprising: obtaining an electronic document (102) and an entity identifier associated with the electronic document, the entity identifier pertaining to a predefined class of sensitive information; detecting, within the electronic document (102), a first item (109A) belonging to the predefined class of sensitive information; detecting, within the electronic document, a second item (108B) belonging to the predefined class of sensitive information; matching the first item (108A) to the entity identifier; based on the first item (108A) matching the entity identifier and based on detecting the second item (108B), redacting the second item (108B) from the electronic document (102), producing an edited document (112) comprising the first item (108A); and outputting the edited document (112) comprising the first item (108A).

2. The method of claim 1, comprising: receiving a document search request (231) comprising the entity identifier; wherein the electronic document (202) is obtained from computer-readable storage (234) via document search based on the entity identifier.

3. The method of claim 1, comprising: receiving, at a proxy server (334) from a client device (330), a download request (331) associated with the entity identifier; in response to the download request (331): transmitting, from the proxy server (332) to an upstream server (334), a proxied download request (333), in response to which the electronic document (302) is received at the proxy server (332) from the upstream server (334), detecting the first item and the second item within the electronic document (302) received from the upstream server (334), matching the first item to the entity identifier, redacting the second item, and transmitting, from the proxy server (332) to the client device (330), the edited document (312).

4. The method of claim 3, wherein the entity identifier comprises a user identifier associated with the client device (330).

5. The method of claim 3 or 4, wherein the proxy server (332) is a web proxy server and the download request is received from a web browser executing on the client device.

6. The method of claim 1, comprising: receiving, at a proxy server (432) from a client device (430), a message (431) comprising the electronic document (402), the message (401) being associated with the entity identifier; in response to the message (431) comprising the electronic document (402): detecting the first item and the second item within the electronic document (402) received from the client device (430), matching the first item to the entity identifier, redacting the second item, and transmitting, from the proxy server (432) to an upstream server (434), a proxied message (433) comprising the redacted document (412).

7. The method of claim 6, wherein the entity identifier comprises a user identifier associated with the client device (430).

8. The method of claim 6 or 7, wherein the proxy server (432) is a web proxy server and the message is received from a web browser executing on the client device.

9. The method of claim 8, comprising: receiving, at the proxy server (432) from the client device (430), a content request (500) comprising a resource identifier; in response to the content request (500): retrieving, at the proxy server (432), web content (505) associated with the resource identifier, generating, based on the web content (505), modified web content (505) comprising proxy client code (507), and transmitting the modified web content (505) to the client device (430) for execution of the proxy client code (507) on the client device (430); detecting, in the message (431) comprising the electronic document, tagged data (600) inserted by the client proxy code executing on the client device (430); in response to detecting the tagged data (600): detecting the first item and the second item within the electronic document (402) received from the client device (430), matching the first item to the entity identifier, redacting the second item, and transmitting, from the proxy server (432) to the upstream server, the proxied message comprising the redacted document.

10. The method of claim 1, comprising outputting, via a graphical user interface, an indication of the second item, wherein in response to user input indicating the second item, the second item is redacted from the electronic document (102).

11. The method of claim 10, comprising displaying, via the graphical user interface, the electronic document (102), wherein the indication of the second item comprises a visual marker marking the second item within the electronic document.

12. The method of claim 10 or 11, comprising outputting, in association with the indication of the second item, an indication of the predefined class of sensitive information.

13. The method of any preceding claim, wherein the entity identifier is a person identifier or a group of persons identifier, wherein the predefined class of sensitive information is a predefined class of personal information.

14. A proxy server (432), comprising: at least one memory configured to store computer-readable instructions; at least one processor coupled to the at least one memory and configured to execute the computer-readable instructions, the computer-readable instructions, when executed on the at least one processor, being configured to cause the at least one processor to: receiving, at the proxy server (432) from a client device (430), a content request (500) including a resource identifier; responsive to the content request (500): retrieving web page content (505) associated with the resource identifier, generating modified web page content (505) including proxy client code (507) based on the web page content (505), and transmitting the modified web page content (505) to the client device (430) for execution of the proxy client code (507) on the client device (430); receiving, at the proxy server (432) from the client device (430), an upload request including a document; detecting, in the upload request, marker data (430) inserted by the client proxy code executing on the client device; responsive to detecting the marker data in the upload request, redacting from the document an item determined to belong to a predefined sensitive information category, resulting in an edited document; generating a proxied upload request including the edited document; and transmitting the proxied upload request to an upstream server.

15. The proxy server (432) of claim 14, wherein the computer-readable instructions are configured to cause the at least one processor to: determine an entity identifier based on the upload request; and cause the item to be redacted from the document based on the entity identifier.

16. The proxy server (432) of claim 15, wherein the item is redacted from the document responsive to determining that the item does not match the entity identifier.

17. A computer-readable storage medium configured to store computer-readable instructions that, when executed on at least one processor, are configured to cause the at least one processor to perform operations comprising: receiving a message (231, 331, 431) from a client device (230, 330, 430); determining an entity identifier associated with the message (231, 331, 431); obtaining a document (202, 302, 402) associated with the message (231, 331, 431); and causing an item i) determined to belong to a predefined sensitive information category and ii) determined to not match the entity identifier to be redacted from the document (202, 302, 402), resulting in an edited document (212, 312, 412).

18. The computer-readable storage medium of claim 17, wherein: the message is a download request (331), and obtaining the document comprises transmitting a proxied download request (333) to an upstream server (334) and receiving the document (302) from the upstream server (334) in response, the operations further comprising transmitting a response including the edited document to the client device; or ​ ​ The message (431) includes the document (402), and the operations further include transmitting, to an upstream server (434), a proxied message (433) that includes the edited document (412).

19. The computer-readable storage medium of claim 17, wherein the document (202) is obtained from a document store (234) via a document search performed using the entity identifier.

20. The computer-readable storage medium of any of claims 17 to 19, wherein the entity identifier is a user identifier associated with the message or with the client device, and wherein the predefined sensitive information category is a predefined personal information category.