Using machine learning to enhance data security and access control
By using machine learning models to evaluate data requests and elements and dynamically control data access, it addresses the flexibility and privacy issues in existing systems and enables secure and beneficial data sharing.
Patent Information
- Application Number
- CN202080082630.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-03
- Filing Date
- 2020-11-24
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2040-11-24
AI Technical Summary
Existing data access systems lack flexibility and predictability, making it difficult to provide beneficial data access without violating privacy, resulting in information silos and lost innovation opportunities.
Use machine learning models to automatically control data access. By training multiple sets of models to evaluate data requests, individual data elements, and aggregated data elements, dynamic decisions are made based on defined access rules and customized reports are generated.
It enables flexible and secure data access, ensuring that data sharing is beneficial to human life without compromising privacy, while improving the efficiency and accuracy of data access.
Smart Images

Figure CN114787809B_ABST
Abstract
Description
[0001] introduction
[0002] Aspects of the present disclosure relate to data access and security, and more particularly, to using machine learning to drive data visibility, control, access, and security decisions. Background Art
[0003] A wide variety of global systems are used to collect and store data about any number of data subjects, such as patients, users, or any other individual or entity described by the data being processed and stored. For example, healthcare data is typically maintained for each patient at a given institution. This healthcare data may include any number and variety of data elements, such as diagnoses, genetic information, clinical records, medications the patient is currently taking (or previously prescribed), surgical inpatient or outpatient procedures, or other surgical procedures that have been performed or recommended, etc. Typically, this data is subject to various protections and requirements related to data security and user privacy. However, in many cases, access to the data can be beneficial to the data subject or others without causing any harm to the data subject or infringing their privacy.
[0004] Existing systems often make data access difficult and create significant confusion about which data elements are (or can be) exposed and which are protected. In many fields, such as healthcare, data access and security are primarily controlled by various technologies that provide inflexibility and little predictability regarding which data elements can be shared, even when sharing could benefit others. For example, in some cases, a patient may tragically pass away due to complications or underlying medical conditions. During the patient's lifetime, healthcare data, such as genetic information and diagnoses, may have been collected and correlated. For example, this information could be beneficial to the patient's living relatives, such as determining whether they also have genetic markers that may predispose them to the same condition. If relatives had access to this information, it could help them address the condition earlier and improve their quality of life. However, current systems may be so restrictive that relatives cannot access this information because they lack the patient's consent, which, unfortunately, is often unavailable at this stage.
[0005] In other industries, partners and competitors are similarly concerned about losing proprietary knowledge and often block data access entirely, even if this means missing out on opportunities for innovation and growth. At a high level, this can negatively impact society's progress in innovation, as valuable data that others could use to advance innovation is often unavailable due to access restrictions that fail to consider the impact of such restrictions. Therefore, a flexible and intelligent system is needed to more finely control data access, as an alternative to the existing binary framework of all-or-nothing data disclosure. Summary of the Invention
[0006] Certain embodiments provide a method for automatically controlling data access using one or more machine learning models. The method generally includes: receiving a first request from a first user for data associated with a second user; automatically determining whether the first request satisfies one or more data access rules by processing the first request using a first set of one or more trained machine learning models; after determining that the first request satisfies the one or more data access rules, automatically retrieving a first plurality of data elements based on the first request; automatically determining whether each of the first plurality of data elements satisfies the one or more data access rules by processing each of the first plurality of data elements using a second set of one or more trained machine learning models; after determining that a first group of data elements from the first plurality of data elements satisfies the one or more data access rules, determining whether the first group of data elements satisfies the one or more data access rules by processing the first group of data elements using a third set of one or more trained machine learning models; and after determining that the first group of data elements satisfies the one or more data access rules, generating a customized report including the first group of data elements.
[0007] Certain embodiments provide a method for training one or more machine learning models to control data accessibility. The method generally includes: generating a first training data set from a set of historical access records, wherein each corresponding access record in the first training data set corresponds to a corresponding data request and includes information identifying whether the corresponding request satisfies one or more data access rules; generating a second training data set from a set of data records, wherein each corresponding data record in the second training data set corresponds to a corresponding data element and includes information identifying whether the corresponding data element satisfies the one or more data access rules; generating a third training data set from the set of historical access records, wherein each corresponding access record in the third training data set corresponds to a corresponding set of aggregated data elements and includes information identifying whether the corresponding set of aggregated data elements satisfies the one or more data access rules; training the one or more machine learning models based on the first training data set, the second training data set, and the third training data set to generate an output identifying whether the data request should be granted; and deploying the one or more machine learning models to one or more computing systems.
[0008] Aspects of the present disclosure provide apparatuses, devices, processors, and computer-readable media for performing the methods described herein.
[0009] To accomplish the foregoing and related ends, the one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and accompanying drawings set forth in detail certain illustrative features of the one or more aspects. However, these features are indicative of only a few of the various ways in which the principles of the various aspects may be employed. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings depict certain aspects of one or more embodiments and, therefore, should not be considered limiting the scope of the disclosure.
[0011] Figure 1 Depicted is an environment including an analytics server configured to control data access using machine learning, according to some embodiments disclosed herein.
[0012] Figure 2 Depicted is a workflow for controlling data access using various access rules according to some embodiments disclosed herein.
[0013] Figure 3 is a flowchart depicting a method for training a machine learning model to control data access based on characteristics of a data request according to some embodiments disclosed herein.
[0014] Figure 4 is a flowchart depicting a method for training a machine learning model to control data access based on characteristics of individual data elements according to some embodiments disclosed herein.
[0015] Figure 5 is a flowchart depicting a method for training a machine learning model to control data access based on characteristics of aggregated data elements that individually satisfy access rules according to some embodiments disclosed herein.
[0016] Figure 6 is a flowchart depicting a method for controlling data access using a trained machine learning model according to some embodiments disclosed herein.
[0017] Figure 7 Depicted is a graphical user interface (GUI) for enhancing data access control and notification according to some embodiments disclosed herein.
[0018] Figure 8 is a flowchart depicting a method for automatically controlling data access using one or more machine learning models according to some embodiments disclosed herein.
[0019] Figure 9 is a flowchart depicting a method for training one or more machine learning models to control data accessibility according to some embodiments disclosed herein.
[0020] Figure 10 is a block diagram depicting a computing device configured to train and use a machine learning model to control data access according to some embodiments disclosed herein.
[0021] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the figures. It is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further recitation. DETAILED DESCRIPTION
[0022] Embodiments of the present disclosure provide techniques for performing effective data access control that ensure that data privacy and security are maintained while also enabling flexible access in situations where data access is beneficial without causing corresponding harm. Advantageously, such a system can automatically provide access to valuable data in a defined manner so that data that can benefit from sharing is shared, while data sharing that may cause harm (e.g., by identifying specific individuals and thereby causing privacy issues) is not performed. Such fine-tuned data sharing is simply not possible with existing methods that utilize an all-or-nothing approach to data sharing. For example, current mechanisms for manually determining what data to share, based on the large amounts of data collected by existing systems, make it virtually impossible to provide the level of flexibility in data sharing discussed herein. Consequently, such manual mechanisms may be overly cautious and overly restrict access to data by not sharing most data.
[0023] In some embodiments, to enable this flexible data sharing, data requests and data elements are evaluated at multiple layers or steps using defined access rules before any data is provided (or refrained from providing) (e.g., evaluating the request individually, evaluating each individual data element, and evaluating the aggregated data element). In some embodiments, a set of one or more machine learning models are trained to classify requests and data elements based on such access rules. In doing so, embodiments of the present disclosure allow for rapid evaluation and response to data requests while maintaining data security. Additionally, because the system utilizes objective models to evaluate requests (whereas current systems rely on subjective decisions), data integrity is ensured.
[0024] In some embodiments, a set of access rules are utilized to drive data access decisions. In some embodiments, there can be any number of access rules in a given deployment. In some embodiments, the system can utilize a basic set of access rules regardless of industry, and allow additional rules to be added or modified based on the specific requirements and expectations of a given industry or deployment. In some embodiments, access rules are used to control data access during an initial manual phase, and decisions made based on the rules (e.g., by subject matter experts or other users) are used to train machine learning models to automatically provide analysis. That is, in the process referred to above as the initial manual phase, human reviewers can evaluate requests and data elements based on the access rules to determine whether the request should be granted and / or whether the data should be shared. These manual decisions can be stored in a record that includes the request and / or data details and marked as manual decisions. Machine learning algorithms can use such records to train machine learning models to automatically perform similar analyses.
[0025] In some embodiments, access rules (and therefore models trained based on access rules) are used to ensure that data can be shared when access is deemed beneficial while protecting privacy and security. In some embodiments, the system uses a set of three rules: a first access rule states that data (if provided) can only be used to improve human life or society (without harming the data subject), a second access rule states that the requesting entity must have a legitimate intent with the data, and a third access rule states that the data must continue to be protected and kept secure to the greatest extent possible without conflicting with other rules. Based on this framework, models can be trained to effectively provide a dynamic data access system that allows and restricts access to data in an intelligent and flexible manner while complying with the access rules.
[0026] Figure 1 An environment 100 is depicted that includes an analysis server 110 configured to use machine learning to control data access according to some embodiments disclosed herein. In the illustrated embodiment, a requesting user 105 can provide a request to the analysis server 110. The request typically identifies at least the requested data and the intended use of the data. In some embodiments, the request includes metadata or other additional data that identifies the requesting user 105 or provides other additional information to give context to the request. In some embodiments, the request is associated with a requester profile that specifies, for example, the identity of the requesting user 105 (e.g., the requesting user's name or other identifying information), the reason or purpose of the request, a timeline for when the data is needed, or any other documentation that supplements or provides context to the request.
[0027] For example, suppose a user wishes to determine whether their risk of suffering from a particular condition (such as deep vein thrombosis (DVT)) is increased based on their family history. In some embodiments, a user (acting as requesting user 105) can provide a request including this information to analysis server 110 directly or via another device (such as through a network). Analysis server 110 can be any suitable server in any suitable environment (e.g., an internal environment, an environment associated with an entity, an environment in the cloud, etc.). In some embodiments, the request specifically identifies one or more data subjects. For example, requesting user 105 can identify its family members as data subjects (e.g., "Does anyone in my family have a history of DVT?"). In another embodiment, analysis server 110 evaluates the request to identify the relevant data subject. For example, based on the request (e.g., "Do I have a high genetic risk of suffering from DVT?"), the system can determine that the relevant data subject is a relative of the requesting user 105. This can, for example, be accomplished using natural language processing (NLP).
[0028] Additionally, in some embodiments, the request specifically identifies the desired data elements (eg, identifying a particular report, test, or other data element). In another embodiment, the analysis server 110 identifies relevant data elements based on analyzing the request using NLP or other techniques.
[0029] In the illustrated environment 100, the analysis server 110 includes a data sharing component 115 and a custom report generator 120. The data sharing component 115 typically evaluates requests to determine whether the request should be granted (in whole or in part) based on a set of access rules discussed herein, and additionally identifies, retrieves, and evaluates relevant data elements to determine whether these relevant data elements should be provided to the data requester based on the same set of access rules discussed herein. In some embodiments, the data sharing component 115 evaluates requests and data at three levels: a first level evaluating the request against the access rules, a second level evaluating each individual data element against the access rules, and a third level evaluating the aggregated set of data elements against the access rules. Only data elements that pass all levels are included in the final report. The custom report generator 120 typically constructs a custom report for the requesting user 105. The custom report can include any data elements approved for sharing by the data sharing component. In some embodiments, the custom report can further include reasons why any portion of the request (or the entire request) was rejected or why any data was excluded, as discussed in more detail below. Although the data sharing component 115 and the custom report generator 120 are illustrated as discrete components for conceptual clarity, in embodiments, operations may be combined or distributed across any number of components and devices.
[0030] In some embodiments, the data sharing component 115 may first evaluate the request to determine whether the request should be directly denied, such as based on defined access rules. In some embodiments described herein, this may be referred to as the "first layer." For example, the data sharing component 115 may determine whether allowing access to the requested data will improve human life without harming the data subject, whether the request is for legitimate purposes, whether the data will continue to be protected, etc., as specified in the access rules. In some embodiments, the data sharing component 115 uses one or more trained models to achieve this, as discussed in more detail below. For example, the data sharing component 115 may extract features of the request and process these features using one or more models trained based on access data labeled from previous requests, as discussed further herein. These features may include, but are not limited to, the identity of the requesting user 105 and / or (multiple) data subjects, the relationship between the requesting user and the data subject, the indicated purpose of the requested data (which may be explicitly stated or may be determined based on processing the request using, for example, NLP), etc.
[0031] If the data sharing component 115 determines that the request should be denied (e.g., because the requesting user 105 intends to use the data commercially for advertising purposes, which is not a legitimate use to improve human life), the data sharing component 115 can deny the request. The custom report generator 120 can then generate a report indicating that the request was denied and including the reason for the determination (e.g., indicating which access rule was failed).
[0032] In some embodiments, if the data sharing component 115 determines (e.g., using a trained model) that the request satisfies the access rules, the data sharing component 115 may begin a second layer of analysis by identifying relevant data elements and retrieving these relevant data elements from one or more data repositories 125. For example, the data sharing component 115 identifies data repositories 125 that may contain the data. For example, the data repositories 125 may be identified based on the identity of the requesting user 105, the identity of the data subject(s), the nature or context of the request (e.g., the specific type of data being requested), etc. The data sharing component 115 may then transmit a query to each identified repository to retrieve relevant data.
[0033] In some embodiments, the data sharing component 115 then evaluates each individual data element based on the access rules (e.g., using one or more trained models). In some embodiments, the data sharing component 115 can utilize the same model used to evaluate the request, or can use a different set (or multiple) of models trained to evaluate data elements. For each individual data element, if the data sharing component 115 determines that the data element meets the access rules, the data sharing component 115 can add it to the set of data elements that will potentially be granted access to the requesting user 105. For any data that does not pass, the data sharing component 115 can avoid disclosing the data. In some embodiments, the custom report generator 120 will include the reasons why a particular data element was excluded.
[0034] If multiple data elements are determined to meet the criteria, in some embodiments, the data sharing component 115 can then evaluate the multiple data elements as a whole to determine whether the data sharing component as a whole should be shared. For example, a group of data elements may meet the rules individually (e.g., because the group of data elements is being used to improve humanity without identifying or harming the data subject), but when evaluated collectively, the data elements may not meet the rules (e.g., because the data elements may collectively be used to identify and / or harm the data subject). For example, each of gender, date of birth, and place of work may not be sufficient to identify the data subject individually because there are many data subjects who individually match such definitions, but in the aggregate, such information may only be relevant to a small group or even a single data subject.
[0035] In some embodiments, based on the evaluation by the data sharing component 115, the custom report generator 120 then generates a custom report and returns it to the requesting user 105. In some embodiments, the custom report generator 120 may additionally provide a notification to the data subject(s) 130 indicating what data was shared. In some embodiments, the notification further indicates the reason or purpose of the request, the identity of the requesting user 105, etc. In certain embodiments, the notification may additionally indicate any data elements that were retained.
[0036] In the illustrated embodiment, data subject 130 can provide feedback related to the data access to analysis server 110. For example, data subject 130 can indicate that one or more specific data elements are not desired to be shared, or indicate that one or more retained data elements should still be shared. In some embodiments, the system can improve the trained model(s) based on the feedback.
[0037] In the illustrated embodiment, data sharing component 115 utilizes a trained model provided by training server 135. Although training server 135 and analysis server 110 are depicted as separate servers for conceptual clarity, in some embodiments, the training server and analysis server can operate as a single server. That is, the model can be trained and used by a single server, or it can be trained by one or more servers and deployed on one or more other servers for use.
[0038] As illustrated, training server 135 includes training data generator 140 and model trainer 145. Although training data generator 140 and model trainer 145 are depicted as separate components for conceptual clarity, in some embodiments, the operations of the training data generator and model trainer may be combined or distributed across any number of components and devices.
[0039] The training data generator 140 typically uses the historical access records 150 to generate a training data set that the model trainer 145 uses to train one or more machine learning models. In some embodiments, the historical access records 150 relate to previous decisions regarding data sharing. For example, each record in the historical access records 150 may correspond to a specific request, and the record may indicate whether the request was granted. In some embodiments, for each denied request, the corresponding record may also indicate the reason why the request was denied. In some embodiments, for each approved request, the corresponding record may indicate the related data elements, whether each individual data element was approved for release, whether the aggregated data set was approved, etc.
[0040] In some embodiments, the training data generator 140 generates a separate training data set for each model to be trained. For example, the model trainer 145 can train a separate model for each layer of analysis: a first set of one or more models for evaluating whether a request satisfies an access rule, a second set of one or more models for evaluating whether each individual data element satisfies an access rule, and a third set of one or more models for evaluating whether an aggregated data element satisfies an access rule. Similarly, for each layer, the model trainer 145 can train a separate model for each access rule. For example, the model trainer 145 can train a first model to determine whether a request satisfies a first rule (e.g., whether a request will benefit a human), train a second model to determine whether a request satisfies a second access rule, and train a third model to determine whether a request satisfies the second access rule. Similarly, the model trainer 145 can train a first model to determine whether an individual data element satisfies the same first rule, train a second model to determine whether an individual data element satisfies a second access rule, and train a third model to determine whether an individual data element satisfies a third access rule. Additionally, model trainer 145 may train a third model to determine whether the aggregated data element satisfies the first rule, train a second model to determine whether the aggregated data element satisfies the second access rule, and train a third model to determine whether the aggregated data element satisfies the third access rule.
[0041] In some embodiments, the generated training dataset can differ based on the target model. For example, to train a model(s) for analyzing the request layer, the training data generator 140 can generate a dataset from historical access records 150, where each training record specifies input features representing various aspects of the request (e.g., the reason for the determination, the identity of the requester, etc.) and corresponding labels indicating whether the request satisfied the access rules and was therefore approved (or whether each individual data access rule was determined to have passed or failed). For example, a human user can evaluate a request to determine whether it satisfies the access rules. The request data (or metadata) can then be recorded along with the user's decision, as an example of a label. For the individual data element layer, the training data generator 140 can generate a set of data records, each of which corresponds to a specific data element that was previously requested and / or shared, and each record specifies input features related to the characteristics of the data element (e.g., the domain to which the data element relates, predefined privacy levels, relevant regulations, etc.) and a label indicating whether access to the data element was determined to satisfy the access rules.
[0042] In the illustrated embodiment, the model trainer 145 uses the generated training data set to train a set (of models). Typically, training each model includes, for one or more training records, providing the indicated (of multiple) input features as input to the model (which may be started with random parameters). The generated output is then compared with the label of the training record, and the model trainer 145 can calculate a loss based on the difference between the generated output and the provided label. This loss can then be used to modify the internal parameters or weights of the model (e.g., via backpropagation). By iteratively processing each training record, the model is iteratively improved to generate accurate access decisions based on the input features.
[0043] As shown, the training server 135 deploys these trained model(s) to the analysis server 110 for use during runtime. In some embodiments, the training server 135 also receives update(s) from the analysis server 110 (e.g., in the form of feedback from data subjects or subject matter experts). These updates can be used to further refine the model(s).
[0044] Figure 2 A workflow 200 for controlling access to data using various access rules according to some embodiments disclosed herein is depicted. The workflow 200 begins when a request 205 is received. The request typically indicates or identifies the data desired. The indication can be of any specificity, including specifying a particular data element (e.g., identifying a particular record), identifying a type of data (e.g., "test results"), indicating what content is desired (e.g., "records related to DVT"), etc. In some embodiments, the request 205 also identifies one or more data subjects. As discussed, a data subject is a person to whom the requested data relates. In some embodiments, identifying the data subject(s) can also be of any level of specificity, including identifying a particular person or indicating a group of people (e.g., "my relatives," "males under sixty years of age," etc.).
[0045] In some embodiments, request 205 further identifies the requesting entity, the purpose of the data, and the like. In some embodiments, indicating the purpose of the data may include explicitly stating a reason, selecting a predefined purpose, and the like. In certain embodiments, request 205 includes natural language text. In one such embodiment, the system may use NLP to identify the requested data, relevant data subjects, and / or the purpose of the request. For example, the analysis server may use NLP to extract concepts from the text and, based on the concepts identified in request 205, determine relevant industries or fields (e.g., healthcare), desired data elements, relevant data subjects, and the like.
[0046] At block 210, the analysis server determines whether the request 205 satisfies one or more defined access rules. In some embodiments, without using a machine learning model, the determination includes comparing the concept identified (or specified) in the request 205 to one or more defined lookup tables that specify acceptable or legitimate purposes for the data. In some embodiments, these lookup tables are industry-specific, such that a given purpose may be acceptable for some industries but unacceptable for other industries.
[0047] In some embodiments, at block 210, the analysis server instead utilizes a trained machine learning model to determine whether the request 205 satisfies the access rules. In some embodiments, as discussed above, one or more machine learning models are trained based on manually curated access records that include labeled training data. Further, as discussed, the labeled training data can be used to train one or more machine learning models to automatically perform analysis of whether the request 205 satisfies the access rules. In some embodiments, the access rules relate to defined data handling ethics, as discussed above.
[0048] In some embodiments, a separate model is trained for each access rule. In some such embodiments, if the analysis server utilizes three access rules, box 210 will include passing the request 205 through three separate machine learning models. The input to the model (and therefore, the input used to train the model) typically includes characteristics of the request 205, such as the indicated purpose of the data or the field or industry to which the data relates. For example, a request indicating that the data will be anonymized and used to promote medical research may be approved, while a request indicating that the data will be sold to an advertising agency may be denied. In general, the request characteristics that are evaluated can include any number and variety of concepts extracted from the request 205. As discussed above, in some embodiments, the characteristics can include the identity of the requesting user and / or (multiple) data subjects, the relationship between the requesting user and the data subjects, the indicated purpose of the requested data (which can be explicitly stated or can be determined based on processing the request using, for example, NLP), the field or industry to which the request or data belongs, etc.
[0049] As shown, if request 205 fails any access rule (e.g., as indicated by a trained model that rejected request 205), workflow 200 continues to block 250, where the analysis server generates a customized report including one or more reasons for the rejection. In some embodiments, the analysis server can indicate which access rule(s) request 205 failed to pass. The analysis server can identify the failed rule(s) based on the model(s) that rejected the request. For any trained model that rejected the request, the analysis server can indicate that the corresponding access rule(s) were not satisfied.
[0050] If the request 205 satisfies all access rules (e.g., if the request is approved by all machine learning models at this stage), the analysis server generates a data query 215, which is transmitted to one or more data repositories 125. In some embodiments, the analysis server generates the data query 215 based on the requested data indicated by the request 205. For example, if the request inquires about the requester's familial risk of developing DVT, the analysis server can generate a query to retrieve data records corresponding to a data subject associated with the requester and related to DVT (e.g., diagnosis, test results, genetic markers, name and / or location of medical provider who performed the test(s), accuracy of the test, surgical procedures recommended or completed, etc.).
[0051] In the illustrated embodiment, the data repository 125 stores data created by one or more data generators 220. Data generators 220 generally include any data source, such as a medical or non-medical institution, a specific machine or device, the data subject itself, another person facilitating data collection, etc. For example, in the case of medical data, the data generators 220 may include patients, medical professionals, devices used to retrieve or record data from patients, clinics or institutions that collect data, etc.
[0052] As shown, the data repository 125 returns relevant data elements 225 based on the data query 215. As used herein, a data element 225 is generally a discrete piece of data and can include any number and type of values. For example, a data element 225 can specify a medical test result, indicate the test(s) performed, specify the accuracy of the test(s), indicate the institution that performed the test(s), etc.
[0053] At block 230, each of these data elements 225 is individually evaluated to determine whether it satisfies the access rules. In some embodiments, the analysis server utilizes one or more trained machine learning models to perform this review. In some embodiments, the analysis server utilizes models trained specifically for evaluating individual data elements. That is, while the model(s) used in block 210 are trained to evaluate request features, the model(s) used in block 230 can be trained to evaluate data elements. In some embodiments, the models used in block 230 are trained similarly to the models discussed above with reference to block 210. For example, the system can retrieve historical access records or data records that indicate, for each data element in the data record, whether a human user approved access (or whether the user determined that the data element passed a particular access rule).
[0054] Each such record can specify one or more characteristics of the data as input features. In an embodiment, these features may include the type of data, the specificity of the data, the source of the data, the field to which the data relates, whether the data specifically identifies a data subject, etc. In some embodiments, one or more of these features are specified in the metadata associated with the data element 225. In some embodiments, if the data element 225 includes natural language text (e.g., clinical notes), the analysis server can utilize natural language processing to extract concepts to be used as input features. Additionally, in some embodiments, each record is tagged with an indication of whether a human user has determined that the data element passes (or fails) one or more data access rules. In some embodiments, the model is then trained based on the records.
[0055] If a given data element fails any access rule, it is excluded from the report at block 250. In some embodiments, the analysis server may also include one or more rejection reasons (e.g., indicating that one or more data elements are retained because they fail a particular rule). In some embodiments, the analysis server may additionally generate and transmit to the data subject a notification indicating which data element(s) were retained and which were released.
[0056] In the illustrated workflow 200, any data elements 225 that are determined to satisfy the access rule(s) are combined to form a set of aggregated data elements ("aggregated data") 235. As illustrated, the aggregated data 235 is then evaluated at block 240 to determine whether the aggregated data 235 satisfies the data access rule(s). For example, two or more data elements 225 may pass an access rule when evaluated individually at step 230 but fail when aggregated because, when combined, the elements allow identification of the underlying data subject.
[0057] In some embodiments, the evaluation at block 240 is similarly performed using one or more trained machine learning models. In some embodiments, the analysis server utilizes models trained specifically for the aggregate evaluation. That is, the analysis server may utilize a first set of one or more models to perform the request evaluation at block 210, a second set of one or more models to perform the data evaluation at block 230, and a third set of one or more models to perform the aggregate data evaluation at block 240. In some embodiments, the features evaluated at block 240 may mirror the features utilized in block 230.
[0058] As shown, if the aggregated data passes the access rules, the analysis server generates a custom report using the approved data elements (in block 245). In the illustrated embodiment, if the aggregated set fails one or more access rules, the analysis server generates the custom report while excluding at least some of the data elements. In some embodiments, the analysis server may avoid providing any data elements. In another embodiment, the analysis server may provide a subset of the approved data elements.
[0059] For example, upon determining that aggregated data 235 fails one or more access rules, the analysis server may remove one or more data elements from the collection and re-evaluate the aggregated data set using the model. In some embodiments, the analysis server may iteratively evaluate different combinations of data elements to identify which data elements should be removed from the collection. For example, the analysis server may attempt to find a combination of data elements that pass the rules with the fewest elements removed (so that the analysis server can return as much data as possible).
[0060] Figures 3 to 5 Techniques for training the machine learning models discussed herein are described in further detail, such as techniques for evaluating data requests, data elements, and aggregated data.
[0061] Figure 3 is a flowchart depicting a method 300 for training a machine learning model to control data access based on characteristics of a data request according to some embodiments disclosed herein. In some embodiments, the method 300 is used to train a model to evaluate a request (e.g., Figure 2 Method 300 begins at block 305, where a training server (e.g., training server 135) retrieves a set of historical access records. In some embodiments, each historical access record corresponds to a previous data request and includes a tag indicating whether a human reviewer approved the request (and / or whether the request was determined to satisfy one or more data access rules). For example, during an initial manual / training phase, the training server may collect data as reviewers evaluate and approve or deny data requests. Based on this monitoring, the training server may build a training dataset of historical access records.
[0062] At box 310, the training server selects one of the historical access records. In an embodiment, when the training server traverses each historical access record in the training set, the selection can utilize any suitable criteria (e.g., starting with the oldest record, starting with the most recent record, etc.). Method 300 then continues to box 315, where the training server extracts one or more features of the request corresponding to the selected record. These features will be used as input features for (multiple) machine learning models. This can include extracting concepts from the request, such as the purpose of the request, the field or industry to which the request relates, etc. For example, the training server can determine whether the request is related to health or happiness, economic benefits, etc. In some embodiments, the training server extracts these features from the request using NLP. In some embodiments, the request may have been pre-evaluated to extract features, and these features may be stored in the access record. In some embodiments, each access record is further associated with a label indicating whether the request is approved or rejected.
[0063] Method 300 then continues to block 320, where the training server trains one or more machine learning models based on the selected records. In some embodiments, the training server does this by providing the features (extracted in block 315) as input to the model. The model can be a new model initialized with random weights and parameters, or it can be a partially or fully pre-trained model (e.g., based on a previous training round). Based on the input features, the model in training generates some output (e.g., a classification as "pass" or "fail") for one or more access rules. In an embodiment, the training server can compare this generated classification with the actual label of the record (as indicated in the record) to calculate a loss based on the difference between the actual result and the generated result. This loss is then used to improve one or more internal weights and parameters of the model (e.g., via backpropagation) so that the model learns to classify requests more accurately.
[0064] In some embodiments, the training server trains the model to analyze requests according to a set of access rules. That is, the training server can train the model to evaluate requests according to all access rules simultaneously and output a binary "pass" or "fail" (or a set of determinations, one for each rule) based on whether the request passes all access rules or fails at least one access rule. In other embodiments, as discussed above, the training server trains a separate model for each access rule.
[0065] Method 300 then proceeds to block 325, where the training server determines whether additional training is required. This may include evaluating any termination criteria, such as whether any additional historical access records remain in the training dataset. In various embodiments, other termination criteria may include, but are not limited to, whether a predefined amount of time or computing resources has been expended to train the model, whether the model has achieved a predefined minimum accuracy, etc. If additional training remains to be completed, method 300 returns to block 310.
[0066] If not, the method 300 continues to block 330, where the training server deploys the trained model(s) to analyze incoming data requests during runtime. In some embodiments, this includes transmitting some indication of the trained model(s) (e.g., weight vectors) that can be used to instantiate the model(s) on another device. For example, the training server can transmit the weights of the trained model(s) to the analysis server. These models can then be used to evaluate newly received data requests.
[0067] Figure 4 is a flowchart depicting a method 400 for training a machine learning model to control data access based on characteristics of individual data elements according to some embodiments disclosed herein. In some embodiments, the method 400 can be used to train a model to evaluate individual data elements (e.g., Figure 2 Method 400 begins at block 405, where the training server retrieves one or more historical access records, each corresponding to a previous data request. In some embodiments, the training server selects access records for which requests were approved. That is, because data is not retrieved or analyzed for rejected requests, the training server can only retrieve approved requests for which at least one data element was retrieved and evaluated by a human reviewer. In some embodiments, each historical access record is associated with one or more data records, each corresponding to a respective data element retrieved based on the request.
[0068] Method 400 then proceeds to block 410, where the training server selects a historical access record from a set of training access records. In embodiments, as the training server iterates through each historical access record in the training set, the selection may utilize any suitable criteria (e.g., starting with the oldest record, starting with the most recent record, etc.). At block 415, the training server identifies data record(s) associated with the selected access record. In some embodiments, each data record corresponds to a data element retrieved in response to a request corresponding to the selected access record. For example, assume that the request corresponding to the selected access record results in ten data elements being retrieved from the data repository. In some embodiments, the access record will therefore include, or be linked to or otherwise associated with, the ten data records (one for each data element). In some embodiments, each data record includes characteristics of the corresponding data element and a tag indicating whether the data element satisfies one or more access rules.
[0069] At block 420, the training server selects one of the identified data records. Method 400 then proceeds to block 425, where the training server extracts one or more features of the data element corresponding to the selected record. These features typically correspond to characteristics of the data element, such as the data type, data source, predefined sensitivity or privacy level of the data, and the like. In some embodiments, the features include a data profile for the data element, where a data profile is a metadata structure that specifies the relevant features. In certain embodiments, the training server also extracts one or more data source profiles for the data element. A data source profile is typically a metadata structure that specifies characteristics of the source of the data element. For example, if the data element was collected by a specific medical institution, the data source profile may specify characteristics of the institution (such as its name, location, and the like). Similarly, if the data element was collected using a specific device, the profile may specify the identity and type of the device, maintenance records, accuracy of the device, and the like. In some embodiments, each data record may be associated with any number of profiles corresponding to the entities involved in collecting the data and forwarding it to the data repository.
[0070] Method 400 then continues to block 430, where the training server trains one or more machine learning models based on the selected data records. In some embodiments, the training server does this by providing the features (extracted in block 425) as input to the model. The model can be a new model initialized with random weights and parameters, or it can be a partially or fully pre-trained model (e.g., based on a previous training round). Based on the input features, the model in training generates some output (e.g., a classification of "pass" or "fail") for one or more access rules. In an embodiment, the training server can compare this generated classification with the actual label (included by the data record) to calculate a loss based on the difference between the actual result and the generated result. This loss is then used to improve one or more internal weights and parameters of the model (e.g., via backpropagation) so that the model learns to more accurately classify individual data elements.
[0071] In some embodiments, the training server trains the model to analyze data elements according to a set of access rules. That is, the training server can train the model to evaluate data elements according to all access rules simultaneously and output a binary "pass" or "fail" (or a set of determinations, one for each rule) based on whether the data element passes all access rules or fails at least one access rule. In other embodiments, as discussed above, the training server trains a separate model for each access rule.
[0072] Method 400 then proceeds to block 435, where the training server determines whether the selected access records include at least one additional data record that has not yet been evaluated. If so, method 400 returns to block 420. If not, method 400 continues to block 440, where the training server determines whether additional training is required. This may include evaluating any termination criteria, such as whether any additional historical access records remain in the training dataset. In various embodiments, other termination criteria may include, but are not limited to, whether a predefined amount of time or computing resources has been expended to train the model, whether the model has achieved a predefined minimum accuracy, and the like. If additional training remains to be completed, method 400 returns to block 410.
[0073] If not, the method 400 continues to block 445, where the training server deploys the trained model(s) to analyze the individual data elements retrieved during runtime. In some embodiments, this includes transmitting some indication of the trained model(s) (e.g., weight vectors) that can be used to instantiate the model(s) on another device. For example, the training server can transmit the weights of the trained model(s) to the analysis server. The model(s) can then be used to evaluate the data elements retrieved in response to newly received data requests.
[0074] Figure 5 is a flow chart depicting a method 500 for training a machine learning model to control data access based on characteristics of aggregated data elements that individually satisfy access rules, according to some embodiments disclosed herein. In some embodiments, the method 500 can be used to train a model to evaluate aggregated data elements corresponding to aggregated data (e.g., Figure 2 240). Method 500 begins at block 505, where the training server retrieves one or more historical access records, each corresponding to a previous data request. In some embodiments, the training server selects access records for which requests were approved. That is, because data is not retrieved or analyzed for rejected requests, the training server can only retrieve approved requests for which at least one data element was retrieved and evaluated by a human reviewer. In some embodiments, the training server only retrieves records for which at least two data elements were retrieved (e.g., so that aggregated data may result in different results than individual evaluations). In some embodiments, each historical access record is associated with one or more data records, each corresponding to a respective data element retrieved based on the request.
[0075] Method 500 then proceeds to block 510, where the training server selects a historical access record from a set of training records. In some embodiments, as the training server iterates through each historical access record in the training set, the selection can utilize any suitable criteria (e.g., starting with the oldest record, starting with the most recent record, etc.). At block 515, the training server identifies the data record(s) associated with the selected access record that was determined to satisfy the access rule. In other words, the training server can identify which data elements, if any, are considered to satisfy the access rule individually. For example, assume that the system retrieves ten data elements based on the request, and three of the data elements fail one or more data access rules when evaluated individually. In some embodiments, the training server can identify a subset of the data elements that pass the individual review (e.g., the remaining seven).
[0076] At block 520, the training server selects one of the identified data records that passed the individual review. Method 500 then proceeds to block 525, where the training server extracts one or more features of the data element corresponding to the selected record. As discussed above, these features typically correspond to characteristics of the data element, such as the data type, data source, predefined sensitivity or privacy level of the data, and the like. In some embodiments, the features include a data profile for the data element, where a data profile is a metadata structure that specifies the relevant features. In certain embodiments, the training server also extracts one or more data source profiles for the element. A data source profile is typically a metadata structure that specifies the characteristics of the data source. For example, if the data was collected by a specific medical institution, the data source profile may specify the characteristics of the institution (such as its name, location, etc.). Similarly, if the data was collected using a specific device, the profile may specify the identity and type of the device, maintenance records, the accuracy of the device, and the like. In some embodiments, each data record may be associated with any number of profiles corresponding to the entities involved in collecting the data and forwarding it to the data repository.
[0077] At block 530, the training server determines whether to include the data element in the generated data report. If the data is excluded, then the human has determined that including the data would cause the aggregate set to violate one or more data access rules. In contrast, if the data is included, then the reviewer has determined that the selected element still satisfies the access rules when combined with the other included elements.
[0078] Method 500 then proceeds to block 535, where the training server determines whether the selected access record includes at least one additional data record that has not yet been evaluated. If so, method 500 returns to block 520. If not, method 500 proceeds to block 540, where the training server trains one or more machine learning models based on the identified data records that individually satisfy the access rules. In some embodiments, the training server does this by providing the features of each data record (extracted in block 525) as input to the model. The model can be a new model initialized with random weights and parameters, or it can be a partially or fully pre-trained model (e.g., based on a previous training round). Based on the input features, the model in training generates some output for one or more access rules (e.g., classifying the aggregate set as "pass" or "fail"). In embodiments, the training server can compare this generated classification with the actual results determined in block 530 (e.g., the actual set of data elements included in the report) to calculate a loss based on the difference between the actual results and the generated results. This loss is then used to improve (e.g., via backpropagation) one or more internal weights and parameters of the model so that the model learns to more accurately classify the set of aggregated data elements.
[0079] In some embodiments, the training server trains the model to analyze the aggregated data according to a set of access rules. That is, the training server can train the model to evaluate the aggregated data according to all access rules simultaneously and output a binary "pass" or "fail" (or a set of determinations, one for each rule) based on whether the aggregated set passes all access rules or fails at least one access rule. In other embodiments, as discussed above, the training server trains a separate model for each access rule.
[0080] Method 500 then proceeds to block 545 , where the training server determines whether additional training is required. This may include evaluating any termination criteria, such as whether any additional historical access records remain in the training dataset. In various embodiments, other termination criteria may include, but are not limited to, whether a predefined amount of time or computing resources has been expended to train the model, whether the model has achieved a predefined minimum accuracy, and the like. If additional training remains to be completed, method 500 returns to block 510 .
[0081] If not, the method 500 continues to block 550, where the training server deploys the trained model(s) to analyze the set of aggregated data elements retrieved during runtime. In some embodiments, this includes transmitting some indication of the trained model(s) (e.g., weight vectors) that can be used to instantiate the model(s) on another device. For example, the training server can transmit the weights of the trained model(s) to the analysis server. The model can then be used to evaluate the set of aggregated data elements that are determined to individually satisfy the rule.
[0082] Figure 6 is a flow chart depicting a method 600 for controlling data access using a trained machine learning model according to some embodiments disclosed herein. In an embodiment, the method 600 utilizes machine learning and / or a rules engine to provide a common approach across industries to act as a trusted source for retrieving relevant data based on valid requests.
[0083] Method 600 begins at box 605, where an analysis server (e.g., analysis server 110) receives a request for access to data. As discussed above, the request typically indicates the desired data by explicit reference, by providing characteristics that can be used to filter or identify the data, etc. In addition, in embodiments, the request typically indicates the purpose or reason for the request. In some embodiments, the request includes a natural language text description of the requested data and / or the proposed use. For example, the request may include questions such as "Do I have an increased risk for DVT due to my family history? If so, what markers should we screen for?" In some embodiments, the request may additionally include other fields, such as a timeline for which the data is needed (or expected), and any additional supporting documentation that can be provided. In some embodiments, these request features are included in a metadata structure called a request profile (provided directly or generated based on evaluating the request using NLP).
[0084] Method 600 then continues to box 610, where the analysis server processes the request profile using a first set of one or more trained machine learning models. In some embodiments, as discussed above, these models are typically trained to determine whether a request satisfies one or more access rules. For example, to determine whether a request improves human life without harming the data subject, the analysis server may determine whether the request relates to health or well-being (indicating that it benefits human life), whether the use involves commercial gain (indicating that it does not involve commercial gain), etc. In addition, the model can be used to determine whether the proposed use is legitimate (e.g., for clinical or medical use, or whether the user is merely curious or intends to use the data for bad purposes). Similarly, the model can be used to determine whether the data is protected (e.g., whether the data will remain confidential). In some embodiments, as discussed above, a separate trained model is used to evaluate the request according to each separate access rule.
[0085] At block 615, the analysis server determines whether the request passes the access rules based on the classification(s) provided by the model. For example, if the request is for business gain, the analysis server may deny the request.
[0086] If the request is not approved, method 600 continues to block 660, where the analysis server generates a custom report denying the request. In some embodiments, the report includes the reason(s) for the request being denied (e.g., specifying the rule(s) violated). If the request satisfies the access rules, method 600 continues to block 620.
[0087] At block 620, the analysis server retrieves the requested data from one or more data repositories. Method 600 then proceeds to block 625, where the analysis server processes one of the retrieved data elements using a second set of one or more trained models. That is, the analysis server processes each data element individually. In some embodiments, the analysis server uses a single model to evaluate each data element. In another embodiment, the analysis server uses a set of models (e.g., one model per data access rule).
[0088] In some embodiments, processing the data element includes extracting features or characteristics of the data element (e.g., data and / or one or more data profiles of a data source or data generator). These features are then used as input to one or more models. At block 630, the analysis server determines whether the selected data element satisfies all access rules. If not, the method 600 continues to block 635, where the analysis server blocks the selected data element (e.g., marks it as excluded from the custom report, discards it, or otherwise stops processing or considering it). If the analysis server determines that the data element passes the rules, the analysis server adds it to the subset of approved data elements, and the method 600 continues to block 640.
[0089] Continuing with the example of the DVT-related request above, the analysis server may determine that the use of some data elements could improve human life without harming or identifying the data subject, such as the DVT marker being tested and / or identified, the requester's family history, diagnoses of relatives, the type of test performed, etc. In contrast, some example data elements that may not pass this rule because they do not improve human life or may harm the data subject include a doctor's note, the specific identity of a family member who has or has had DVT, etc.
[0090] Similarly, as examples of elements whose uses may be considered legitimate, the analysis server may determine that data such as DVT labeling, diagnosis, and the type of test used have legitimate uses. In contrast, the analysis server may determine, upon request, that data elements such as any non-DVT-related history, tests not related to DVT, and the like do not have legitimate uses. These elements may be restricted. Furthermore, as examples of elements for which the analysis server may determine that data is unprotected, the analysis server may determine that DVT labeling and diagnosis are acceptable, while data elements such as a specific patient's name, date of birth, and non-DVT diagnoses should be excluded.
[0091] Back to Figure 6 At block 640 , the analysis server determines whether any additional data elements have been retrieved but not yet evaluated. If so, the method 600 returns to block 625 . Otherwise, the method 600 continues to block 645 .
[0092] At box 645, the analysis server processes the remaining set of data elements of the aggregation using a third set of one or more machine learning models (e.g., to find data elements that individually satisfy the rules). As discussed above, this can include using the third set(s) of models to provide an aggregated feature set (from each data element in the approved element set). At box 650, the analysis server determines whether the aggregated data passes the data access rules. If so, method 600 continues to box 660, where the analysis server generates a report including the aggregated data. In some embodiments, if any elements are excluded (e.g., at box 635), the analysis server can include an explanation (e.g., identifying the rule(s) that each excluded data element fails to pass).
[0093] If, at block 650, the analysis server determines that the aggregated data fails a set of rules, method 600 proceeds to block 655, where the analysis server excludes at least one data element from the final report. For example, a data element identifying the medical professional or facility location involved may pass an access rule on its own, but when combined with other approved data elements, it may result in the data subject being identified or potentially violate some other access rule. In some embodiments, the analysis server may iteratively remove one or more data elements from the aggregated data and reprocess the remaining set until a satisfactory set of aggregated data elements is found. Method 600 then proceeds to block 660.
[0094] In some embodiments, a notification may also be sent to (multiple) data subjects, thereby informing them what data has been shared. In some embodiments, the notification also indicates the requester, the reason for the request, etc.
[0095] As another example of the evaluation of block 615, assume that an adopted individual requests to know the current location of their biological parent(s) in order to receive information about their medical history. In one embodiment, such a request may be denied at block 615 because the indicated purpose (receiving and viewing the medical history) can be satisfied by a less intrusive request (e.g., a request specifically for data rather than a request for the parents' location).
[0096] As another example, assume that an adopted individual requests general information about their birth parents in order to review their medical history. In an embodiment, the request may pass block 615 (e.g., because the requester is valid and the requested data is consistent with the stated purpose), and data may be retrieved from one or more sources (e.g., a relevant adoption agency) at block 620. At block 630, some data (e.g., the parents' names, adoption date, family history, basic medical history, etc.) may pass the access rules. In contrast, data such as the parents' current contact information, the parents' social security number, etc., will not pass because such data violates the access rules.
[0097] As an example of data elements that, when aggregated, may not pass the evaluation at block 650, consider the example of an adopted child. While data such as the names of the parents and the date or place of adoption may individually pass the rules (at block 630), such data may not pass the evaluation at block 650 (e.g., because such data enables identification of the parents). In contrast, data such as underlying medical history may still pass such an aggregated evaluation.
[0098] As yet another example, suppose an individual already knows the identities of their biological parents and requests that their parents' health insurance company release medical genetic testing information to determine their genetic risk factors. In one embodiment, such a request may pass the evaluation at block 615 because it is a valid request for a valid purpose that satisfies the access rules. At block 630, data such as the parents' physical attributes (e.g., height, weight, BMI, etc.), insurance information, individual responses to surveys or questionnaires (e.g., medication use), and the identification of the company(ies) performing the test may be rejected. In contrast, at block 630, data such as the date the test was performed, the location(s) of the testing facility, and the specific genetic biomarker values found would satisfy the rules. However, at block 650, data such as the facility location, test date, and doctor's notes would not pass the aggregate analysis, while data such as identified biomarkers would.
[0099] As yet another example of the evaluation at block 615, suppose a local government official requests report cards or performance information for all students in the county in order to improve educational outcomes and prevent students from dropping out. In an embodiment, such a request would fail the evaluation at block 615 because the intent (improving outcomes and reducing dropouts) could be met by a less intrusive request that does not share such data.
[0100] Alternatively, suppose a government official wishes to improve educational outcomes and requests information about parents who, out of concern for their children's education, require additional assistance. The request may indicate that the official wishes to increase or change the policy regarding tutoring and / or curriculum for these individuals of concern in order to improve their educational outcomes. In an embodiment, such a request may pass the evaluation at block 615 because it is valid and limited to the least intrusive data necessary to satisfy the intent.
[0101] In an embodiment, data such as parent-teacher notes, parent's name, subject of interest, student's age, teacher's name, tutor, and study techniques being used may satisfy the rule at block 630. This data is relevant and does not harm the subject or otherwise violate the rule. In contrast, data such as the student's specific transcript, the parent's financial situation, the student's specific identity, etc. may not pass the evaluation at block 630 because the data may harm the data subject or is otherwise not necessary to meet the intent.
[0102] Continuing with the above example of a government official requesting information about a parent or student requesting additional assistance, some data such as the student's name (e.g., included in a note between a parent and a teacher), specific scores received on a given test, the name of the parent or teacher, etc., may not pass the evaluation at block 650. Such data may be harmful to the subject as a whole. In contrast, data such as subject of concern, known learning disabilities, age group or range, etc., may pass the evaluation and be included in the report.
[0103] As yet another example of the application of method 600, assume that a government official requests information about the individual taxpayers for each household in a county in order to provide a tax refund specific to the taxpayer. In an embodiment, this request may pass the evaluation at box 615 because the requester identity and request / intention are valid. At box 630, data such as the taxpayer's (multiple) social security numbers, total income, number of dependents, zip code, etc. may each pass the individual analysis at box 630 because the data can satisfy the request without harm. In contrast, data such as the individual's citizenship, identity, disability status, etc. may not pass. Overall, at box 650, data such as social security numbers, number of dependents, total income, etc. may not pass because the data may harm the subject. In contrast, data such as the number of taxpayers in the area may pass.
[0104] As an additional example, suppose a government official requests information about the number of individuals eligible for publicly offered insurance coverage in order to provide coverage to all residents. Such a request may pass the evaluation at block 615. At block 630, data such as each subject's household income, the ZIP code where they live, and pre-existing health conditions may not pass the access rules. In contrast, data such as their taxpayer information, age, Social Security number, residence, employment status, etc. may pass this individual evaluation because the data can be used to serve the request without harm. However, at block 650, data such as their Social Security number, age, marital status, etc. will not pass the aggregate review, while data such as their insurance coverage eligibility, name, etc. may pass.
[0105] As yet another example, suppose an airline requests the identities of all individuals who came into contact with a contagious individual within a specific time period in order to minimize the risk of disease transmission and inform relevant passengers of the concern. Such a request may not pass the evaluation at block 615 because it could be resolved with a less intrusive request.
[0106] Continuing with the above example, assume that instead the airline requests a determination as to whether any passengers have been in contact with an infectious individual (without specifically identifying the passenger). In an embodiment, this request can pass the evaluation at box 615. At box 620, relevant data such as the passenger's identity, location(s) (e.g., using social media or GPS), calendar, relevant testing agencies, and laboratory results can be retrieved. At box 630, data such as the passenger's personal name or identity, age, pre-existing conditions, etc. can be excluded. However, data such as contact tracing information (e.g., location data), current health results, etc. can be included. At box 650, in general, data such as the names of those who have been in contact with the passenger, the passenger's age, the current location of the potentially infectious individual, etc. can be excluded. In contrast, data such as a binary "yes" or "no" indication as to whether someone has been in contact with an infectious person, and whether the contact was within a predefined time, can pass the rules.
[0107] As another example, suppose a patient (or potential patient) requests information from one or more institutions about patients who have undergone retinal detachment surgery in order to select a treatment plan and medical provider for their own surgery. Such a request would fail the evaluation at block 615 because the intent could be satisfied by a less intrusive request or data.
[0108] Alternatively, suppose the patient requests information about the success rate of such surgery or any permanent damage or injury resulting from the surgery. This request may pass the evaluation at block 615. At block 630, data such as the specific location of the practitioner, specific patient information, etc., will not pass the review. In contrast, data such as a list of the practitioners and / or surgeons who performed the surgery, an indication of factors affecting the success rate, a list of medical equipment used in the surgery, the patient's eye measurements, or other data may pass the evaluation.
[0109] However, data such as specific practitioners or individual surgeons with low success rates, specific equipment used in surgery, etc. may be excluded at block 650. In contrast, data such as a list of surgeons with high success rates, indications of complications or injuries, etc. may be included.
[0110] As another example, suppose a government official or contracted nonprofit organization requests data related to a clinical trial of a vaccine currently under development in order to evaluate the results. At block 615, the request may be evaluated and the data retrieved. At block 630, data such as vaccine composition, side effects, dates or times the vaccine was administered, dosage, reports of concern, independent evaluations of the vaccine, development phase, trial phase, number of participants, reported adverse events, and indications for patients who withdrew from the trial may all pass the access rules. In contrast, data such as specific patient names, locations, and addresses may not pass this review.
[0111] Continuing with this example, at block 650, the name and location of the specific trial, the vaccine's pricing structure, cost, etc., can all be excluded overall. In contrast, data such as vaccine efficacy, antibody or immune response by age group, reported side effects, etc., can be included in the report through this aggregate review.
[0112] As another example, suppose a researcher requests access to raw image data of patients undergoing surgery in order to analyze morphological changes around the surgical site over time. Such a request may pass the evaluation at block 615. At block 630, data such as the hospital or location where the surgery was performed, the equipment used during the surgery, the imaging device used to collect the data, physician notes, patient complaints, side effects, etc. may pass a separate evaluation. In contrast, data such as the location of the institution(s), the patient's name, the patient's medical history, etc., would not pass this review.
[0113] At block 650, data such as the raw data, a summary of pre-existing conditions related to the image or procedure, etc. can be reviewed through the aggregation. In contrast, data such as physician notes not relevant to the image analysis, the physician's name or identity, the specific medical device used to capture the image, side effects not related to the image, etc. can be excluded.
[0114] As yet another example of the application of method 600, suppose a school teacher requests information from a health service provider regarding all records of any health treatment a given student has received because the student is suspected of being abused. In an embodiment, this request will fail the evaluation at block 615 because, although the requester is legitimate, more records are requested than are needed. For example, if the request targets the frequency of absences or tardiness, the number of medical visits, etc. (which does not require a request for specific health data), the request may satisfy the rule.
[0115] As another example, suppose a professor requests a student's location data because he suspects the student is falsely reporting a family emergency in order to skip class. In an embodiment, such a request would fail the evaluation at block 615 because the request would be harmful to the data subject (and could be satisfied if different data were requested).
[0116] As yet another example, assume that an adult child of a deceased individual requests access to the deceased individual's social network account in order to download photos and videos for a video memorial. Such a request may be evaluated at block 615. At block 630, information such as publicly available pictures and videos from the social media account, friend lists, etc. may be reviewed separately. In contrast, data such as non-public information, saved posts or content, and private conversations of the deceased may be excluded.
[0117] As yet another example, suppose someone requests access to a deceased relative's social network account in order to determine whether the deceased relative engaged in illegal activity. At block 615, such a request may be denied because the intent does not satisfy the access rules (e.g., the intent does not improve human life or may harm the data subject).
[0118] As another example, suppose an individual wishes to obtain diagnostic information for a family member in order to obtain a second opinion. Such a request may pass the evaluation at block 615. At block 630, data such as laboratory results relevant to the diagnosis, genetic predisposition, or symptoms may pass the evaluation, while data such as information not relevant to the diagnosis (e.g., blood type) may be excluded. At block 650, information such as the patient's name, doctor's name, hospital identification or location may be excluded, while data such as relevant laboratory results may pass the access rules.
[0119] Figure 7A graphical user interface (GUI) 705 for enhancing data access control and notification according to some embodiments disclosed herein is depicted. In the illustrated embodiment, the GUI 705 includes a series of data elements 710A-J and an indication of whether the data element is shared (or shareable) for each data element 710. In the illustrated embodiment, the GUI 705 uses a sliding indicator with one position corresponding to blocked data elements / data that are not shared (e.g., the leftmost side of the slider), one position corresponding to data that is sometimes or limitedly shared (e.g., on a case-by-case basis, depending on a particular request) (e.g., the center of the slider), and one position corresponding to data that is always or freely shared (e.g., the rightmost side of the slider). In some embodiments, each of the data elements 710 is associated with other visual aids, such as color coding (e.g., red, yellow, and green).
[0120] In some embodiments, the user can provide preferences or selections to the analysis server using the GUI 705. For example, the user can specify that although an element is shared (or selectively shared), they would prefer that the element always be locked. Alternatively, the user can indicate that although a data element is blocked, they would like the data element to be shareable (at least selectively). In some embodiments, such user feedback can be used to iteratively improve the model used to make access decisions.
[0121] Figure 8is a flow chart depicting a method 800 for automatically controlling data access using one or more machine learning models, according to some embodiments disclosed herein. Method 800 begins at block 805, where an analysis server receives a first request from a first user for data related to a second user. At block 810, the analysis server automatically determines whether the first request satisfies one or more data access rules by processing the first request using a first set of one or more trained machine learning models. Method 800 then continues to block 815, where, after determining that the first request satisfies the one or more data access rules, the analysis server automatically retrieves a first plurality of data elements based on the first request. At block 820, the analysis server automatically determines whether each data element in the first plurality of data elements satisfies the one or more data access rules by processing each data element in the first plurality of data elements using a second set of one or more trained machine learning models. Additionally, at block 825, after determining that each data element in the first set of data elements from the first plurality of data elements individually satisfies the one or more data access rules, the analysis server determines whether the first set of data elements collectively satisfies the one or more data access rules by processing the first set of data elements using a third set of one or more trained machine learning models. At block 830 , upon determining that the first set of data elements satisfies the one or more data access rules, the analytics server generates a customized report that includes the first set of data elements.
[0122] Figure 9is a flow chart depicting a method 900 for training one or more machine learning models to control data accessibility, according to some embodiments disclosed herein. Method 900 begins at block 905, where a training server generates a first training dataset from a set of historical access records, wherein each respective access record in the first training dataset corresponds to a respective data request and includes information identifying whether the respective request satisfies one or more data access rules. At block 910, the training server generates a second training dataset from the set of data records, wherein each respective data record in the second training dataset corresponds to a respective data element and includes information identifying whether the respective data element satisfies one or more data access rules. Additionally, at block 915, the training server generates a third training dataset from the set of historical access records, wherein each respective access record in the third training dataset corresponds to a respective set of aggregated data elements and includes information identifying whether the respective set of aggregated data elements satisfies one or more data access rules. Method 900 then continues to block 920, where the training server trains one or more machine learning models based on the first, second, and third training datasets to generate an output identifying whether the data request should be granted. At box 925, the training server then deploys the one or more machine learning models to one or more computing systems.
[0123] Example system for training and using machine learning models with controlled data access
[0124] Figure 10 is a block diagram depicting a computing device 1000 configured to train and use a machine learning model to control data access according to some embodiments disclosed herein. For example, the computing device 1000 may include Figure 1 One or more of the illustrated analysis servers 110 and / or training servers 135. The computing device 1000 may be configured to perform various techniques disclosed herein, such as those described in connection with FIG. Figures 2 to 9 Describe the methods and techniques.
[0125] As shown, the computing device 1000 includes a central processing unit (CPU) 1005, one or more I / O device interfaces 1020 that can allow various I / O devices 1035 (e.g., a keyboard, a display, a mouse device, a pen input, etc.) to be connected to the computing device 1000, a network interface 1025 through which the computing device 1000 can be connected to one or more networks (which can include a local network, an intranet, the Internet, or any other group of computing devices communicatively connected to each other), memory 1010, storage 1015, and interconnects 1030.
[0126] CPU 1005 can retrieve and execute programming instructions stored in memory 1010. Similarly, CPU 1005 can retrieve and store application data present in memory 1010. Interconnect 1030 transfers programming instructions and application data between CPU 1005, I / O device interface 1020, network interface 1025, memory 1010, and storage 1015.
[0127] CPU 1005 is included to represent a single CPU, multiple CPUs, a single CPU with multiple processing cores, and the like.
[0128] Memory 1010 represents a volatile memory such as random access memory or a non-volatile memory such as non-volatile random access memory, phase change random access memory, etc. As shown, memory 1010 includes a data sharing component 115, a custom report generator 120, a training data generator 140, and a model trainer 145.
[0129] The data sharing component 115 is generally configured to evaluate requests and data elements to determine whether the request and data elements should be shared (e.g., whether access should be granted to the requesting entity). In embodiments, the data sharing component 115 does this based in part on a set of access rules that define ethical and acceptable data security and access practices. In some embodiments, the data sharing component 115 utilizes a machine learning model that has been trained on historical access records 150.
[0130] The custom report generator 120 generally generates a data report based on the decision returned by the data sharing component 115. That is, the custom report generator 120 generates a report that includes any data elements that are approved for sharing (individually and collectively). In some embodiments, for any excluded data, the custom report generator 120 may include an indication of which rule(s) the element(s) failed to satisfy (e.g., based on a particular model that classified the data element as unsatisfied).
[0131] The training data generator 140 typically generates a training dataset from historical access records. Each record in the training dataset indicates a set of input features (of corresponding historical requests or data elements) and a target output label (e.g., whether the historical requests or data elements satisfy an access rule).
[0132] The model trainer 145 typically uses a training dataset to train a set(s) of trained models 1050 , which are used by the data sharing component 115 to drive data access decisions.
[0133] Additional considerations
[0134] The foregoing description is provided to enable any person skilled in the art to practice the various embodiments described herein. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments. For example, without departing from the scope of this disclosure, the function and arrangement of the elements discussed may be changed. Various examples may appropriately omit, replace, or add various programs or components. In addition, the features described with respect to some examples may be combined in some other examples. For example, any number of aspects set forth herein may be used to implement a device or practice method. In addition, the scope of this disclosure is intended to cover devices or methods practiced using other structures, functions, or structures and functions in addition to or different from the various aspects of this disclosure set forth herein. It should be understood that any aspect of this disclosure disclosed herein may be embodied by one or more elements of the claims.
[0135] As used herein, the phrase "at least one of" a list of items refers to any combination of those items, including single members. For example, "at least one of a, b, or c" is intended to encompass any combination of a, b, c, ab, ac, bc, and abc, as well as multiples of the same element (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other order of a, b, and c).
[0136] As used herein, the term "determining" includes a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, database, or other data structure), ascertaining, etc. Furthermore, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Furthermore, "determining" may include resolving, selecting, choosing, establishing, etc.
[0137] The method disclosed herein includes one or more steps or actions for implementing the method. Without departing from the scope of the claims, the method steps and / or actions can be interchangeable with each other. In other words, unless a specific step or action sequence is specified, the order and / or use of the specific steps and / or actions can be modified without departing from the scope of the claims. Further, the various operations of the above method can be performed by any suitable device that can perform the corresponding function. The device may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs) or processors. Typically, where operations are shown in the figures, those operations may have corresponding devices with similar numbers plus functional components.
[0138] The various illustrative logical blocks, modules, and circuits described in conjunction with the present disclosure may be implemented or executed with a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but in the alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0139] The processing system can be implemented using a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus can include any number of interconnecting buses and bridges. The bus can link together various circuits including a processor, machine-readable media, and input / output devices. User interfaces (e.g., keyboard, display, mouse, joystick, etc.) can also be connected to the bus. The bus can also link various other circuits such as timing sources, peripherals, voltage regulators, power management circuits, etc., which are well known in the art and therefore will not be described further. The processor can be implemented using one or more general and / or special purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuit systems that can execute software. Those skilled in the art will recognize how to best implement the described functions of the processing system based on the specific application and the overall design constraints imposed on the entire system.
[0140] If implemented in software, the functions may be stored as one or more instructions or codes on or transmitted via a computer-readable medium. Software should be broadly interpreted as instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. Computer-readable media include both computer storage media and communication media (e.g., any medium that facilitates the transfer of computer programs from one place to another). The processor may be responsible for managing the bus and general processing, including the execution of software modules stored on the computer-readable storage medium. The computer-readable storage medium may be coupled to the processor so that the processor can read information from and write information to the storage medium. In an alternative embodiment, the storage medium may be integral to the processor. For example, the computer-readable medium may include a transmission line, a carrier modulated by data, and / or a computer-readable storage medium on which instructions are stored that are separate from the wireless node, all of which can be accessed by the processor via a bus interface. Alternatively or in addition, the computer-readable medium or any portion thereof may be integrated into the processor, as may be the case with a cache and / or general register file. For example, examples of machine-readable storage media may include RAM (random access memory), flash memory, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), registers, magnetic disks, optical disks, hard drives, or any other suitable storage media, or any combination thereof. Machine-readable media may be embodied in a computer program product.
[0141] A software module may include a single instruction or multiple instructions and may be distributed across several different code segments, between different programs, and across multiple storage media. A computer-readable medium may include multiple software modules. A software module includes instructions that, when executed by a device such as a processor, cause a processing system to perform various functions. A software module may include a transmission module and a reception module. Each software module may reside in a single storage device or be distributed across multiple storage devices. For example, when a triggering event occurs, a software module may be loaded from a hard drive into RAM. During execution of a software module, the processor may load some instructions into a cache to increase access speed. One or more cache lines may then be loaded into a general register file for execution by the processor. When referring to the functionality of a software module, it should be understood that such functionality is implemented by the processor when executing instructions from that software module.
[0142] The following claims are not intended to be limited to the embodiments shown herein, but are to be given the full scope consistent with the language of the claims. In the claims, unless otherwise specified, references to singular elements are not intended to mean "one and only one", but "one or more". Unless otherwise specifically stated, the term "some" refers to one or more. Pursuant to 35 U.S.C. § 112 (f), no element of any claim will be interpreted unless the element is expressly recited using the phrase "means for..." or, in the case of a method claim, the element is recited using the phrase "step for...". All structural and functional equivalents of the elements of the various aspects described throughout this disclosure that are known or will later be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, regardless of whether such disclosure is explicitly recited in the claims, the content disclosed herein is not intended to be committed to the public.
Claims
1. A method for automatically controlling data access using one or more machine learning models, the method comprising: receiving a first request from a first user for data related to a second user; automatically determining whether the first request satisfies one or more data access rules by processing the first request using a first set of one or more trained machine learning models; automatically retrieving a first plurality of data elements based on the first request after determining that the first request satisfies the one or more data access rules; automatically determining whether each data element in the first plurality of data elements satisfies the one or more data access rules by individually processing each data element in the first plurality of data elements using a second set of one or more trained machine learning models; after determining that each data element in a first set of data elements from the first plurality of data elements individually satisfies the one or more data access rules, determining whether the first set of data elements collectively satisfies the one or more data access rules by processing the first set of data elements using a third set of one or more trained machine learning models; upon determining that the first set of data elements satisfies the one or more data access rules, generating a customized report including the first set of data elements; receiving a second request; automatically retrieving a second plurality of data elements based on the second request; automatically determining that a second set of data elements from the second plurality of data elements satisfies the one or more data access rules by processing each data element in the second plurality of data elements using the second set of one or more trained machine learning models; determining whether the second set of data elements collectively satisfy the one or more data access rules by processing the second set of data elements using the third set of one or more trained machine learning models; as well as Upon determining that the second set of data elements do not collectively satisfy the one or more data access rules, providing at least one data element from the second set of data elements is avoided.
2. The method of claim 1, further comprising: receiving a second request; determining whether the second request satisfies the one or more data access rules by processing the second request using the first set of one or more trained machine learning models; as well as Upon determining that the second request does not satisfy the one or more data access rules, retrieving data for the second request is avoided.
3. The method of claim 1, further comprising: Upon determining that a second set of data elements from the first plurality of data elements does not satisfy the one or more data access rules, providing the second set of data elements is avoided.
4. The method of claim 1, further comprising: Notification is transmitted to the second user that the first set of data elements was accessed by the first user.
5. The method of claim 1, further comprising: receiving a second request; determining that the second request does not satisfy the one or more data access rules; as well as Generate a custom report specifying the reason why the second request was denied.
6. A method for training one or more machine learning models to control data accessibility, the method comprising: generating a first training data set from a set of historical access records, wherein each corresponding access record in the first training data set corresponds to a corresponding data request and includes information identifying whether the corresponding request satisfies one or more data access rules; generating a second training data set from a set of data records, wherein each respective data record in the second training data set corresponds to a respective data element and includes information identifying whether the respective data element satisfies the one or more data access rules; generating a third training data set from the set of historical access records, wherein each respective access record in the third training data set corresponds to a respective set of aggregated data elements and includes information identifying whether the respective set of aggregated data elements satisfies the one or more data access rules; training the one or more machine learning models based on the first training data set, the second training data set, and the third training data set to generate an output identifying whether a data request should be granted; and Deploy the one or more machine learning models to one or more computing systems.
7. The method according to claim 6, wherein: Training the one or more machine learning models based on the first training dataset, the second training dataset, and the third training dataset includes: training a first set of machine learning models among the one or more machine learning models based on the first training dataset; training a second set of machine learning models among the one or more machine learning models based on the second training dataset; and A third group of machine learning models among the one or more machine learning models is trained based on the third training data set.
8. The method of claim 7, wherein: The one or more data access rules include: (i) First Rule; (ii) the second rule; and (iii) The third rule.
9. The method of claim 8, wherein: Training a first set of the one or more machine learning models includes: training a first machine learning model based on the first training data set and the first rule, training a second machine learning model based on the first training dataset and the second rule, and training a third machine learning model based on the first training data set and the third rule, Training a second set of the one or more machine learning models includes: training a fourth machine learning model based on the second training data set and the first rule, training a fifth machine learning model based on the second training dataset and the second rule, and training a sixth machine learning model based on the second training data set and the third rule, and Training a third set of the one or more machine learning models includes: training a seventh machine learning model based on the third training data set and the first rule, training an eighth machine learning model based on the third training dataset and the second rule, and A ninth machine learning model is trained based on the third training dataset and the third rule.
10. The method of claim 8, wherein: The first rule states that data should be accessed only when access to the data can improve human life. The second rule provides that data may only be accessed if its intended use is legitimate, and The third rule stipulates that data can only be accessed while they are still protected.
11. The method according to claim 6, wherein: Each corresponding access record in the first training dataset further includes information identifying the following: the purpose of the corresponding request; and One or more data elements associated with the corresponding request.
12. The method of claim 6, wherein: Each corresponding data record in the second training data set further includes information identifying: One or more characteristics of the corresponding data element.
13. The method of claim 6, wherein: Each corresponding data record in the second training data set further includes information identifying: A data profile of a source of each data element in the corresponding set of aggregated data elements.
14. The method of claim 6, wherein: Each corresponding data record in the third training data set further includes information identifying: The data configuration file of the source of the corresponding data element.
Citation Information
Patent Citations
Smart access control system for implementing access restrictions of regulated database records based on machine learning of trends
US20190258818A1
Automated access control management for computing systems
US20190327271A1