Secure Data Analysis

By identifying and masking sensitive data fields, and providing masked datasets to analytics providers, the method ensures privacy and maintains control over data usage, ensuring accurate analytical results without exposing the original data.

JP7828701B2Active Publication Date: 2026-03-12INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

The challenge arises when external analytics providers have access to sensitive data, posing a risk of data misuse or unauthorized access, and existing methods fail to maintain privacy and control over data masking while ensuring analytical accuracy.

Method used

A method is implemented to identify sensitive data fields, apply masking techniques, and provide a masked dataset to analytics providers, along with a request for analytical processing, allowing the generation of analytical functions that can be executed on the original dataset while maintaining privacy and accuracy.

Benefits of technology

This approach ensures privacy of sensitive data, maintains control over masking methods, and preserves analytical accuracy by generating functions that produce consistent results on the original dataset, thus preventing unauthorized data access and misuse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007828701000001
    Figure 0007828701000001
  • Figure 0007828701000002
    Figure 0007828701000002
  • Figure 0007828701000003
    Figure 0007828701000003
Patent Text Reader

Abstract

Secure data analysis is provided through a process of identifying sensitive data fields of an initial dataset and mappings between the sensitive data fields and other data fields of the dataset, where an analytical operation is to be performed on the initial dataset; then selecting and applying a masking method to the initial dataset to mask the sensitive data fields based on expected data fields of the initial dataset and based on the identified sensitive data fields to be used in performing the analytical operation to produce a masked dataset; providing the masked dataset along with a request for the analytical operation to an analysis provider in response to which receiving a generated analytical function generated based on the masked dataset configured to perform the analytical operation; and invoking the generated analytical function on the initial dataset to perform the analytical operation on the initial dataset.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] An analytics provider is an organization with the expertise to analyze data, especially large data sets, to extract insights and other valuable information from the data. While some data owners perform their own analyses, it is more common for data owners to hire external data analytics providers to perform analyses on their data. This traditionally requires the data owner to provide the data, or at least access to the data, to the analytics provider in order for the provider to perform the analysis. Summary of the Invention

[0002] Shortcomings of the prior art are overcome and additional advantages are provided through the provision of a computer-implemented method. The method identifies sensitive data fields in an initial dataset on which an analytical process is to be performed and a mapping between the sensitive data fields and other data fields in the dataset. Based on the expected data fields of the initial dataset to be used in performing the analytical process on the initial dataset and further based on the identified sensitive data fields, the method selects and applies a masking method to the initial dataset to mask the sensitive data fields of the initial dataset to produce a masked dataset. The method provides the masked dataset to an analysis provider along with a request for the analytical process and receives a generated analytical function in response, the generated analytical function being generated based on the masked dataset and configured to perform the analytical process on the initial dataset. The method also invokes the generated analytical function on the initial dataset to perform the analytical process on the initial dataset.

[0003] Also provided is a computer system including a memory and a processor in communication with the memory, the computer system configured to perform a method. The method identifies sensitive data fields in an initial dataset on which an analytical operation is to be performed and mappings between the sensitive data fields and other data fields in the dataset. Based on expected data fields in the initial dataset to be used in performing the analytical operation on the initial dataset and further based on the identified sensitive data fields, the method selects and applies a masking method to the initial dataset to mask the sensitive data fields of the initial dataset to produce a masked dataset. The method provides the masked dataset to an analysis provider along with a request for analytical operation and receives a generated analytical function in response, the generated analytical function being generated based on the masked dataset and configured to perform the analytical operation on the initial dataset. The method also invokes the generated analytical function on the initial dataset to perform the analytical operation on the initial dataset.

[0004] Additionally, a computer program product is provided for implementing a method, the computer program product including a computer-readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit. The method identifies sensitive data fields in an initial dataset on which an analytical operation is to be performed and a mapping between the sensitive data fields and other data fields in the dataset. Based on the expected data fields of the initial dataset to be used in performing the analytical operation on the initial dataset and further based on the identified sensitive data fields, the method selects and applies a masking method to the initial dataset to mask the sensitive data fields of the initial dataset to produce a masked dataset. The method provides the masked dataset to an analysis provider along with a request for analytical operation and receives a generated analytical function in response, the generated analytical function being generated based on the masked dataset and configured to perform the analytical operation on the initial dataset. The method also invokes the generated analytical function on the initial dataset to perform the analytical operation on the initial dataset.

[0005] Additional features and advantages are realized through the concepts described herein.

[0006] The aspects described herein are particularly pointed out and distinctly claimed as examples in the claims at the conclusion of this specification. The foregoing and other objects, features, and advantages of the present disclosure will become apparent from the following detailed description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 illustrates an example environment for incorporating and using aspects described herein. [Figure 2] FIG. 1 is a diagram of an example of an initial data set. [Figure 3] FIG. 1 is a diagram of an example masked dataset. [Figure 4]FIG. 1 illustrates another example environment for incorporating and using aspects described herein. [Figure 5] FIG. 1 is an illustration of an example methodology for secure data analysis according to aspects described herein. [Figure 6] FIG. 1 illustrates an example of a computer system and associated devices for incorporating and / or using aspects described herein. [Figure 7] 1 is a diagram of a cloud computing environment in accordance with an embodiment of the present invention. [Figure 8] FIG. 2 is a diagram of abstraction model layers according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0008] An approach for secure data analysis is described herein. Concerns can arise when an external analytics provider (AP) is contracted to perform analysis on a dataset held or controlled by another entity, such as a data owner or provider. Typically, the AP will have the authority to analyze the dataset indiscriminately, which may, and often is, a large enterprise dataset containing sensitive data. The dataset may contain private / sensitive information, such as personally identifiable information (PII), also known as sensitive personal information, personally identifying information, personally identifiable information, personally identifiable information, personal information, or personal data, or a combination thereof, typically abbreviated as "PII" or "SPI." AP access to the private information of users or other entities to which the data pertains poses a risk. For convenience, such users / entities are referred to herein as "subjects" or "subject entities." An AP or another entity that gains access to the data may abuse this access, for example, by selling private information to a third party not specified in the customer consent, by taking some action using the data, or by deriving some other illicit benefit from the data, or a combination thereof.

[0009] Thus, at least the following challenges are identified: Protection may be desired to maintain the privacy of subjects reflected in the data of the dataset to be processed and to prevent the analytics provider from gaining unwanted insights (from the data owner's perspective) into the subject's private data; It may be desired for the data owner of the dataset, who is the customer that requested the analysis to be performed on the dataset by the AP, to maintain control over the methods used to mask the data based on one or more queries performed by the analytics provider; It may further be desired for the foregoing to be performed using functional programming, in which case the logic for running and performing the analysis can also be provided to the customer to run on the live data, along with the data, to help ensure that the accuracy of the analytical model is not affected, that subject privacy is largely maintained, and that the conclusions or results of the analysis are still the same regardless of any masking performed on the dataset before providing it to the analytics provider.

[0010] FIG. 1 depicts an example environment for incorporating and using aspects described herein. Referring to FIG. 1, the environment 100 includes a data owner 102 that holds / owns a dataset of data related to a subject entity. An analysis provider 104 is a provider of data analytics, meaning that the provider receives the dataset from its customer (the data owner) and can then build and implement an analytical process to process the data according to any analysis or insights requested by the customer. Conceptually, between the data owner 102 and the analysis provider 104 is a network 106, which can include not only a network infrastructure for network communication but also a cloud environment. In some examples, the data owner 102 holds the dataset on-premises and sends the dataset over the network 106 to the analysis provider 104, who performs the analysis on its own computer system and then sends the results back to the data owner 102. In other examples, some or all aspects of these exchanges are facilitated using a cloud platform. For example, the dataset may be stored on a customer's private cloud (present as part of network 106), or the analysis provider 104 may perform its analytical processing on / using a cloud platform present as part of network 106. In either case, the data owner, sometimes referred to herein as the analysis provider's "customer," controls access to and provision of the dataset. On the analysis provider side, the analysis provider receives requests from the analysis provider's customer (data owner) to perform analysis on the dataset, and the analysis provider receives access rights to the dataset from the analysis provider's customer.

[0011] 1 form a wired or wireless network of devices, and communication between the devices occurs via wired or wireless communication links 112 for communicating data between the devices. Figure 1 is only one example of an environment for incorporating and using aspects described herein, and many other examples are possible and contemplated as being compatible with the capabilities described herein.

[0012] Aspects described herein provide dynamic maintenance of privacy dependencies and curation mechanisms among relational data in a multi-tier framework for secure data analysis. As part of this, and at the data owner's end, various characteristics of the dataset can be identified automatically, manually, or both based on the dataset itself and metadata (e.g., schema information) stored near the dataset. Such characteristics can include, by way of example, characteristics of data fields within the dataset and the mappings and dependencies between data fields. Data fields correspond to categories of data points stored for subjects reflected in the dataset. In a relational database, each column or "attribute" might correspond to a respective data field, for example.

[0013] As an example, a data owner's computer system (owned by or acting on behalf of the data owner) may identify sensitive data fields, i.e., fields containing sensitive / private data. For example, it may be desirable to mask aspects of the sensitive data fields, such as (i) masking the actual data in these fields (e.g., social security numbers) or (ii) anonymizing / altering any mappings between the sensitive fields and any personal identifiers (IDs) of the subject entities (e.g., users), or both. In this regard, mappings and dependencies may exist between data fields to link together data corresponding to a particular subject entity. One option to help anonymize data is to modify a mapping that correlates all of the data for an identified subject entity in a dataset so that the mapping no longer correlates all of the entity's data to this particular entity. An example of a modification is remapping or deleting the mapping. Additionally or alternatively, actual subject entity attribute data (username, Social Security Number (SSN), address, phone number, etc.) may be masked, particularly if the changes to the mapping are not sufficient to anonymize or decorrelate one or more attributes from the subject entity.

[0014] On the data owner side, the system can also identify non-sensitive fields, i.e., fields that do not contain data for which association to a specific entity is desired to be masked. In this regard, it may be acceptable to maintain any correlation of subject-entity identifiers, such as user IDs, to non-sensitive data fields. Whether a field is sensitive or non-sensitive may also depend on the context. For example, in a dataset of all of a company's U.S. customers, a data field about shipping country may be considered non-sensitive because each person represented in the dataset has a matching shipping country (U.S.), which does not identify any particular customer. In contrast, when dealing with a dataset reflecting just one individual from each of several different countries, the shipping country may be considered sensitive.

[0015] As part of identifying sensitive and non-sensitive data fields and the mappings and dependencies between them, the system can provide the data owner with recommendations about which data fields in the dataset are sensitive, which are non-sensitive, or both. This can be done along with recommending a list of candidate data fields that should be sensitive, masked, or both. The recommended list can be modified and further customized by the data owner to make a final determination about which data fields should be considered sensitive for this particular dataset. For example, the data owner can add and / or remove sensitive fields from the list. Often, the data owner will have a reasonable understanding of the stimuli for the analysis, their current and past analytical requirements, and which data fields have been or should be masked historically. This can be useful in determining which data fields are considered sensitive and should be masked to perform the specific analysis desired by the data owner.

[0016] In some examples, information representing analytical requirements can be stored in a specific format (e.g., Extensible Markup Language (XML) or JavaScript Object Notation (JSON) format) for input to an artificial intelligence (AI)-based intelligent search and text analysis platform, such as the discovery platform provided by International Business Machines Corporation of Armonk, New York, USA. Custom rule-based machine learning / natural language processing models can be established based on past analysis and current and past usage to train the model to predict relationships between corresponding fields, including those that should be masked because it is anticipated that analytics will require access to data in those fields to perform analytical processing. Models specific to individual data holders can be trained, or aggregate models based on information from a population of different data holders can be trained. Models can be trained based on past analysis requirements and masking to provide suggestions / recommendations on which data fields should be masked and which should be retained within a given dataset.

[0017] In addition to identifying the sensitivity of data fields and the mappings that exist between them, the data owner's system also understands the masking method to be applied to mask sensitive data fields of a dataset to produce a masked dataset. The dataset discussed above containing sensitive data is referred to herein as the "initial" dataset. According to aspects described herein, at least one masking method is selected and applied to the initial dataset to mask sensitive fields of the initial dataset to produce another dataset, referred to herein as the "masked" dataset. The data owner can provide the masked dataset to an analysis provider, which can use it to build an analysis function configured to perform the requested analysis. Analysis providers often require a set of data to generate the correct analysis function (which may consist of several functions / queries) that performs the desired analysis on the data. The data owner provides the analysis provider with data by providing the masked dataset rather than the initial dataset to generate the appropriate analysis function without revealing the sensitive data of the initial dataset. The generated analysis function can then be made available to the data owner to run against the initial dataset containing the real (unmasked) data, thus delivering the requested analysis results specific to the actual data on which the analysis is intended to be performed.

[0018] Thus, the masking method can be selected / specified by the analytics provider's data owner / customer. While the customer does not know the analytical functions to be generated or which data / data fields the analytics provider will need access to in generating those analytical functions, past experience and machine learning can inform the specific functions that may be performed and / or the data fields that may be required to generate the analytical functions, and therefore which fields may need to be passed and / or masked. In some instances, the selection of the masking method is based in part on a prediction of the type of analysis / function the analytics provider will perform on the dataset, the specific sensitive and / or non-sensitive data fields identified as discussed above, and / or this understanding of the current / historical relationship between the analytical requirements (particularly as informed by the analysis requested by the customer) and the corresponding data fields to be masked, or a combination thereof. Thus, the customer's selection and application of the masking method can be based at least in part on (i) which data fields will be identified as sensitive and (ii) which data fields from the initial dataset will ultimately be used in performing the requested analysis. If any data fields that are to be used are sensitive, they can be masked. Other sensitive data fields in the initial dataset can be similarly masked or omitted entirely from the masked dataset.

[0019] The specific masking method for masking data and / or mappings between data fields and subject entity (e.g., user) IDs, in some instances a hash function, remains secret to the data owner, meaning that the analysis provider does not know what was masked and / or how it was masked. Based on the above identification information about sensitive / non-sensitive attributes and which attributes are expected to be used in generating the analysis features, the system can mask sensitive data fields in the initial dataset to be sent to the analysis provider as part of the masked dataset. Masking a data field refers to (i) modifying (including deleting, editing, expanding, randomizing, etc.) the subject entity data for this data field / attribute, or (ii) modifying (deleting / removing / remapping) one or more mappings between this data field and one or more other data fields in the dataset, or both. In some instances, masking involves masking data fields by masking the attributes of each entity, i.e., by changing these attribute values ​​(e.g., changing a 9-digit Social Security number to be "111-11-1111" or "XXX-XX-XXXX"), or by randomly remapping any mappings and relationships that the data fields have from one set of fields to a different set of fields using equations, or both. In other instances of changing mappings, mappings are removed / deleted so that the association between attributes no longer exists.

[0020] Note that the particular masking method used by the data owner can be changed as desired. The data owner can select a different masking method each time the data owner requests analysis by the AP, whether for the same data set or a different data set. New and / or different randomization and mapping strategies can be used to prevent the AP from understanding the data owner's randomized inputs over time, thus further ensuring the entity's privacy.

[0021] As an illustrative example, we first refer to Figure 2, which depicts an example initial dataset. The example dataset in this example is very basic, showing only three records for three different subject entities and only five data fields / attributes for each of the records. Each record includes a first name, a last name, birth data, an identification number, and a locale. The identification number is a unique number assigned to the subject entity's locale, in this example, a country, and uniquely identifies the subject entity in this locale. In this basic example, we assume that only the identification number data field is considered sensitive.

[0022] Figure 3 depicts an example masked dataset, which is the initial dataset from Figure 2, but with a masking method applied to mask the data in the sensitive data fields, and (conceptually) illustrates the remapping after the masking method was applied. The masking method, here, masks the identification numbers for each of the three subject entities, replacing each identification number with similarly formatted random numbers / letters. Additionally, the remapping remapped some of the data fields. The first name and last name data fields for the Tobias Baumer entity were remapped to the birth date, masked identification number, and locale of the John Smith entity. The first name and last name fields for the John Smith entity were remapped to the birth data, masked identification number, and locale of the Ryan James entity, and the first name and last name fields of the Ryan James entity were remapped to the birth date, identification number, and locale of the Tobias Baumer entity. First name, last name, date of birth, and locale are not considered sensitive in this example, and the data owner in this example is satisfied that the initial data set has been sufficiently anonymized by masking the identification number attributes in combination with a remapping that changed the mapping between each entity's identification number / date of birth / locale attributes and these entities' first name / last name attributes. In this example, the masking method involves both modifying entity data (the identification number data fields) and remapping some data fields to other data fields. In practice, the masking method can be as complex or simple as the data owner desires.

[0023] Once the masked dataset is created, it is then provided to the analytics provider along with a request for analytical processing. Since what is sent is the masked dataset and information about the masking method, and which data in this dataset was masked and which was not is unknown to the AP and cannot be determined from the masked dataset itself, the data owner securely conveys the data, allowing the AP to apply its analytical expertise, which, when run against the initial dataset, generates the appropriate analytical functions to perform the requested analysis. Again, on the AP side, using a computer system to implement the aspects discussed herein, the AP system interprets the masked dataset as if it were real, unmasked data. In fact, the AP may be completely unaware that it is dealing with masked data, but in some cases, the AP may wish to be alerted that the data has been masked in order to advise actors on the AP side and discourage them from misusing the data. Meanwhile, the AP system also understands the specific analytical requirements based on the customer's analytical processing request, and based on these analytical requirements, the AP generates the logic required to perform the requested analysis. The AP system then derives / generates analytical functions based on the masked data that will be used to accomplish the requested analysis. The goal is for the AP to provide analytical functions that, from the data owner's perspective, perform well on the real data (initial dataset). The goal may also be for the analytical output, despite being generated based on the masked dataset, to remain statistically consistent (measured in terms of accuracy, precision, etc.) when the generated functions are applied to the initial dataset. In other words, the AP is expected to be able to generate analytical functions that are consistent in terms of what the analytical functions produce on the masked dataset compared to what the analytical functions would have produced if the dataset provided to the AP had been the initial dataset. "Consistent" in this context does not necessarily refer to the literal same output value, but rather to the same building blocks for producing the correct analytical output given the input dataset.In this regard, efforts can be made to help ensure that AP systems generate analytical functions that do not rely on actual (unmasked) data fields.

[0024] According to another aspect, the AP system communicates generated analytical functions, generated based on the masked dataset and configured to perform desired analytical processing on the initial dataset maintained by the customer, to the customer system, and performs the analytical processing in response to the customer's analytical processing request. Thus, rather than the AP performing the analytical function on the initial dataset to which the AP does not have access, the AP passes the analytical function to the data owner. The generated function can follow a functional programming paradigm, where the generated function can be generated as logic (or functions) written in functional programming that can be invoked and executed on the initial dataset to perform the analytical processing on the initial dataset. In other words, the logic of the generated function can be passed or made available to the customer for execution by the customer on the initial dataset, allowing the customer to retain the initial dataset and thus avoid exposing the actual data. This is in contrast to a conventional situation in which the AP performs the function on the actual data and sends the data results back to the data owner.

[0025] In some examples, the AP passes the generated functions in encrypted form. The customer can then decrypt and prepare the generated functions to be executed against the customer's initial dataset. Because the analytical functions can be proprietary in nature, the AP may desire that the logic in decrypted form not be viewable by the customer. In these embodiments, a secure engine is provided by the AP as a software module and made available to the customer. The secure engine can take the generated functions and perform the decryption while keeping the decrypted logic secure so that it is not viewable by the customer. The customer can then use the engine to execute functions (which in practice may be tens or hundreds of functional programming functions) against the customer's initial dataset and obtain the results of the analytical processing.

[0026] Further details of an example sequence of events are provided to illustrate aspects described herein. In a situation where a data owner desires an analytics provider to perform analytics on data stored by the data owner, the data owner's computer system identifies a set / number of data fields of the dataset to be analyzed, a list of sensitive data fields within this set of data fields, and a list of uniquely identifying data fields within this set of data fields. By uniquely identifying, it is intended that the field hold data that uniquely identifies a unique subject entity, such as an individual. Examples of uniquely identifying data fields are an employee or user ID, an employee name, and an employee address, which can be used to uniquely identify a user / person. An example of a sensitive data field is a government-issued identification number, such as a Social Security Number (SSN) or taxpayer identification number. A sensitive data field (e.g., an SSN) can also be a uniquely identifying data field, and vice versa, although this is not always the case. For example, information about an individual is still considered sensitive information even if this information does not by itself uniquely identify the individual. The data owner's system also identifies mappings of sensitive data fields to other data fields, such as fields that map to / from uniquely identified subject entities. Finally, the data owner's system selects and implements data randomization / masking logic capable of randomizing / masking sensitive data and / or relationships between sensitive data fields and other data fields (sensitive or non-sensitive) such that the results of the analytical function delivered back to the customer remain accurate when performed on the actual data in the initial dataset. The analytical provider does not know which dataset fields were masked, let alone the logic / masking method used to mask the data; instead, that knowledge resides with the requesting customer.

[0027] It is noted that the masking method may involve only masking data in one or more fields, only remapping some fields, a combination of both, or any other modification to anonymize / mask data in the dataset. Additionally or alternatively, the masked dataset may include homomorphically encrypted versions of the data fields that maintain the referential integrity of the relationships / mappings but do not reveal personal information. In either case, the masked dataset sent to the AP may be encrypted / randomized / masked in such a way that relationships across data fields / columns are maintained and recoverable at the data owner's end when the data owner receives the generated analytical functions from the AP.

[0028] The functional programming paradigm can be used in various aspects of the exchange between data owners and APs. One embodiment provides a collection of application programming interfaces (APIs) for invocation by the data owner and a collection of APIs for invocation by the AP. On the data owner side, for example, there may be APIs for the data owner to pass masked data to the API provider system, call / select / specify / invoke randomization logic or other masking operations to randomize / mask data passed to the API provider and / or data already accessible to the API provider, or retrieve or invoke pre-generated analytical functions on a dataset. In this regard, masking / randomization can be achieved using snippets of functional programming-based logic employed for masking / randomization, which can be modified to prevent the AP from knowing the masking method applied. On the AP side, for example, there may be APIs for the AP to retrieve a dataset, invoke processing on the dataset to generate analytical functions, or provide pre-generated analytical functions / logic. As mentioned above, the generated analytical function can be generated logic written in functional programming that is passed to the data owner via an API for invocation on the API provider system or elsewhere.

[0029] FIG. 4 depicts another example environment for incorporating and using aspects described herein. In this example, a cloud provider 400 (which may encompass one or more private / public clouds) hosts and exposes an API 410 to a data owner 402 and an analysis provider 404. The API, as described above, facilitates at least communication / data exchange between the data owner 402 and the analysis provider 404. In this particular example, the cloud provider 400 is a trusted cloud (e.g., the data owner's private cloud) that hosts datasets 414, 416 and an analysis engine 420. The data owner's initial dataset 414 is hosted on the cloud 400, and the data owner invokes one or more of the APIs 400 to select and apply a masking method to the initial dataset 414, thereby transforming the initial dataset into a masked dataset 416. The data owner 402 also requests analysis processing from the analysis provider via the API or through a separate channel. The analysis provider 404 accesses the masked dataset 416 in generating the analysis function. The processing to generate the analytical functions can be performed partially or fully by the cloud provider, by the analytics provider 404, or both. In some examples, the analytics provider 404 pulls a copy of the masked dataset 416, or portions thereof, onto the analytics provider's local system and generates the analytical functions entirely on the local system. In other examples, the analytics provider 404 invokes one or more of the APIs 400 to cause the cloud provider to perform various processing to generate the analytical functions on the cloud. In further examples, the generation is a combination of the two.

[0030] In either case, the generated analytical function is "received" by the data owner 402 - in this example, the function is made available for access / use by the data owner to run against the initial dataset 414. The analytical engine 420 facilitates the execution of the analytical function against the initial dataset and therefore includes the software necessary to execute the logic of the generated analytical function. The generated function can be kept hidden from the data owner 402 and / or encrypted, as desired, so that the data owner 402 cannot see the logic of the generated function and the analytical engine 420 can execute the function to perform the analysis against the initial dataset 414.

[0031] In the particular example of FIG. 4 , a cloud service provider holds customer data (e.g., an initial dataset), and the customer invokes a masking method on the initial dataset to provide a masked dataset. The service provider then works with the analysis provider to provide the masked dataset to the AP. The service provider can, for example, send the dataset to the analysis provider or provide the analysis provider with access to the data even if the data is not shipped from the cloud provider to the analysis provider. The analysis provider generates an analysis function (on the service provider's cloud or elsewhere), and the service provider either ships the analysis function to its customer 402 or, alternatively, applies the analysis function to the initial dataset hosted by the service provider and then provides the analysis results to the data owner, the service provider's customer.

[0032] In alternative embodiments, and for security or other reasons, some aspects described in FIG. 4 as being performed on the cloud may instead be performed by the data owner 402 or the analysis provider 404. As one example, the initial dataset is not hosted on the cloud; instead, a masked dataset is generated on the data owner's system and uploaded to the cloud provider 400. Additionally or alternatively, the analysis provider 404 may download the masked dataset 416, generate analytical functions on the analysis provider's 404 system, and then upload the functions to the cloud 400. Additionally or alternatively, the generated analytical functions may be pulled to the data owner's system for execution against the initial dataset hosted by the data owner rather than on the cloud.

[0033] Yet another embodiment of the aspects described herein, which can be optionally combined with the above-mentioned aspects, is as follows: A data owner generates a fake dataset based on and modeled after a real dataset and sends the fake dataset to an analysis provider so that the analysis provider can understand the structure and schema of the fake dataset. The data owner requests that the analysis provider perform a specific analysis on the fake dataset (A1 and A2 as the analysis "requirements"). The analysis company understands A1 and A2 and, based on this, sends the data owner, via an API call, a list of data fields (e.g., f1, f2, f3, f4, and f5) from all fields (e.g., f1-f10) that need to be accessed for the analysis, as well as one or more dataset queries Q to be performed on the dataset when performing the analysis. This essentially tells the data owner which fields are of interest in light of the requested analysis, allowing the data owner to focus on which data to later pass to the analysis provider (rather than passing the entire dataset as would traditionally be done). This can also be used for training, as the data owner learns from the analysis provider which data fields are required for each requested analysis, e.g., which fields are required for A1, which fields are required for A2, etc. The data owner can learn to proactively qualify future requests to include only these fields when interacting with the AP for analysis A1 or A2.

[0034] The data owner receives information from the AP through an API call and can then, based on what the AP provides to indicate which data fields are needed, crop, mask, or randomize the actual dataset, specifically only the relevant portions, into a masked dataset so that the final conclusion of Q remains the same. In other words, the query results remain blind to the final result analysis or pattern. The data owner sends the masked dataset to the AP, which receives the masked dataset and uses it to generate analytical functions for the data owner's customer to use against the actual dataset. For example, the logic is sent to the data owner / customer to run on the data owner's / customer's own systems or in a cloud environment, such as the data owner's private cloud. The final analysis on the initial / actual dataset is then performed (for example) on the data owner's private cloud, but to protect the analytical methods employed by the analytical provider, the analytical methods and logic remain protected and not directly accessible by the customer. Meanwhile, the analytical provider can delete the masked dataset in its possession, and the data owner can delete any data related to the generated analytical functions. This retention / deletion of data or functionality may be governed by the service agreement in place between the data owner and the analytics provider.

[0035] 5 depicts an example process for secure data analysis according to aspects described herein. In some examples, the process is performed by one or more computer systems, such as those described herein, which may include one or more computer systems of a data owner, one or more cloud servers, or one or more other computer systems, or a combination thereof.

[0036] The process includes identifying (502) characteristics of an initial dataset upon which analytical processing will be performed. Example characteristics may include sensitive and / or non-sensitive data fields of the initial dataset, mappings and / or dependencies between sensitive data fields and other data fields (sensitive or non-sensitive) of the dataset, and any other characteristics of the dataset. The initial dataset is the original or "real" data, some of which may be sensitive and / or uniquely identify subject entities represented within the dataset. Some sensitive data fields may also include data fields that uniquely identify corresponding unique people in data records of the initial dataset. In some examples, the characteristics of the initial dataset may be informed, at least in part, by metadata about the initial dataset, such as the dataset's schema.

[0037] In certain examples, identifying sensitive or other data fields and mappings for the initial dataset includes the system generating and recommending to the initial dataset owner a list of candidate data fields for the initial dataset, the list of candidate data fields indicating candidate data fields to consider sensitive, candidate data fields to consider non-sensitive, or both. The list may be modified by the initial dataset owner for a final determination of which data fields are identified as sensitive data fields.

[0038] In some embodiments, identification of sensitive data fields and / or data fields expected to be used in the analytical process to be performed can be facilitated by inputting the characteristics of the dataset and an indication of the analytical requirements (informed by which analysis the data owner requests) into a machine learning model trained to predict which data fields may be used in the requested analytical process. The machine learning model outputs an indication of which data fields in the initial dataset the model predicts will be used by the analysis provider and / or which data fields are sensitive data fields that the data owner may want to mask, so that these fields can be generated and recommended to the data owner. Thus, as part of generating / recommending a list of candidate data fields, the process includes inputting the indication of the analytical process requirements into a machine learning model trained based on knowledge of the analytical requirements and the uses of the data fields corresponding to those analytical requirements, and then receiving as output from the machine learning model predictions and identification of the data fields in the initial dataset that are expected to be used in the analytical process based on the indication of the analytical process requirements. Identified sensitive data fields (those that the data owner has ultimately determined to be sensitive data fields that should be masked) can then be identified based at least in part on this output of the machine learning model.

[0039] Whether or not a model is used to inform the data fields of the initial dataset that are expected to be used when performing analytical processing on the initial dataset, based on the expected data fields to be used and the identified sensitive data fields, the method selects and applies a masking method to the initial dataset to mask the identified sensitive data fields from the initial dataset to produce a masked dataset (504). By way of example, the specific masking method can be entirely defined or specified by the data owner, or in some cases, can be selected from predetermined masking methods, perhaps customized / parameterized by the data owner, to randomize the masking that occurs according to this predetermined method. Additionally / alternatively, the specific masking method can be randomly selected from a collection of possible masking methods. Depending on the method specifications, the masking method can be as simple as a single masking strategy applied to all data fields to be masked, or as complex as applying different functions to different data fields. A single masking method can therefore encompass several complex functions applied to different data and fields of the dataset.

[0040] In one example embodiment, selecting and applying (504) can select and apply a masking method that includes modifying mappings between sensitive data fields and other data fields, which unlinks various data about a single entity from other data about the entity and optionally remaps it to data about another entity. Sensitive health data about an individual may be declassified, for example, if it is disassociated from the individual. Modifying the mappings between sensitive data fields and other data fields can include (i) removing any mappings that the sensitive data field has from at least one of the sensitive data fields to any other data fields in the initial dataset, or (ii) randomizing at least some of the mappings to randomize the relationships between the sensitive data fields and other data fields, or both.

[0041] In another example embodiment, additionally or as an alternative to modifying the mapping, selecting and applying (504) selects and applies a masking method that includes modifying / randomizing data in at least some of the identified sensitive data fields, which may include modifying unique attribute data of one or more entities in the dataset.

[0042] In yet another example embodiment of selecting and applying (504), selecting and applying (504) selects and applies a masking method that includes homomorphically encrypting data in the sensitive data field, where the homomorphically encrypting maintains referential integrity of mappings between the sensitive data field and other data fields but does not uniquely identify corresponding entities (e.g., people) in data records of the initial dataset.

[0043] Applying the masking method to the initial dataset results in a masked dataset. The process of Figure 5 then provides this masked dataset to the analysis provider (506) along with a request for analysis processing. The request for analysis processing informs the analysis provider of the analysis that the data owner is specifically requesting. The data owner can optionally alert the analysis provider that the masked dataset is a dataset of masked data that has been masked from the initial data.

[0044] The analysis provider at this point generates an analytical function that, when executed, performs the requested analysis on the input dataset. In practice, the generated analytical function may include a collection of functional programming functions. The analytical function is generated by the analysis provider based in part on the analysis provider's use of the masked dataset to configure the function to perform the requested analysis. More importantly from the data owner's perspective, the analytical function is configured to perform the analytical processing the data owner desires to be performed on the initial dataset held by the data owner when executed on the initial dataset. In response to providing the masked dataset, the data owner receives the analytical function generated by the analysis provider based on the masked dataset (508), which the data owner can invoke on the initial, unmasked dataset. The process then proceeds with invoking the generated analytical function on the initial dataset (510) to perform the analytical processing on the initial dataset. The analytical processing provides some output, typically in the form of output data, for use by the data owner. Invocation of the analytical function can be triggered automatically or by the data owner.

[0045] In certain examples, the generated analytical function includes logic written in functional programming, and invoking 510 the generated analytical function includes executing the functional programming to perform analytical processing on the initial dataset. In some cases, the generated analytical function is provided or received in an encrypted form. This may be desirable for security reasons, due to the proprietary nature of the analytical function, or both. In this case, invoking the generated analytical function includes decrypting the encrypted form of the analytical function into a decrypted form for execution. In some examples, this can be facilitated by an analytical engine configured to utilize the analytical function in encrypted or unencrypted form and the initial dataset as inputs and execute the analytical function on the dataset to output the results of the analytical processing. The analytical engine can be secure code distributed to the data owner's system, hosted on a cloud, or both, and can perform its functions in a "black box" manner, receiving inputs and delivering outputs but not exposing the analytical engine's processing and functionality in any manner observable to external software / entities, such as the data owner in the external software / entity's system. In some instances, the engine is built by AP and provided as client software for data owner systems and / or provided in a secure cloud environment for launch.

[0046] Whatever masking method is selected and applied, the selection can be made to mask information that uniquely identifies corresponding people (or other entities) in the data records of the initial dataset while both maintaining the accuracy of the generated analytic functions generated based on the masked dataset when the generated analytic functions are used for analytical processing on the initial dataset and minimizing the reverse engineerability of the masked dataset to reveal the masking method.

[0047] In some embodiments, the initial dataset is hosted on a cloud computing environment that also hosts an application programming interface (API). The selection and application of the masking method, the provision of the masked dataset and receipt of the generated analytical function, or the invocation of the generated analytical function, or a combination thereof, can be performed individually or all by the cloud computing environment based on invocation of the API by the owner of the initial dataset. Additionally or alternatively, the API can expose an interface for invocation by the analysis provider to perform aspects of their involvement, such as generating or sending / providing an analytical function to the data owner. In some examples, both the data owner and the analysis provider perform their respective steps of the process by making parameterized API calls to the cloud environment that instruct the cloud environment to perform aspects discussed herein. Control over data access / usage in the cloud environment can be achieved by controlling privileges surrounding the use of different API calls.

[0048] Although various examples are provided, variations are possible without departing from the spirit of the claimed aspects.

[0049] The processes described herein may be performed by one or more computer systems, singly or collectively. FIG. 6 depicts an example of such a computer system and associated devices for incorporating and / or using aspects described herein. The computer system may also be referred to herein as a data processing device / system, a computing device / system / node, or simply a computer. The computer system may be based on one or more of a variety of system architectures and / or instruction set architectures, such as those provided by International Business Machines Corporation (Armonk, NY, USA), Intel Corporation (Santa Clara, CA, USA), or Arm Holdings Plc (Cambridge, England, UK), by way of example.

[0050] FIG. 6 illustrates a computer system 600 in communication with an external device 612. The computer system 600 includes one or more processors 602, such as, for example, a central processing unit (CPU). The processor may include functional components used in executing instructions, such as functional components for fetching program instructions from a location such as a cache or main memory, decoding and executing the program instructions, accessing memory for instruction execution, and writing the results of executed instructions. The processor 602 also includes registers to be used by one or more of the functional components. The computer system 600 also includes memory 604, input / output (I / O) devices 608, and I / O interfaces 610, which may be coupled to the processor 602 and each other via one or more buses and / or other connections. The bus connections represent one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and without limitation, such architectures include Industry Standard Architecture (ISA), Micro Channel Architecture (MCA), Enhanced ISA (EISA), Video Electronics Standards Association (VESA) Local Bus, and Peripheral Component Interconnect (PCI).

[0051] Memory 604 may be, or may include, for example, main or system memory (e.g., random access memory), a storage device such as a hard drive, flash media, or optical media, for example, or cache memory, or a combination thereof, used during execution of program instructions. Memory 604 may include, for example, a cache, such as a shared cache, which may be coupled to a local cache of processor 602 (examples include an L1 cache, an L2 cache, etc.). Additionally, memory 604 may be, or may include, at least one computer program product having a set (e.g., at least one) of program modules, instructions, code, or the like, configured, when executed by one or more processors, to perform the functions of embodiments described herein.

[0052] The memory 604 may store an operating system 605 and other computer programs 606, such as one or more computer programs / applications that execute to implement aspects described herein. In particular, the programs / applications may include computer-readable program instructions that may be configured to perform the functions of embodiments of the aspects described herein.

[0053] Examples of I / O devices 608 include, but are not limited to, microphones, speakers, global positioning system (GPS) devices, cameras, lights, accelerometers, gyroscopes, magnetometers, sensor devices (configured to sense light, proximity, heart rate, body and / or ambient temperature, blood pressure, or skin resistance, or a combination thereof), and activity monitors. The I / O devices may be integrated into the computer system as shown, although in some embodiments, the I / O devices may be considered external devices (612) coupled to the computer system through one or more I / O interfaces 610.

[0054] Computer system 600 may communicate with one or more external devices 612 through one or more I / O interfaces 610. Examples of external devices include a keyboard, pointing device, display, or any other device that allows a user to interact with computer system 600, or a combination thereof. Other examples of external devices include any device that allows computer system 600 to communicate with one or more other computing systems or peripheral devices, such as a printer. A network interface / adapter is an example of an I / O interface that allows computer system 600 to communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, that provide communication with other computing devices or systems, storage devices, or the like. Ethernet-based interfaces (such as Wi-Fi) and Bluetooth adapters are just some examples of currently available types of network adapters used in computer systems (BLUETOOTH is a registered trademark of Bluetooth SIG, Inc., Kirkland, Washington, USA).

[0055] Communication between I / O interface 610 and external device 612 can occur across a wired and / or wireless communication link 611, such as an Ethernet-based wired or wireless connection. Examples of wireless connections include cellular, Wi-Fi, Bluetooth, proximity-based, near-field, or other types of wireless connections. More generally, communication link 611 may be any suitable wireless and / or wired communication link for communicating data.

[0056] Specific external device 612 may include one or more data storage devices that may store one or more programs, one or more computer-readable program instructions, or data, or a combination thereof. Computer system 600 may include, or be coupled to and in communication with, removable / non-removable, volatile / non-volatile computer system storage media (e.g., as a computer system external device), or both. For example, computer system 600 may include, or be coupled to, non-removable, non-volatile magnetic media (typically referred to as a "hard drive"), a magnetic disk drive for reading and writing removable, non-volatile magnetic disks (e.g., "floppy disks"), or an optical disk drive for reading and writing removable, non-volatile optical disks, such as CD-ROMs, DVD-ROMs, or other optical media, or a combination thereof.

[0057] Computer system 600 may be operational with numerous other general purpose or special purpose computing system environments or configurations. Computer system 600 may take any of a variety of forms, familiar examples of which include, but are not limited to, personal computer (PC) systems, server computer systems such as messaging servers, thin clients, thick clients, workstations, laptops, handheld devices, mobile devices / computers (such as smartphones, tablets, and wearable devices), multiprocessor systems, microprocessor-based systems, telephony devices, network equipment (such as edge appliances), virtualization devices, storage controllers, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices, and the like.

[0058] Although this disclosure includes detailed descriptions of cloud computing, it should be understood that implementations of the teachings recited herein are not limited to cloud computing environments. Rather, embodiments of the present invention are capable of being practiced in conjunction with any other type of computing environment now known or later developed.

[0059] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be quickly provisioned and released with minimal administrative effort or interaction with the service provider. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.

[0060] The characteristics are as follows:

[0061] On-demand self-service: Cloud users can unilaterally provide computing capacity, such as server time and network storage, automatically as needed, without the need for human interaction with the service provider.

[0062] Broad Network Access: Capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0063] Resource Pooling: Provider computing resources are pooled to serve multiple consumers using a multi-tenant model, with various physical and virtual resources dynamically allocated and reallocated according to demand. Location independence is meant in that consumers generally have no control or knowledge of the exact location of the resources provided, but may specify location at a higher level of abstraction (e.g., country, state, or data center).

[0064] Rapid Elasticity: Capacity is rapidly and elastically provisioned, sometimes automatically, to quickly scale out, and can be rapidly released to quickly scale in. To the consumer, the capacity available for provisioning often appears unlimited, and can be purchased in any quantity at any time.

[0065] Metered Services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource utilization can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services used.

[0066] The service model is as follows:

[0067] Software as a Service (SaaS): The consumer is provided with the ability to use a provider's applications running on a cloud infrastructure. The applications are accessible from a variety of client devices through a thin-client interface, such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or possibly individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0068] Platform as a Service (PaaS): The ability provided to a consumer is to deploy consumer-created or acquired applications, created using programming languages ​​and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but rather exercises control over the deployed applications and, in some cases, the application hosting environment configuration.

[0069] Infrastructure as a Service (IaaS): The capability provided to a customer is the provision of processing, storage, networking, and other basic computing resources on which the customer can deploy and run any software, which may include operating systems and applications. The customer does not manage or control the underlying cloud infrastructure, but rather exercises control over the operating systems, storage, deployed applications, and in some cases, limited control over selected networking components (e.g., host firewalls).

[0070] The deployment model is as follows:

[0071] Private Cloud: The cloud infrastructure is operated solely for the organization. The cloud infrastructure may be managed by the organization or a third party and may be on-premise or off-premise.

[0072] Community Cloud: Cloud infrastructure is shared by several organizations to support a unique community of shared concerns (e.g., mission, security requirements, policies, and compliance considerations). The cloud infrastructure may be managed by these organizations or a third party and may be on-premises or off-premises.

[0073] Public Cloud: Cloud infrastructure is made available to the general public or large industry groups and is owned by organizations that sell cloud services.

[0074] Hybrid Cloud: A cloud infrastructure is a composite of two or more clouds (private, community, or public) that remain unique entities but are tied together by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that allow for data and application portability.

[0075] Cloud computing environments are service-oriented, with a focus on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0076] Referring now to FIG. 7, an illustrative cloud computing environment 50 is depicted. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud users, such as, for example, a personal digital assistant (PDA) or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, or an automobile computer system 54N, or combinations thereof, may communicate. The nodes 10 may also communicate with each other. The nodes 10 may be physically or virtually grouped in one or more networks (not shown), such as private, community, public, or hybrid clouds, or combinations thereof, as described above. This enables the cloud computing environment 50 to provide infrastructure, platform, or software as a service, or combinations thereof, without the cloud user having to maintain resources on their local computing devices. The types of computing devices 54A-N shown in FIG. 7 are intended to be illustrative only, and it is understood that computing node 10 and cloud computing environment 50 can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).

[0077] Referring now to Figure 8, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 7) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 8 are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As depicted, the following layers and corresponding functions are provided:

[0078] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (reduced instruction set computer) architecture-based servers 62, servers 63, blade servers 64, storage devices 65, and network and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0079] The virtualization layer 70 provides an abstraction layer within which examples of virtual entities such as virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75 can be provided.

[0080] In one example, management layer 80 may provide the functions described below. Resource provisioning 81 dynamically procures computing and other resources utilized to perform tasks within the cloud computing environment. Metering and pricing 82 tracks costs as resources are utilized within the cloud computing environment and bills or invoices for the usage of these resources. In one example, these resources may include application software licenses. Security verifies the identity of cloud users and tasks and protects data and other resources. User portal 83 provides users and system administrators with access to the cloud computing environment. Service level management 84 allocates and manages cloud computing resources to ensure required service levels are met. Service level agreement (SLA) planning and fulfillment 85 pre-provisions and procures cloud computing resources to anticipate future requirements according to SLAs.

[0081] The workload layer 90 provides examples of functions for which a cloud computing environment may be utilized. Examples of workloads and functions that may be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and secure data analytics 96.

[0082] The present invention may be a system, method, or computer program product, or combination thereof, at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to carry out aspects of the present invention.

[0083] A computer-readable storage medium may be any tangible device capable of retaining and storing instructions for use by an instruction-execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or ridge-in-groove structures having instructions recorded thereon, and any suitable combination of the foregoing media. Computer-readable storage media as used herein should not be construed as being signals that are transitory in nature, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted through wires.

[0084] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may comprise copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within the respective computing / processing device.

[0085] Computer-readable program instructions for carrying out operations of the present invention may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine language instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object-oriented programming languages ​​such as Smalltalk®, C++, or the like, and procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer readable program instructions by utilizing state information of the computer readable program instructions to individualize the electronic circuitry to implement aspects of the present invention.

[0086] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0087] These computer-readable program instructions may be provided to a computer or other programmable data processing apparatus processor to produce a machine, the instructions of which execute on the processor of the computer or other programmable data processing apparatus processor to produce means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium such that the computer-readable storage medium comprises an article of manufacture including instructions for performing aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams, and may direct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner.

[0088] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to perform a series of operational steps on the computer, other programmable apparatus, or other device to produce a computer-executed process, the instructions executing on the computer, other programmable apparatus, or other device to perform the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.

[0089] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing specified logical functions. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may actually be performed as a single step performed concurrently, substantially concurrently, partially, or fully in a time-overlapping manner, or the blocks may sometimes be performed in the reverse order, depending on the functionality involved. It is also noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified function or operation or executes a combination of dedicated hardware and computer instructions.

[0090] In addition to the above, one or more aspects may be provided, offered, implemented, managed, serviced, etc. by a service provider offering to manage a customer environment. For example, a service provider may create, maintain, support, etc. computer code and / or computer infrastructure that implements one or more aspects for one or more customers. Alternatively, the service provider may receive payment from the customer, for example, based on a subscription or fee agreement or both. Additionally or alternatively, the service provider may receive payment from the sale of advertising content to one or more third parties.

[0091] In one aspect, an application for implementing one or more embodiments may be deployed, including, by way of example, providing a computer infrastructure operable to implement one or more embodiments.

[0092] As a further aspect, a computing infrastructure may be implemented that includes integrating computer readable code into a computing system, the code in combination with the computing system being capable of implementing one or more embodiments.

[0093] In a further aspect, a process for integrating a computing infrastructure may be provided, including integrating computer-readable code into a computer system, the computer system comprising a computer-readable medium, the computer medium comprising one or more embodiments, the code in combination with the computer system capable of implementing one or more embodiments.

[0094] Although various embodiments have been described above, these are merely examples.

[0095] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used herein, specify the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.

[0096] Corresponding structure, materials, acts, and equivalents of all means or step and functional elements in the following claims are intended to include any structure, material, or acts for performing a function, if any, in combination with other claim elements as specifically claimed. The description of one or more embodiments has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the precise form disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described to best explain various aspects and practical applications, and to enable those skilled in the art to recognize various embodiments with various modifications as suited to the particular uses intended.

Claims

1. identifying sensitive data fields of an initial dataset and mappings between the sensitive data fields and other data fields of the initial dataset, wherein an analytical process is to be performed on the initial dataset; selecting and applying a masking method to the initial dataset based on expected data fields of the initial dataset that will be used in performing the analytical processing on the initial dataset and based on the identified sensitive data fields to mask the sensitive data fields of the initial dataset to produce a masked dataset; providing the masked dataset to an analysis provider along with a request for the analytical operation, and receiving, in response to the providing, a generated analytical function generated based on the masked dataset, the generated analytical function being configured to perform the analytical operation on the initial dataset; invoking the generated analytical function on the initial data set to perform the analytical processing on the initial data set; A method executed by information processing of a computer, comprising:

2. 2. The method of claim 1, wherein the generated analytical function includes logic written in functional programming, and wherein invoking the generated analytical function includes executing the functional programming to perform the analytical processing.

3. 2. The method of claim 1, wherein the identifying the sensitive data fields and mappings includes generating and recommending a list of candidate data fields for the initial dataset to an owner of the initial dataset, the list of candidate data fields being modified by the owner of the initial dataset for final determination of the identified sensitive data fields.

4. The generating and recommending includes: inputting instructions about requirements for the analytical process into a machine learning model, wherein the machine learning model is trained based on knowledge of analytical requirements and data field usage corresponding to the analytical requirements; receiving, as an output of the machine learning model, predictions and identifications of data fields of the initial dataset based on the indication of requirements of the analytical process anticipated to be used in the analytical process, wherein the identified sensitive data fields are identified based at least in part on the output of the machine learning model; The method of claim 3, comprising:

5. The method of claim 1 , wherein said selecting and applying comprises selecting and applying a masking method that includes modifying the mapping between the sensitive data field and the other data field.

6. 6. The method of claim 5, wherein said modifying said mappings comprises at least one selected from the group consisting of: (i) removing any mappings that at least one of said sensitive data fields has from said sensitive data field to any other data fields of said initial data set; and (ii) randomizing said mappings to randomize relationships between said sensitive data fields and said other data fields.

7. 2. The method of claim 1, wherein said selecting and applying comprises selecting and applying a masking method that includes randomizing data in at least some of the identified sensitive data fields.

8. 2. The method of claim 1, wherein said selecting and applying comprises selecting and applying a masking method that includes homomorphically encrypting data in the sensitive data field, wherein said homomorphically encrypting maintains referential integrity of a mapping between the sensitive data field and the other data fields without uniquely identifying corresponding people in data records of the initial dataset.

9. 2. The method of claim 1, wherein the selected masking method is selected to mask information that uniquely identifies corresponding people in data records of the initial dataset while (i) maintaining accuracy of the generated analytical function generated based on the masked dataset when the generated analytical function is used for analytical processing on the initial dataset, and (ii) minimizing the possibility of reverse engineering of the masked dataset to reveal the masking method.

10. 2. The method of claim 1, wherein the initial dataset is hosted on a cloud computing environment that also hosts an application programming interface (API), and wherein the selecting and applying the masking method, the providing the masked dataset and the receiving the generated analytical function, and the invoking of the generated analytical function are performed by the cloud computing environment based on invocation of the API by an owner of the initial dataset.

11. The method of claim 1 , wherein the sensitive data fields comprise data fields that uniquely identify corresponding people of data records in the initial data set.

12. 2. The method of claim 1, wherein the generated analytical function is received in an encrypted form, and wherein invoking the generated analytical function comprises decrypting the analytical function in encrypted form into a decrypted form for execution.

13. 10. The method of claim 1, further comprising: identifying characteristics of the initial dataset using a schema of the initial dataset, the characteristics including mappings and dependencies between data fields of the initial dataset.

14. 10. The method of claim 1, further comprising alerting the analysis provider that the masked dataset is a dataset of masked data masked from initial data.

15. 1. A computer system comprising: Memory and a processor in communication with said memory; The computer system comprises: identifying sensitive data fields of an initial dataset and mappings between the sensitive data fields and other data fields of the initial dataset, wherein an analytical process is to be performed on the initial dataset; selecting and applying a masking method to the initial dataset based on expected data fields of the initial dataset that will be used in performing the analytical processing on the initial dataset and based on the identified sensitive data fields to mask the sensitive data fields of the initial dataset to produce a masked dataset; providing the masked dataset to an analysis provider along with a request for the analytical operation, and receiving, in response to the providing, a generated analytical function generated based on the masked dataset, the generated analytical function being configured to perform the analytical operation on the initial dataset, wherein the generated analytical function includes logic written in functional programming; invoking the generated analytical function on the initial data set to perform the analytical operation on the initial data set, wherein the invoking includes executing the functional programming to perform the analytical operation; 10. A computer system configured to perform a method comprising:

16. In the computer system, identifying sensitive data fields of an initial dataset and mappings between the sensitive data fields and other data fields of the initial dataset, wherein an analytical process is to be performed on the initial dataset; selecting and applying a masking method to the initial dataset based on expected data fields of the initial dataset that will be used in performing the analytical processing on the initial dataset and based on the identified sensitive data fields to mask the sensitive data fields of the initial dataset to produce a masked dataset; providing the masked dataset to an analysis provider along with a request for the analytical operation, and receiving, in response to the providing, a generated analytical function generated based on the masked dataset, the generated analytical function being configured to perform the analytical operation on the initial dataset, wherein the generated analytical function includes logic written in functional programming; invoking the generated analytical function on the initial data set to perform the analytical operation on the initial data set, wherein the invoking includes executing the functional programming to perform the analytical operation; A computer program for executing a method comprising:

17. 17. A computer-readable storage medium having recorded thereon the computer program of claim 16.

Citation Information

Patent Citations

  • Data kind detector and data kind detection method

    JP2009116680A

  • Analysis device system, analysis system, and analysis method

    JP2010204838A