Data Privacy Pipeline for Providing Collaborative Intelligence and Constrained Computing
By building data privacy pipelines and constrained computing technology in the trustee environment, privacy and confidentiality problems in data sharing are solved, collaborative intelligence is derived without leaking data privacy, and the development of data sharing and collaborative intelligence is promoted.
Patent Information
- Application Number
- CN202080028567.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-18
- Filing Date
- 2020-03-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-03-17
AI Technical Summary
The existing technology is difficult to effectively solve data privacy and confidentiality problems when sharing data, resulting in limited data sharing and many obstacles to controlling and accessing shared data, hindering the development and progress of collaborative intelligence.
By building data privacy pipelines in the trustee environment, leveraging constrained computing and query technologies, collaborative intelligence is derived to ensure data privacy is not leaked, while achieving configurable control and access to collaborative data.
Without disclosing the original data, effectively derive collaborative intelligence to ensure data privacy, realize efficient data sharing and control, and promote the development of collaborative intelligence.
Smart Images

Figure CN113678117B_ABST
Abstract
Description
Background Art
[0001] Business and technology are increasingly dependent on data. Many types of data can be observed, collected, derived, and analyzed to gain insights that inspire scientific and technological progress. In many cases, valuable intelligence can be derived from data sets, and useful products and services can be developed based on that intelligence. This type of intelligence can help drive industries such as banking, education, government, healthcare, manufacturing, retail, and almost any other industry. However, in many cases, the data sets owned or available to a particular data owner are incomplete or limited in some fundamental way. Information sharing is one way to bridge the data set gap, and sharing data has become an increasingly common practice. There are many benefits to sharing data. However, there are also many concerns and obstacles. Summary of the Invention
[0002] Embodiments of the present disclosure relate to techniques for deriving collaborative intelligence based on constrained computing and querying. At a high level, a data trustee can operate a trustee environment that derives collaborative intelligence subject to configurable constraints without sharing the original data. The trustee environment can include a data privacy pipeline through which data can be ingested, fused, derived, and cleaned to generate collaborative data without compromising data privacy. The collaborative data can be stored and queried to provide collaborative intelligence subject to configurable constraints. In some embodiments, the data privacy pipeline is provided as a distributed computing or cloud computing service (cloud service) implemented in the trustee environment and can be accelerated and decelerated as needed.
[0003] To implement constrained computing and querying, a constraint manager can monitor and orchestrate the use and generation of collaborative data in the constrained trustee environment. As used herein, collaborative data refers to data derived from shared input data (e.g., data from different users). The shared input data can come from any number of sources (e.g., different users) and can be processed to generate intermediate data, which in turn can be processed to generate collaborative data. The collaborative data can include a publicly shareable portion that is allowed to be shared and a restricted portion that is not allowed to be shared. Although the restricted portion of the collaborative data may not be shared, it can include an actionable portion that can be used to derive collaborative intelligence that can be shared. In some embodiments, the collaborative intelligence can be derived from publicly available data and / or restricted data, and the collaborative intelligence can be provided without disclosing the restricted data.
[0004] A user interface may be provided to enable a tenant (such as a customer, enterprise, user) to specify desired calculations and constraints for the use and access of data in a fiduciary environment, including qualified data sources and how their data may be processed or shared. Any number of various types of constraints may be implemented, including data access constraints, data processing constraints, data aggregation constraints, and data cleansing constraints, to name a few. A constraint manager may communicate with various components in the fiduciary environment to enforce the constraints. For example, requests to execute executable logic units such as commands or function calls may be issued to the constraint manager, which may grant or deny permissions. Permissions may be granted based on one or more conditions for enforcing the constraints, such as requiring a constraint executable logic unit to replace a specific executable logic unit. When constraints are applied to generate collaborative data and intermediate data, any combination of schema, constraints, and / or property metadata may be associated with the data. Thus, the constraint manager may orchestrate constraint calculations in the fiduciary environment.
[0005] In some embodiments, constraint queries may be applied to allow data consumers associated with a fiduciary environment to query collaborative data in a constrained fiduciary environment. Constraint queries may allow data consumers to access collaborative data or derive collaborative intelligence while enforcing constraints to prevent disclosure of specified data (data that has been identified for enforcing specific constraints). Constraints may be applied in a variety of ways in response to a query, including reformulating the query before execution, applying the constraints after the query is executed, constraining the qualified query for execution, applying access constraints before execution, etc. To reformulate a query according to a constraint, the query may be parsed into an execution tree including executable logic units arranged hierarchically, which implement the query when executed. The execution tree may be reformatted into a constraint execution tree by replacing executable logic units inconsistent with a specific constraint with custom executable logic units consistent with the constraint. The constraint execution tree may be translated into a language used by the fiduciary environment and forwarded for execution.
[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used in isolation to assist in determining the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The present invention is described in detail below with reference to the accompanying drawings, in which:
[0008] Figure 1 is a block diagram of an example collaborative intelligence environment according to embodiments described herein;
[0009] Figure 2 is a block diagram of an example constraint query component according to embodiments described herein;
[0010] Figure 3A are examples of queries issued, and Figure 3B are examples of corresponding execution trees according to embodiments described herein;
[0011] Figure 4A are examples of constrained execution trees, and Figure 4B are examples of corresponding queries according to embodiments described herein;
[0012] Figure 5 is a flowchart showing an example method for generating collaborative data according to embodiments described herein;
[0013] Figure 6 is a flowchart showing an example method for generating collaborative data according to embodiments described herein;
[0014] Figure 7 is a flowchart showing an example method for providing constrained computation for collaborative data in a data trustee environment according to embodiments described herein;
[0015] Figure 8 is a flowchart showing an example method for providing constrained access to collaborative data in a data trustee environment according to embodiments described herein;
[0016] Figure 9 is a flowchart showing an example method for constrained queries according to embodiments described herein;
[0017] Figure 10 is a flowchart showing an example method for constrained queries according to embodiments described herein;
[0018] Figure 11 is a block diagram of an example computing environment suitable for implementing embodiments described herein; and
[0019] Figure 12 is a block diagram of an example computing environment suitable for implementing embodiments described herein. DETAILED DESCRIPTION
[0020] Overview
[0021] There are many benefits to sharing data. For example, sharing data generally results in more complete data sets, encourages collaborative efforts, and produces better intelligence (e.g., understanding or knowledge of events or situations, or information, relationships, and facts about different types of entities). Researchers benefit from having more data available. Further, sharing can stimulate interest in research and can motivate the production of higher data quality. Generally, sharing may result in synergies and efficiencies in research and development.
[0022] However, there are also many concerns and obstacles with shared data. In fact, the ability and willingness to share data vary across different industries. Data privacy and confidentiality issues are fundamental to many industries such as healthcare and banking. In many cases, laws, regulations, and consumer demands impose restrictions on the ability to share data. Additionally, the actions of observing, collecting, deriving, and analyzing data sets are often expensive and labor-intensive activities, and many people are concerned that sharing data will result in a loss of competitive advantage. Even when there is sufficient motivation to share data, the issues of controlling and accessing shared data are often barriers to sharing. In fact, these barriers often prevent data sharing and the resulting opportunities for progress. Therefore, data sharing technologies are needed to facilitate the development of collaborative intelligence while ensuring data privacy and promoting control and access to shared data.
[0023] Accordingly, embodiments of the present disclosure relate to techniques for deriving collaborative intelligence based on constrained computing and constrained queries. At a high level, a data trustee can operate a trustee environment that is configured to derive collaborative intelligence for tenants subject to configurable constraints without exposing the underlying raw data provided by the tenants or the collaborative data shielded by the trustee environment. As used herein, collaborative data refers to data derived from shared input data (e.g., data from different users). The shared input data can come from any number of sources (e.g., different users) and can be processed to generate intermediate data, which in turn can be processed to generate collaborative data. The collaborative data can include an openly shareable portion and a restricted portion that is not allowed to be shared. Although the restricted portion of the collaborative data may not be shared, it can include actionable portions that can be used to derive collaborative intelligence that can be shared. In some embodiments, the collaborative intelligence can be derived from openly shareable data and / or restricted data, and the collaborative intelligence can be provided without exposing the restricted data. For example, the configurable constraints can programmatically manage restrictions (e.g., allowing some operations but not others) on certain underlying data (e.g., personally identifiable information, some other sensitive information, or any other specified information that is collected, stored, or used) and how the underlying data can and cannot be accessed, used, stored, or displayed (or its variations). Further, the configurable constraints can programmatically support operations on the collaborative intelligence of the accessible data (e.g., deriving aggregate statistics) without displaying the individual data entries being operated on.
[0024] By relying on trustee computations to perform data processing, tenants can derive collaborative intelligence from each other's data without compromising data privacy. To achieve this, the trustee environment can include one or more data privacy pipelines through which data can be ingested, fused, derived, and / or cleansed to generate collaborative data. The data privacy pipelines can be provided as distributed computing or cloud computing services (cloud services) implemented in the trustee environment and can be accelerated and decelerated as needed. In some embodiments, tenants that provide data to the data privacy pipelines do not have access to the pipelines. Instead, the pipeline outputs collaborative data that is subject to constraints provided by one or more tenants. Depending on the specified constraints, the collaborative data can be output from the trustee environment (e.g., because it has been cleansed according to the specified constraints) and / or can be stored in the trustee environment and shielded by the trustee environment. The shielded collaborative data can be queried to derive collaborative intelligence subject to configurable constraints (e.g., without disclosing the shielded collaborative data).
[0025] Generally, the data privacy pipelines can accept data provided by one or more tenants. Initially, the data privacy pipelines can determine whether the input data is federated data according to a contract or other tenant agreement with one or more tenants. Data determined to be federated data can be ingested, and data determined not to be federated data can be discarded. In this regard, federated data refers to any shared data that is specified for ingestion when generating collaborative data (e.g., specified in a tenant agreement with one or more tenants or otherwise identified). Ingesting data can include data from multiple sources, so the data privacy pipelines can fuse data from multiple sources according to the computations and constraints specified in the tenant agreement. For example, constrained data fusion can implement one or more constraints to combine the ingested data in any number of ways to form fused federated data, including using one or more join operations (e.g., left, right, inner, outer, anti), custom joins (e.g., via imperative scripts), data appending, normalization operations, some combination thereof, etc.
[0026] In some embodiments, a data privacy pipeline may perform constrained computations to generate derived joint data. Constrained computations may obtain data from a source (e.g., ingested data, fused joint data) and perform any number of specified computations (e.g., arithmetic operations, aggregations, generalizations, filtering, sorting, bounding). A simple example of a constrained computation is calculating the average age per city, where the calculation is only performed for a city if the underlying data set includes entries for at least five people in the city. Additionally or alternatively, the data privacy pipeline may perform data cleaning to generate collaborative data that enforces constraints on storage, access, precision, etc. For example, data cleaning may enforce constraints specified in a tenant agreement, specifying whether collaborative data (e.g., stored in a trustee environment) should be masked, whether collaborative data can be exported, whether exported collaborative data should be restricted (e.g., not export emails, credit card numbers, and parts thereof), etc. Thus, the data privacy pipeline may generate collaborative data from data provided by one or more tenants and provide agreed access to the collaborative data without sharing the underlying raw data with all tenants.
[0027] In some embodiments, to enable constrained computations and queries, the use and generation of collaborative data in a trustee environment may be monitored and orchestrated according to configurable constraints. At a high level, the constraints may be provided via a user interface to enable tenants (e.g., customers, enterprises, users) to specify the computations and constraints required for the use and access to their data in the trustee environment, including the eligible data sources and how their data may be processed or shared. Any number of various types of constraints may be implemented, including data access constraints, data processing constraints, data aggregation constraints, and data cleaning constraints.
[0028] For example, data access constraints may be specified to allow or prohibit access (e.g., to specific users, accounts, organizations). In some embodiments, the specified constraints may be general, such that the constraints apply to all potential data consumers (e.g., only allow access to the average age regardless of the data consumer). In some embodiments, the specified constraints may be applied to specified users, accounts, organizations, etc. (e.g., do not allow group A to access salary data, but allow group B to access it). Generally, tenants may specify constraints that define how tenant data may be combined with a specified data set or parts thereof, constraints that limit the data patterns read from tenant data (e.g., specify horizontal filtering to be applied to tenant data), constraints that limit the size of ingested data (e.g., specify storage limits, subsampling of tenant data, vertical filtering applied to tenant data), constraints that limit the collaborative data patterns that may be output, constraints that define the ownership of collaborative data, constraints that define whether collaborative data should be open, encrypted, or masked (e.g., stored in a trustee environment), etc.
[0029] In some embodiments, various types of data processing constraints can be specified, such as constraints that specify what operations can be performed (e.g., allowable computations and restricted computations, binary checks), constraints that limit comparison precision (e.g., for numerical data, geographical data, date and time data), constraints that limit cumulative precision (e.g., for geographical data, numerical data, date or time data), constraints that limit location boundary precision (e.g., restricting the determination of allowable geofences to specific grids, minimum geographical divisions such as blocks, counties, cities, states, or countries, etc.), and constraints on other precision and / or data processing requirements.
[0030] Additionally or alternatively, one or more data aggregation constraints can be specified, such as constraints that require a minimum aggregation amount (e.g., at least N rows or distinct field values), constraints that require some statistical distribution conditions to be valid (e.g., minimum standard deviation), constraints that define allowable aggregation functions (e.g., allow minimum, maximum, average, but not percentile), to name just a few.
[0031] In some embodiments, one or more data cleansing constraints can be specified, such as constraints that require cleansing of personally identifiable information (e.g., removing email, name, ID, credit card number), constraints that require lower precision cleansing (e.g., reducing numerical, data, and time and / or geographical precision), constraints that require cleansing of values from specific fields (which may require tracking the transformations applied in the data privacy pipeline), constraints that require custom cleansing (e.g., requiring one or more custom and / or third - party cleansing scripts), constraints that require data masking (e.g., outputting certain data such as phone numbers, credit cards, dates, but masking a portion of the digits), etc.
[0032] In addition to or as an alternative to the constraints listed above, one or more constraints can be specified to limit the amount of queries and / or data access allowed per unit of time (e.g., minutes, hours, days). Such constraints can operate by asking a slightly different set of questions within a relatively small time window to reduce the risk of brute - force attempts to reverse - engineer masked data. Generally, one or more custom constraints can be specified, such as constraints that require some specified attributes to match some specified criteria. These and other types of constraints are contemplated within the present disclosure.
[0033] In some embodiments, a constraint manager can monitor and orchestrate data flows, generation, and access according to specified constraints. For example, the constraint manager can communicate with various components in a fiduciary environment (such as a data privacy pipeline) to enforce the constraints, which can be maintained in a contract database accessible to the constraint manager. In some embodiments, components can publish requests to the constraint manager to obtain permission to execute a specific command, function call, or other executable logical unit. The constraint manager can evaluate the requests and grant or deny the permissions. In some cases, permissions can be granted according to one or more conditions corresponding to one or more constraints. By way of non-limiting example, some possible conditions that can be implemented include operations that require shifting, filtering, or shaping of data (such as the application of comparison constraints, such as only allowing merging with a specific precision), replacing one or more executable logical units (such as commands or operations) with executable logical units of one or more constraints (such as replacing an average with the average of the constraints), etc.
[0034] Generally, constraints can be checked, verified, or otherwise enforced at any time or step (such as associated with any part of the data privacy pipeline, constraint queries). Thus, the corresponding functionality for enforcing constraints can be applied at any step or multiple steps. In some embodiments, the enforcement of certain constraints can be assigned to certain parts of the data privacy pipeline (such as data access constraints applied during ingestion, processing and aggregation constraints applied during data fusion and / or constraint calculation, cleaning constraints applied during data cleaning). In another example, a specific data access constraint (such as only passing data of patients participating in at least five different studies) can be applied during data fusion. These are only intended as examples, and any suitable constraint enforcement regime can be implemented within the present disclosure.
[0035] Enforcing constraints (such as precision or aggregation constraints) on a specific executable logical unit (such as for a specified calculation, requested operation) can result in any number of scenarios. In one example, a specific executable logical unit can be completely rejected. In another example, a specific executable logical unit can be allowed, but the result is filtered (such as no value being returned for a specific row or data entry). In yet another example, a specific executable logical unit can be allowed, but the result is changed (such as reduced precision, the answer to a question being false). These and other variations can be implemented.
[0036] When constraints are applied to generate collaborative data, any combination of schema, constraints, and / or property metadata can be associated with the collaborative data, intermediate data used to obtain the collaborative data, and so on. Generally, constraints can be enforced across multiple steps and computations. Thus, in some embodiments, the applicable and / or satisfied constraints for each step can be tracked and / or associated with the data produced by a given step. Taking an aggregation constraint as an example, once the aggregation constraint has been satisfied during a particular step, subsequent steps no longer need to consider that constraint. In another example where different constraints have been specified for different data sets to be merged, the merge operation may only need to apply the more stringent constraints. Generally, when data flows through the data privacy pipeline, appropriate assignments or combinations of constraints can be applied and / or tracked. Such tracking can facilitate verifying whether a particular constraint has been applied to a particular data. Thus, when constraints are applied and data is generated, the corresponding schema, applicable or satisfied constraints, and / or property metadata indicating ownership or provenance can be associated with the data set or corresponding entry, row, field, or other data element. In some embodiments, any intermediate data used to obtain the collaborative data (e.g., ingestion data, fused federated data, derived federated data) can be deleted, and the collaborative data can be stored in a trustee environment and / or provided as output, depending on the applicable constraints.
[0037] In some embodiments, constraint queries can be applied to allow data consumers to query collaborative data in a trustee environment subject to configurable constraints. At a high level, the constraint query can operate as a search engine, allowing data consumers to access collaborative intelligence or derive collaborative intelligence from the collaborative data without exposing the underlying raw data provided by the tenant or the collaborative data shielded by the trustee environment. Constraints can be applied in any number of ways in response to a query, including reformulating the query before execution, applying constraints after query execution, constraining the eligible query for execution, applying access constraints before execution, and so on.
[0038] By way of non-limiting example, the issued query can be verified against the specified aggregation constraint by ensuring that the query contains at least one aggregation element and ensuring that the (multiple) aggregation elements are consistent with the aggregation constraint. In another example, the execution plan corresponding to the issued query can be executed, and the results can be verified against the aggregation constraint and / or (multiple) aggregation elements of the query (e.g., confirm that the results correspond to the requested number of different rows, fields, statistical distributions). In some embodiments, the constraints can enforce the corresponding elements of the query by modifying the elements based on the constraints (e.g., restricting the corresponding number of different rows, fields, statistical distributions), by executing the modifying elements before other elements of the query, some combination thereof, or other means.
[0039] Through the background, a query is typically not executable code. To execute the query, it is typically converted into an executable execution plan. In some embodiments, to enforce constraints on a received query, the query can be parsed into a corresponding execution tree that includes hierarchically arranged executable logic units that implement the query when executed. Applicable constraints can be accessed, and the executable logic units can be verified against the constraints. In some embodiments, if one or more executable logic units are not permitted, the query can be effectively reformatted by changing one or more executable logic units based on one or more constraints. More specifically, by traversing the execution tree and replacing executable logic units that are inconsistent with a particular constraint with custom executable logic units that are consistent with the constraint, the execution tree corresponding to the query can be reformatted into a constrained execution tree. Additionally or alternatively, one or more executable logic units can be added to the constrained execution tree to enforce constraints on the output (e.g., precision constraints). These are only intended as examples, and any suitable technique for generating the constrained execution tree can be implemented.
[0040] Generally, the executable logic units of the execution tree can be verified against a corresponding constraint context that includes applicable access constraints and runtime information, such as information identifying the requesting data consumer that issued the query, information identifying the applicable tenant agreement, information identifying the target collaborative data to be operated on, etc. The verification of the executable logic units can involve the verification of considerations of the constituent commands or operations, one or more constituent parameters, and / or other parts of the execution tree. The verification of the executable logic units can result in many possible outcomes. For example, an executable logic unit can be permitted (e.g., the executable logic unit can be copied into the constrained execution tree), an executable logic unit can be prohibited (e.g., the query can be prohibited entirely), or an executable logic unit can be permitted but with changes (e.g., copying the executable logic unit corresponding to the constraint into the constrained execution tree). In some embodiments, the resulting constrained execution tree is translated into a language used by a trustee environment. The resulting execution tree can be executed (e.g., by traversing the hierarchy of the executable logic units of the execution tree), and the results can be returned to the requesting data consumer.
[0041] Thus, using the implementations described herein, users can efficiently and effectively share data through data trustees that allow them to derive collaborative intelligence, while ensuring data privacy and providing configurable control and access to the shared data.
[0042] Example collaborative intelligent environment
[0043] Now refer to Figure 1, a block diagram of an example collaborative intelligent environment 100 suitable for implementing embodiments of the present invention is shown. Generally, the collaborative intelligent environment 100 is suitable for generating collaborative intelligence and, among other things, facilitating constrained computing and constrained queries. The collaborative intelligent environment 100 or a part thereof (such as the data trustee environment 110) may, but need not, be implemented in a distributed computing environment (such as the distributed computing environment 1100), as discussed below with respect to Figure 11 Any or all components of the collaborative intelligent environment 100 may be implemented as any kind of computing device or some parts thereof. For example, in an embodiment, the tenant devices 101a to 101n and the data consumer devices 103a to 103n may be computing devices such as the computing device 1200, respectively, as described below with reference to Figure 12 Further, the data trustee environment may be implemented using one or more such computing devices. In an embodiment, these devices may be any combination of personal computers (PCs), laptops, workstations, servers, mobile computing devices, PDAs, mobile phones, etc. The components of the collaborative intelligent environment 100 may communicate with each other via one or more networks, which may include, but are not limited to, one or more local area networks (LANs) and / or wide area networks (WANs). Such networking environments are common in offices, enterprise-wide computer networks, intranets, and the Internet.
[0044] The collaborative intelligent environment 100 includes a data trustee environment 110 that is capable of deriving collaborative data and / or collaborative intelligence from raw data provided by data owners or providers (such as tenants) subject to configurable constraints without sharing the raw data. Generally, any number of tenants may input their data (such as data sets 105a to 105n) into the data trustee environment 110 and specify one or more constraints (such as from one of the tenant devices 101a to 101n). The data trustee environment 110 may derive collaborative data (such as collaborative data sets 107a to 107n, masked collaborative data sets 160) based on one or more constraints. Any number of data consumers (such as one of the operating data consumer devices 103a to 103n) may issue queries against the masked collaborative data sets 160, and the data trustee environment 110 may derive collaborative intelligence from the masked collaborative data sets 160, subject to one or more constraints. In some cases, the authorized data consumers (such as those that may be defined by one or more constraints) may be the same person or entity that owns or provides the raw data (such as one or more of the data sets 105a to 105n) or owns the derived collaborative data (such as the masked collaborative data sets 160). In some cases, the authorized data consumers may be some other person or entity.
[0045] In Figure 1In the illustrated embodiment, data trustee environment 110 includes a constraint manager 115. At a high level, a tenant seeking to share data can provide one or more desired computations and constraints (which may be implemented in a contract agreement) to the constraint manager 115 via the user interface of data trustee environment 110. The user interface can enable the tenant to specify the desired computations and constraints that will govern the use of its data within data trustee environment 110, including the qualified data sources (e.g., one or more of datasets 105a through 105n) and how its data may be processed or shared. Various types of constraints can be implemented, including data access constraints, data processing constraints, data aggregation constraints, data cleansing constraints, some combination thereof, etc. The specified computations and constraints, along with other characteristics of the tenant agreement, can be stored in a contact database (not depicted) accessible to the constraint manager 115.
[0046] In Figure 1 In the illustrated embodiment, data trustee environment 110 includes a data privacy pipeline 120. At a high level, data privacy pipeline 120 can receive data from one or more specified sources (e.g., one or more of datasets 105a through 105n). The data can be ingested, fused, derived, and / or cleansed based on one or more specified computations and / or constraints to generate collaborative data (e.g., one or more of collaborative datasets 107a through 107n, masked collaborative dataset 160). Data privacy pipeline 120 can be provided as a distributed computing or cloud computing service (cloud service) implemented within trustee environment 110 and can be accelerated and decelerated as needed. In some embodiments, the tenant providing data to data privacy pipeline 120 does not have access to the pipeline. Instead, the pipeline outputs collaborative data subject to the applicable constraints. Depending on the specified constraints, the collaborative data can be output from data trustee environment 110 as one or more of collaborative datasets 107a through 107n (e.g., because it has been cleansed according to the specified constraints) and / or can be masked (e.g., stored as masked collaborative dataset 160) within data trustee environment 110. As explained in more detail below, collaborative dataset 160 can be queried to derive collaborative intelligence subject to configurable constraints.
[0047] In Figure 1In the illustrated embodiment, the data privacy pipeline 120 includes an ingestion component 125 (which generates ingestion data 130), a constraint fusion component 135 (which generates fused federated data 140), a constraint computation component 145 (which generates derived federated data 150), and a cleansing component 155 (which generates collaborative data sets 107a through 107n and 160). Initially, one or more of the data sets 105a through 105 can be provided to the data privacy pipeline 120 (e.g., via a user interface, a programming interface, or some other interface of the data trustee environment). The ingestion component 125 can determine whether the input data or a portion thereof is federated data according to a contract or other tenant agreement. For example, the input data or a portion thereof can be identified in some way, and the ingestion component 125 can communicate with the constraint manager 115 to confirm whether the identified data is federated data according to the tenant agreement represented in the contract database. Data determined to be federated data can be stored as ingestion data 130, and data determined not to be federated data can be discarded.
[0048] The ingestion data can include data from multiple sources, so the constraint fusion component 135 can fuse the ingestion data from multiple sources according to the computations and constraints specified in the tenant agreement. For example, the constraint fusion component 135 can communicate with the constraint manager 115 to obtain, verify, or request the specified fusion operations according to the tenant agreement represented in the contract database. By way of non-limiting example, the constraint fusion component 135 can implement one or more constraints to combine the ingestion data (e.g., ingestion data 130) to form the fused federated data (e.g., fused federated data 140) in any number of ways, including using one or more union operations (e.g., left, right, inner, outer, anti), custom unions (e.g., via imperative scripts), data appends, normalization operations, some combination thereof, and the like.
[0049] Generally, the constraint computation component 145 can perform constraint computations (e.g., on the ingestion data 130, the fused federated data 140) to generate the derived federated data (e.g., derived federated data 150). The constraint computations can involve any number of specified computations (e.g., arithmetic operations, aggregations, generalizations, filters, sorts, bounds). Generally, the constraint computation component 145 can communicate with the constraint manager 115 to obtain, verify, or request the specified computations according to the tenant agreement represented in the contract database. By way of simple example, many retailers may agree to disclose average sales data, so the corresponding computation may involve taking an average. A simple example of a constraint computation is calculating the average age per city, where the computation is only performed for a city if the underlying data set includes entries for at least five people in the city. These are only intended as examples, and any type of computation and / or constraint can be implemented.
[0050] In some embodiments, the cleansing component 155 may perform data cleansing (e.g., on the derived federated data 150) to generate collaborative data (e.g., one or more of the collaborative data sets 107a - 107n, the masked collaborative data set 160) in a manner that meets the constraints for storage, access, accuracy, etc. For example, the cleansing component 155 may communicate with the constraint manager 115 to obtain, verify, or request specified cleansing operations according to the tenant agreements represented in the contract database. Thus, the cleansing component 155 may enforce the constraints specified in the tenant agreement, which specify whether the collaborative data should be masked (e.g., stored as the masked collaborative data set 160 in the data trustee environment 110), whether the collaborative data can be exported (e.g., as one or more of the collaborative data sets 107a - 107n), whether the exported collaborative data should be restricted (e.g., not export emails, credit card numbers, parts thereof), some combination thereof, etc. In some embodiments, any or all of the intermediate data used to obtain the collaborative data (e.g., the ingested data, the fused federated data, the derived federated data) may be deleted, e.g., associated with the de - accelerating data privacy pipeline 120. Thus, the data privacy pipeline 120 may generate collaborative data from the data provided by one or more tenants.
[0051] As explained above, the constraint manager 115 may monitor and orchestrate the use and generation of collaborative data subject to one or more specified constraints. Additionally or alternatively, the constraint manager 115 may monitor and orchestrate access to the constrained collaborative data. Generally, the constraint manager 115 may communicate with various components in the data trustee environment 110 and / or the data privacy pipeline 120 to enforce the specified computations and / or constraints, which may be maintained in a contract database accessible to the constraint manager 115. In some embodiments, components may publish requests to the constraint manager 115 to obtain permission to perform a particular command, function call, or other executable logical unit. The constraint manager 115 may evaluate the request and grant or deny the permission. In some cases, the permission may be granted according to one or more conditions corresponding to one or more constraints. By way of non - limiting example, some possible conditions that may be implemented include operations that require shifting, filtering, or shaping of the data (e.g., application of a comparison constraint, such as only allowing a merge with a specific accuracy), replacement of one or more executable logical units (e.g., commands or operations) with one or more executable logical units of a constraint (e.g., replacing an average value with the average value of the constraint), etc.
[0052] Typically, constraints can be checked, verified, or otherwise enforced at any time or step (e.g., associated with any component of the data privacy pipeline 120, the data trustee environment 110). Accordingly, the corresponding functionality for enforcing constraints can be applied at any step or multiple steps. In some embodiments, the enforcement of certain constraints can be assigned to certain parts of the data privacy pipeline 120 (e.g., data access constraints applied by the ingestion component 125, processing and aggregation constraints applied by the constraint fusion component 135 and / or the constraint computing component 145, cleaning constraints applied by the cleaning component 155). In another example, a specific data access constraint (e.g., passing only data of patients participating in at least five different studies) can be applied by the constraint fusion component 135. These are only intended as examples, and any suitable constraint enforcement regime can be implemented within the present disclosure.
[0053] In some embodiments, the constraint manager 115 can enforce constraints (e.g., precision or aggregation constraints) on a particular executable logic unit (e.g., for a specified computation, a requested operation) by communicating, indicating, or otherwise facilitating any number of dispositions. In one example, the constraint manager 115 can completely reject a particular executable logic unit. In another example, the constraint manager 115 can allow a particular executable logic unit, but require that the result be filtered (e.g., no value is returned for a particular row or data entry). In yet another example, the constraint manager 115 can allow a particular executable logic unit, but require that the result be altered (e.g., reduced precision, the answer to a question is false). These and other variations can be implemented.
[0054] Since constraints are applied to generate collaborative data (e.g., collaborative data sets 107a to 107n, masked collaborative data set 160), any combination of schema, constraints, and / or attribute metadata can be associated with the collaborative data, intermediate data used to obtain the collaborative data (e.g., ingestion data 130, fused federated data 140, derived federated data 150), etc. Generally, constraints can be enforced across multiple steps and computations. Thus, in some embodiments, the applicable and / or satisfied constraints for each step can be traced and / or associated with the data produced by a given component of the data privacy pipeline 120. Taking the aggregation constraint as an example, once the aggregation constraint is satisfied by a particular component of the data privacy pipeline 120, downstream components no longer need to consider that constraint. In another example where different constraints have been specified for different data sets to be merged, the merge operation may only need to apply the more restrictive constraints. Generally, as data flows through the data privacy pipeline 120, appropriate constraint assignment or combination can be applied and / or traced. Such tracing can facilitate verifying whether a particular constraint has been applied to particular data. Thus, when constraints are applied and data is generated, the corresponding schema, applicable or satisfied constraints, and / or attribute metadata indicating the ownership or provenance of the data can be associated with the data set or corresponding entry, row, field, or other data element. Generally, the schema, applicable or satisfied constraints, and / or attribute metadata can be generated according to the tenant agreement represented in the contract database (e.g., via communication with the constraint manager 115). In some embodiments, any or all of the intermediate data used to obtain the collaborative data (e.g., ingestion data 130, fused federated data 140, derived federated data 150) can be deleted, and the collaborative data can be stored in the trustee environment 110 as a masked collaborative data set 160 and / or exported as one or more of the collaborative data sets 107a to 107n, depending on the applicable constraints.
[0055] In some embodiments, the data trustee environment 110 includes a constrained query component 170 that can apply constrained queries to allow a data consumer (e.g., one of the data consumer devices 103a to 103n operating the data consumer device) to query collaborative data (e.g., the masked collaborative data set 160) in the data trustee environment 110 subject to one or more specified constraints. At a high level, the constrained query component 170 can operate as a search engine, allowing the data consumer to access or derive collaborative intelligence from the masked collaborative data set 160 without disclosing the original data provided by the tenant (e.g., one or more of the data sets 105a to 105n), the intermediate data used to generate the masked collaborative data set 160 (e.g., the ingestion data 10, the fused federated data 140, the derived federated data 150), and / or the masked collaborative data set 160. Generally, the constrained query component 170 can communicate with the constraint manager 115 to obtain, validate, or request specified operations in accordance with the tenant agreements represented in the contract database. The constrained query component 170 can enforce constraints in any number of ways in response to a query, including reformulating the query before execution, applying constraints after the query is executed, constraining eligible queries for execution (e.g., only allowing a whitelist of queries), applying access constraints before execution, etc.
[0056] Turning now to Figure 2 , Figure 2 is a block diagram of an example constrained query component 200 according to embodiments described herein. The constrained query component 200 can correspond to Figure 1 the constrained query component 170. At a high level, the constrained query component 200 can operate as a search engine, enabling a data consumer to query collaborative data and derive collaborative intelligence therefrom, subject to one or more constraints specified in the corresponding tenant agreement. By background, a query is generally not executable code. To execute a query, the query is typically transformed into an execution tree, which serves as the basis for an executable execution plan. Generally, the constrained query component 200 can enforce or facilitate the enforcement of constraints by reformulating the execution tree corresponding to the received query to account for any applicable constraints before execution. In a simple example, the constraint may allow the query to compensate for data, but the results must be rounded. Thus, the query and / or its corresponding execution tree can be reformulated before execution such that any returned search results account for the applicable constraints. In Figure 1 the illustrated embodiment, the constrained query component 200 includes an access constraint component 220, a query parser 230, a constrained query formatter 240, a translation component 250, and an execution component 260. This configuration is only intended as an example, and other configurations with similar or different functionality can be implemented in accordance with the present disclosure.
[0057] At a high level, the constrained query component 200 can receive queries from a data consumer (e.g., operatingFigure 1 One of the data consumer devices 103a to 103n (the data consumer device) requests a query 210 for collaborative intelligent publishing based on collaborative data (such as Figure 1 The masked collaborative data set 160). The query 210 can take any suitable form or query language and can include operations for one or more requests for collaborative data. In some embodiments, the query 210 can specify runtime information or otherwise be associated therewith, such as information identifying the requesting data consumer that issued the query, information identifying the applicable tenant agreement, information identifying the target collaborative data to be operated on, etc.
[0058] In some embodiments, the access constraint component 220 can use the runtime information associated with the query 210 to trigger the lookup and enforcement of applicable data access constraints (e.g., via communication with Figure 1 The constraint manager 115). For example, the access constraint component 220 can verify the query 210 against a corresponding constraint context that includes the applicable data access constraints and the runtime information associated with the query 210. Generally, in scenarios where the data consumer is not authorized to access the collaborative data set, the target collaborative data within the collaborative data set (e.g., data for a specific row), or a specific type of requested collaborative intelligence to be derived, the access constraint component 220 can reject the request. In such cases, the access constraint component 220 can return a notification to the publishing data consumer to report that the query requested by the data consumer has been rejected. If the requested access is determined to be authorized and / or compliant with the applicable data access constraints, the query 210 can be passed to the query parser 230.
[0059] Generally, the query parser 230 can parse the query 210 and generate a corresponding execution tree 235. At a high level, the execution tree 235 includes executable logic units arranged hierarchically, which, when executed, implement the query 210. The executable logic units can include any suitable arrangement and combination of commands, operations, function calls, etc. The constraint query formatter 240 can access the applicable constraints (e.g., via communication with Figure 1 The constraint manager 115) and can verify the executable logic units of the execution tree 235 against the constraints. In some embodiments, if one or more executable logic units are not allowed, the query 210 can be effectively reformatted by adding, removing, and / or changing one or more executable logic units based on one or more constraints.
[0060] More specifically, by traversing the execution tree 235 and replacing executable logic units that are inconsistent with a particular constraint with custom executable logic units that are consistent with the constraint, the constraint query formatter 240 can reformat the execution tree 235 into a constraint execution tree 245. Additionally or alternatively, the constraint query formatter 240 can add or remove one or more executable logic units to enforce a constraint (e.g., a precision constraint) on the output. Generally, the constraint query formatter 240 can verify the executable logic units of the execution tree 235 against a corresponding constraint context that includes the applicable constraints and runtime information associated with the query 210. This check can involve verification considering the constituent commands or operations, one or more constituent parameters, and / or other parts of the execution tree 235, and can result in a number of possible outcomes. For example, an executable logic unit can be permitted (e.g., the executable logic unit can be copied into the constraint execution tree 245), an executable logic unit can be prohibited (e.g., the query 210 can be prohibited entirely), or an executable logic unit can be permitted but with changes (e.g., copying the executable logic unit corresponding to the constraint into the constraint execution tree 245). These are only intended as examples, and other variations are contemplated within the present disclosure.
[0061] Thus, the constraint query formatter 240 can evaluate each executable logic unit against the constraint, add or remove executable logic units, and / or replace one or more executable logic units that are inconsistent with the constraint with custom executable logic units that incorporate and / or apply the constraint. The executable logic units, the custom executable logic units, and / or the executable logic units corresponding to one or more constraints (e.g., a list of rules) can be retrieved, accessed, and / or maintained in any suitable manner (e.g., stored locally, accessed via communication with Figure 1 the constraint manager 115, some combination thereof, etc.). The mapping can be one-to-one, one-to-many, or many-to-one.
[0062] In some embodiments, the received query may be in a different query language than that used by the target collaborative data set (e.g., Figure 1 the masked collaborative data set 160). Accordingly, the translation component 250 can translate the constraint execution tree 245 from a first query language to a second query language. That is, the translation component can translate the constraint execution tree 245 into a translated constraint execution tree 255. Any suitable query language can be implemented (e.g., SQL, SparkQL, Kusto Query Language, C# Linq). In some embodiments, the constraint execution tree 245 and / or the translated constraint execution tree 255 can be executed to test for failure, and the failure may result in the rejection of a particular execution, a set of executable logic units, the entire query 210, etc.
[0063] The resulting execution tree (e.g., the constrained execution tree 245 and / or the translated constrained execution tree 255, as appropriate) can be passed to the execution component 260 for execution (e.g., the execution of the corresponding execution plan). Generally, this execution operation derives collaborative intelligence 270 from the collaborative data. In some embodiments, the collaborative intelligence 270 is returned as-is to the requesting data consumer. In some embodiments, one or more constraints can additionally or alternatively enforce the collaborative intelligence 270 before transmission to the requesting data consumer.
[0064] By way of non-limiting example, assume that, under a particular tenant agreement, many retailers have agreed to disclose sales data that includes some sensitive customer information that should not be disclosed. In this example, the tenant agreement specifies a number of constraints, including a requirement that each aggregation have at least 20 unique customers, that the aggregation must span at least 48 hours, that it cannot be aggregated by user id, that the user id cannot be exported, and that numerical results be rounded to the nearest two digits. Further assume that the tenant agreement allows the data consumer to derive the average amount spent by each customer per week in each store. Figure 3A Illustrates an example of a corresponding query 310 in Structured Query Language (SQL). This query language is only intended as an example, and any suitable query structure can be implemented.
[0065] The query 310 can be parsed and transformed into a corresponding execution tree (e.g., by Figure 2 the query parser 230). Figure 3B Illustrates a simplified representation of an example execution tree 320 corresponding to Figure 3A the query 310. Generally, in a query execution tree, each executable logical unit receives data from a previous executable logical unit and one or more parameters for transforming the data. When the execution tree 320 is executed, the data is passed from the bottom to the top along the left branch of the execution tree 320. As the data is passed, each executable logical unit applies one or more associated commands or operations. As will be appreciated by one of ordinary skill in the art, the execution tree 320 includes executable logical units arranged in a hierarchical manner that, if executed, will implement the query 310.
[0066] To account for the applicable constraints, the execution tree 320 can be transformed into Figure 4A a constrained execution tree 410 (e.g., by Figure 2 the constraint query formatter 240). Figure 3B the execution tree 320 of Figure 4AThe differences between the constraint execution tree 410 are illustrated by boxes drawn around different elements. For example, the constraint execution tree 410 includes a rounding operation 415 that implements the above constraints, where the numerical result must be rounded to the nearest two digits. In another example, the constraint execution tree 410 includes a filtering operation 425 that implements the above constraints, where the aggregation must include data for at least 20 unique customers. This configuration of the constraint execution tree 410 is only intended as an example, and any suitable configuration can be implemented. For illustrative purposes, Figure 4B An example of a corresponding query 420 corresponding to the constraint execution tree 410 is illustrated. As will be appreciated, the query 420 includes additional elements not present in the query 310, which are used to enforce the above example constraints. As will be appreciated by those of ordinary skill in the art, the constraint execution tree 410 can be executed by traversing the hierarchy of executable logic units of the tree from the bottom to the top along the left branch. Thus, the constraint execution tree 410 can be executed to derive collaborative intelligence, and the collaborative intelligence can be returned to the requesting data consumer.
[0067] Example flowchart
[0068] Referring to Figures 5 to 10 , a flowchart illustrating various methods related to the generation of collaborative intelligence is provided. These methods can be executed using the collaborative intelligence environment described herein. In an embodiment, when executed by one or more processors, one or more computer storage media having computer-executable instructions implemented thereon can cause the one or more processors to execute the methods in the autonomous upgrade system.
[0069] Now turning to Figure 5 , a flowchart of a method 500 for generating collaborative data is provided. Initially, in block 510, data from multiple input data sets provided by multiple tenants is ingested based on a tenant agreement between the multiple tenants to generate multiple ingested data sets. In block 520, the multiple ingested data sets are fused based on the tenant agreement to generate a fused combined data. In block 530, at least one constraint calculation is performed on the fused combined data based on the tenant agreement to generate derived combined data. In block 540, at least one cleaning calculation is performed on the derived combined data based on the tenant agreement to generate collaborative data. The collaborative data includes a publicly shareable portion derived from the input data sets that is allowed to be shared and a restricted portion derived from the input data sets that is not allowed to be shared.
[0070] Now turning to Figure 6, a flowchart of a method 600 for generating collaborative data is provided. Initially, in block 610, multiple data sets are fused based on at least one specified calculation or constraint to generate fused combined data. In block 620, at least one constraint calculation is performed on the fused combined data based on at least one specified calculation or constraint to generate derived combined data. In block 630, at least one cleansing calculation is performed on the derived combined data based on at least one specified calculation or constraint to generate collaborative data. The collaborative data includes a publicly shareable portion derived from the multiple data sets and a restricted portion derived from the multiple data sets that is not allowed to be shared. In block 640, access to the publicly shareable portion of the collaborative data is provided based on at least one specified calculation or constraint.
[0071] Now turn to Figure 7 , a flowchart of a method 700 for providing constraint calculations for collaborative data in a data trustee environment is provided. Initially, in block 710, a request is received to obtain permission to execute a requested executable logic unit associated with generating collaborative data from multiple input data sets provided by multiple tenants in a data trustee environment. In block 720, in response to receiving the request, at least one constraint associated with the collaborative data is accessed. In block 730, generation of the collaborative data is achieved by resolving the request based on at least one constraint. The collaborative data includes a publicly shareable portion that is allowed to be shared and a restricted portion that is not allowed to be shared. The data trustee environment is configured to provide multiple tenants with access to the publicly shareable portion of the collaborative data without disclosing the restricted portion.
[0072] Now turn to Figure 8 , a flowchart of a method 800 for providing constrained access to collaborative data in a data trustee environment is provided. Initially, in block 810, a request is received to obtain permission to execute a requested executable logic unit associated with accessing collaborative data. The collaborative data is based on multiple input data sets provided by multiple tenants. The collaborative data includes a publicly shareable portion derived from the multiple input data sets and a restricted portion derived from the multiple input data sets that is not allowed to be shared. In block 820, in response to receiving the request, at least one constraint associated with the collaborative data is accessed. In block 830, access to the collaborative data is achieved by resolving the request based on at least one constraint.
[0073] Now turn to Figure 9, a flowchart of a method 900 for constrained query is provided. Initially, at block 910, a query for generating collaborative intelligence from masked collaborative data is received from a data consumer. The masked collaborative data is generated from multiple input data sets provided by multiple tenants. The masked collaborative data includes a publicly shareable portion that is allowed to be shared and a restricted portion that is not allowed to be shared. At block 920, a request to obtain permission to execute at least one executable logic unit corresponding to the query is issued. At block 930, a response to the request is received, resolved based on one or more constraints specified in a tenant agreement among multiple tenants. At block 940, collaborative intelligence is generated from the masked collaborative data based on the query and the response to the resolved request.
[0074] Now turning to Figure 10 , a flowchart of a method 1000 for constrained query is provided. Initially, at block 1010, a query for masked collaborative data stored in a data trustee environment is received from a data consumer. The masked collaborative data is generated from multiple input data sets provided by multiple tenants. The masked collaborative data includes a publicly shareable portion derived from the multiple input data sets that is allowed to be shared and a restricted portion derived from the multiple input data sets that is not allowed to be shared. At block 1020, the query is parsed into an execution tree. At block 1030, a constrained execution tree is generated based on the execution tree and one or more constraints specified in a tenant agreement among multiple tenants. At block 1040, collaborative intelligence is generated from the masked collaborative data based on the constrained execution tree.
[0075] Example distributed computing environment
[0076] Now referring to Figure 11 , Figure 11 FIG. illustrates an example distributed computing environment 1100 in which implementations of the present disclosure may be employed. Specifically, Figure 11 FIG. shows a high-level architecture of an example cloud computing platform 1110 that may host a collaborative intelligence environment or a portion thereof (e.g., a data trustee environment). It should be understood that such and other arrangements described herein are presented only as examples. Further, as described above, many of the elements described herein may be implemented as discrete or distributed components or in combination with other components, and implemented in any suitable combination and location. Other arrangements and elements (e.g., machines, interfaces, functions, orders, and function groupings) may be used in addition to or instead of those shown.
[0077] The data center can support a distributed computing environment 1100, which includes a cloud computing platform 1110, racks 1120, and nodes 1130 (e.g., computing devices, processing units, or blades) within the racks 1120. A collaborative intelligence environment and / or a data trustee environment can be implemented using the cloud computing platform 1110 that runs cloud services across different data centers and geographical regions. The cloud computing platform 1110 can implement a structure controller 1140 component for resource allocation, deployment, upgrade, and management for the provision and management of cloud services. Generally, the cloud computing platform 1110 is used to store data or run service applications in a distributed manner. The cloud computing infrastructure 1110 in the data center can be configured to host and support the operation of endpoints of specific service applications. The cloud computing infrastructure 1110 can be a public cloud, a private cloud, or a dedicated cloud.
[0078] The nodes 1130 can be provisioned with a host 1150 (e.g., an operating system or a runtime environment) that runs a defined software stack on the nodes 1130. The nodes 1130 can also be configured to perform specialized functionality (e.g., a computing node or a storage node) within the cloud computing platform 1110. The nodes 1130 are assigned to run one or more parts of a tenant's service application. A tenant can refer to a customer who uses the resources of the cloud computing platform 1110. The service application components of the cloud computing platform 1110 that support a specific tenant can be referred to as tenant infrastructure or a lease. The terms service application, application, or service can be used interchangeably herein and generally refer to any software or part of software that runs on top of the data center or accesses the storage and computing device locations within the data center.
[0079] When more than one separate service application is supported by the nodes 1130, the nodes 1130 can be partitioned into virtual machines (e.g., virtual machines 1152 and 1154). Physical machines can also concurrently run separate service applications. The virtual machines or physical machines can be configured as personalized computing environments supported by resources 1160 (e.g., hardware resources and software resources) in the cloud computing platform 1110. It is envisioned that the resources can be configured for a specific service application. Further, each service application can be divided into functional parts such that each functional part can run on a separate virtual machine. In the cloud computing platform 1110, multiple servers can be used to run service applications and perform data storage operations in a cluster. Specifically, the servers can perform data operations independently but are exposed as a single device called a cluster. Each server in the cluster can be implemented as a node.
[0080] Client devices 1180 can be linked to service applications in the cloud computing platform 1110. For example, the client devices 1180 can be any type of computing device, which can correspond to the reference Figure 11The described computing device 1100. The client device 1180 may be configured to issue commands to the cloud computing platform 1110. In an embodiment, the client device 1180 may communicate with the service application via a virtual Internet Protocol (IP) and a load balancer or other means of directing communication requests to a specified endpoint in the cloud computing platform 1110. Components of the cloud computing platform 1110 may communicate with each other via a network (not shown), which may include, but is not limited to, one or more local area networks (LANs) and / or wide area networks (WANs).
[0081] Example operating environment
[0082] An overview of embodiments of the present invention has been briefly described. Example operating environments in which embodiments of the present invention may be implemented are described below to provide a general context for various aspects of the present invention. Specifically, initially with reference to Figure 12 , an example operating environment for implementing embodiments of the present invention is shown and is generally designated as computing device 1200. Computing device 1200 is only one example of a suitable computing environment and is not intended to imply any limitation as to the scope of use or functionality of the present invention. Computing device 1200 should also not be construed as having any relevance or requirement associated with any one or combination of the illustrated components.
[0083] The present invention may be described in the general context of computer code or machine - available instructions, including computer - executable instructions, such as program modules executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present invention may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general - purpose computers, more specialized computing devices, etc. The present invention may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network.
[0084] With reference to Figure 12 , the computing device 1200 includes a bus 1210 that directly or indirectly couples the following devices: a memory 1212, one or more processors 1214, one or more presentation components 1216, input / output ports 1218, input / output components 1220, and an illustrative power supply 1222. Bus 1210 represents one or more buses, such as an address bus, a data bus, or a combination thereof. For conceptual clarity, Figure 12 the various boxes in[] are shown with lines, and other arrangements of the described components and / or component functionality are also contemplated. For example, a presentation component, such as a display device, may be considered an I / O component. Also, a processor has a memory. We recognize this as the nature of the art and reiterateFigure 12 The figures merely illustrate example computing devices that may be used in conjunction with one or more embodiments of the present invention. There is no distinction among categories such as "workstation", "server", "laptop computer", "handheld device", etc., as all of these are contemplated to be within the scope of Figure 12 and are referred to as "computing devices".
[0085] Computing device 1200 generally includes various computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 1200 and includes volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.
[0086] Computer storage media includes volatile and non-volatile media, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by computing device 1200. Computer storage media does not itself include signals.
[0087] Communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term "modulated data signal" means a signal whose one or more characteristics are set or changed in such a manner as to encode information in the signal. By way of example and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.
[0088] Memory 1212 includes computer storage media in the form of volatile and / or non-volatile memory. The memory can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. Computing device 1200 includes one or more processors that read data from various entities such as memory 612 or I / O component 1220. The (multiple) presentation components 1216 present data indications to a user or other device. Exemplary presentation components include display devices, speakers, printing components, vibration components, etc.
[0089] The I / O port 1218 allows the computing device 1200 to be logically coupled to other devices including the I / O components 1220, some of which may be built-in. Illustrative components include microphones, joysticks, game pads, satellite dishes, scanners, printers, wireless devices, and the like.
[0090] Referring to the collaborative intelligent environment described herein, the embodiments described herein support constrained computing and / or constrained queries. The components of the collaborative intelligent environment can be integrated components, including a hardware architecture and a software framework that support constrained computing and / or constrained query functionality within the collaborative intelligent system. The hardware architecture refers to the physical components and their interrelationships, and the software framework refers to the software that provides the functionality of the hardware implementation that can be implemented on the device.
[0091] An end-to-end software-based system can operate within system components to operate computer hardware to provide system functionality. At a low level, a hardware processor executes instructions selected from a machine language (also referred to as machine code or native) instruction set for a given processor. The processor recognizes native instructions and performs corresponding low-level functions, such as those related to logic, control, and memory operations. Low-level software written in machine code can provide more complex functionality for higher-level software. As used herein, computer-executable instructions include any software, including low-level software written in machine code, high-level software such as application software, and any combination thereof. In this regard, system components can manage resources and provide services for system functionality. Any other variations and combinations are contemplated by embodiments of the present invention.
[0092] By way of example, a collaborative intelligent system can include an API library that includes specifications of routines, data structures, object classes, and variables that can support the interaction between the hardware architecture of the device and the software framework of the collaborative intelligent system. These APIs include configuration specifications of the collaborative intelligent system such that different components therein can communicate with each other in the collaborative intelligent system described herein.
[0093] The various components used herein have been identified, and it should be understood that any number of components and arrangements can be employed to achieve the desired functionality within the scope of the present disclosure. For example, for clarity of concepts, the components in the embodiments depicted in the figures are shown by lines. Other arrangements of these and other components can also be implemented. For example, although some components are depicted as single components, many of the elements described herein can be implemented as discrete or distributed components or in combination with other components, and implemented in any suitable combination and location. Some elements can be omitted entirely. Moreover, the various functions described herein as being performed by one or more entities can be performed by hardware, firmware, and / or software, as described below. For example, the various functions can be performed by a processor executing instructions stored in a memory. Thus, other arrangements and elements (e.g., machines, interfaces, functions, sequences, and function groupings) can be used in addition to or instead of those shown.
[0094] The embodiments described in the following paragraphs can be combined with one or more of the specifically described alternatives. Specifically, the claimed embodiments can incorporate references to more than one other embodiment in the alternatives. The claimed embodiments can specify further limitations to the claimed subject matter.
[0095] The subject matter of embodiments of the present invention is specifically described herein to meet statutory requirements. However, the description itself is not intended to limit the scope of this patent. On the contrary, the inventors have contemplated that the claimed subject matter can also be implemented in other ways, to include different steps or combinations of steps similar to those described herein, in combination with other existing or future technologies. Moreover, although the terms "step" and / or "block" may be used herein to denote different elements of the methods employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless the order of individual steps is explicitly described and otherwise.
[0096] For the purposes of this disclosure, the word "including" has the same broad meaning as the word "comprising", and the word "access" includes "receiving", "referencing", or "retrieving". Further, the word "communicating" has the same broad meaning as the words "receiving" or "transmitting", which are facilitated by software or hardware buses, receivers, or transmitters using the communication media described herein. Additionally, unless otherwise indicated to the contrary, words such as "a" and "an" include the plural as well as the singular. Thus, for example, the limitation of "a feature" is satisfied when there are one or more features. Moreover, the term "or" includes conjunctive, disjunctive, and both (a or b, thus including a or b as well as a and b).
[0097] For purposes of the foregoing detailed discussion, embodiments of the present invention are described with reference to a distributed computing environment; however, the distributed computing environment described herein is merely exemplary. Components can be configured to perform novel aspects of the embodiments, where the term "configured to" can mean "programmed to" perform a particular task or implement a particular abstract data type using code. Further, while embodiments of the present invention generally can refer to a collaborative intelligent environment and the diagrams described herein, it is to be understood that the techniques described can be extended to other implementation contexts.
[0098] Embodiments of the present invention have been described with respect to specific embodiments that are intended in all respects to be illustrative rather than restrictive. Alternative embodiments will become apparent to those of ordinary skill in the art to which the present invention pertains without departing from its scope.
[0099] From the foregoing it will be seen that the present invention is well adapted to carry out all the objects and purposes hereinabove set forth, together with other advantages that are obvious and inherent to the structure.
[0100] It is to be understood that certain features and subcombinations are useful and may be employed without reference to other features and subcombinations. This is contemplated by and within the scope of the claims.
Claims
1. A system for a data trustee, comprising: One or more hardware processors and a memory, the memory being configured to provide computer program instructions to the one or more hardware processors; And A data privacy pipeline defined by a tenant agreement among multiple tenants of a data trustee environment, the data privacy pipeline being configured to use the one or more hardware processors to derive collaborative data from input data sets provided by the multiple tenants by: ingesting data from the input data sets to generate multiple ingested data sets, fusing the multiple ingested data sets to generate fused combined data, performing at least one computation on the fused combined data to generate derived combined data, and performing at least one cleansing operation on the derived combined data to generate the collaborative data, wherein the data trustee environment is configured to provide the multiple tenants with access to at least the publicly available portion of the collaborative data without disclosing to any of the multiple tenants either the restricted portion of the collaborative data or the multiple ingested data sets.
2. The system according to claim 1, wherein the data privacy pipeline is a cloud service in the data trustee environment.
3. The system according to claim 1, wherein the data privacy pipeline is configured to ingest data from the input data sets based on a determination that the data is combined data designated by the tenant agreement for processing by the data privacy pipeline.
4. The system according to claim 1, wherein the data privacy pipeline is configured to fuse the ingested data sets based on at least one fusion operation specified in the tenant agreement.
5. The system according to claim 1, wherein the data privacy pipeline is configured to perform at least one constrained computation specified in the tenant agreement, wherein the at least one constrained computation includes a baseline computation and a prerequisite for performing the baseline computation.
6. The system according to claim 1, wherein the data privacy pipeline is configured to perform at least one cleansing operation specified in the tenant agreement to omit at least some data from the publicly available portion of the collaborative data.
7. The system according to claim 1, wherein the data privacy pipeline is configured to provide access to the publicly available portion of the collaborative data by deriving the publicly available portion based on restrictions specified in the tenant agreement.
8. One or more computer storage media storing computer-usable instructions, the computer-usable instructions, when used by one or more computing devices, cause the one or more computing devices to perform operations, the operations including: Storing a representation of a data privacy pipeline in a data trustee environment, the data privacy pipeline defined by a tenant agreement among multiple tenants of the data trustee environment, the data privacy pipeline being configured to derive collaborative data from input data sets provided by the multiple tenants without disclosing the input data sets; Ingesting data from the input data sets through a first operation of the data privacy pipeline to generate ingested data from two or more of the multiple tenants; Through a second operation of the data privacy pipeline, fuse the ingested data from the multiple tenants to generate fused federated data; Through a third operation of the data privacy pipeline, perform at least one computation on the fused federated data to generate derived federated data; and Through a fourth operation of the data privacy pipeline, perform at least one cleansing operation on the derived federated data to generate the collaborative data, the collaborative data including a publicly accessible portion accessible to one or more of the multiple tenants and a restricted portion inaccessible to any of the multiple tenants.
9. The one or more computer storage media according to claim 8, wherein the first operation, the second operation, the third operation, and the fourth operation are part of a cloud service of the data trustee environment.
10. The one or more computer storage media according to claim 8, wherein fusing the ingested data includes performing one or more of the following to combine multiple data sets from the input data sets from the multiple tenants: a union operation, a custom union, or data appending.
11. The one or more computer storage media according to claim 8, wherein the at least one computation includes an aggregation operation that implements one or more data aggregation constraints specified in the tenant agreement and requires a subject portion of the collaborative data to meet a condition before the aggregation operation is performed on the subject portion of the collaborative data.
12. The one or more computer storage media according to claim 8, wherein the at least one cleansing operation includes at least one of precision adjustment or data masking specified in the tenant agreement to omit at least some data from the collaborative data.
13. One or more computer storage media according to claim 8, wherein the operation further comprises: Provide access to the collaborative data by storing the collaborative data in the data trustee environment and denying access to the restricted portion of the collaborative data based on the tenant agreement.
14. A method for generating collaborative data, the method comprising: Store, in a data trustee environment, a representation of a data privacy pipeline defined by a tenant agreement between different data owners, the data privacy pipeline configured to derive collaborative data from input data sets provided by the different data owners without exposing the input data sets; Fuse, by the data privacy pipeline, ingested data from the input data sets to generate derived federated data; Perform, by the data privacy pipeline, at least one cleansing operation on the derived federated data to generate the collaborative data, the collaborative data including a publicly accessible portion accessible to one or more of the different data owners and a restricted portion inaccessible to any of the different data owners.
15. The method according to claim 14, wherein the method is a cloud service configured to start and stop running on demand.
16. The method according to claim 14, further comprising: After generating the collaborative data, delete the fused federated data and the derived federated data.
17. The method according to claim 14, wherein fusing the ingestion data from the input data set includes performing one or more of the following to combine the ingestion data from the different data owners: joint operation, customized joint, or data appending.
18. The method according to claim 14, wherein the at least one calculation includes an aggregation operation.
19. The method according to claim 14, wherein the at least one cleaning operation includes at least one precision adjustment.
20. The method according to claim 14, further comprising: Providing the one or more data owners among the different data owners with access to the publicly available portion of the collaborative data is done by: exporting the publicly available portion of the collaborative data, or storing the collaborative data in the data trustee environment and denying any data owner among the different data owners access to the restricted portion of the collaborative data based on the tenancy agreement.
Citation Information
Patent Citations
Apparatus and method for accessing data in a multi-tenant database according to a trust hierarchy
US20090282045A1
Privacy-aware query management system
US20170169253A1