Data trustee systems, computer storage media, and methods of performing a data privacy pipeline

CN116097262BActive Publication Date: 2026-08-11MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-14
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,也有许多担忧和障碍

Benefits of technology

[0007] Therefore, rights can be linked, triggered, and/or enforced within a constrained environment. Whether a right is granted for a dataset, the output of a data privacy pipeline, or an intermediate dataset generated by an intermediate step in a data privacy pipeline, the output of the right can be restricted to the constrained environment and assigned an identifier. Thus, subject to specific access constraints and policies applicable to downstream use, the owner of a protected asset can grant a beneficiary the right to use the protected asset within a constrained environment without exposing the protected asset and without requiring the grantor to explicitly authorize each downstream use. Therefore, an authorized beneficiary can construct pipelines and other computations utilizing any number of rights within a constrained environment without the rights grantor's involvement in the construction of downstream pipelines and other computations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116097262B_ABST
    Figure CN116097262B_ABST
Patent Text Reader

Abstract

The embodiments relate to techniques for implementing rights used in data privacy pipelines. When a data consumer requests to trigger a pipeline that depends on a right, the implementation mechanism can operate to verify that the data consumer's triggering of the pipeline will satisfy the right. The rules engine can access all root entities that claim the right for the pipeline, load all protocols and / or corresponding pipelines referencing one of the root entities, and search for a valid access path through the loaded protocols / pipelines. If multiple protocols and / or multiple access paths allow access to a particular root entity, various conflict rules can be configured to select which protocol and access path to use. If all root entities have valid access paths, the constrained environment can use the identified access path for each root entity to execute the requested pipeline.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Business and technology are increasingly reliant on data. Many types of data can be observed, collected, acquired, and analyzed to gain insights that inspire scientific and technological advancements. In many cases, valuable intelligence can be derived from datasets, and useful products and services can be developed based on this intelligence. This type of intelligence can help advance industries such as banking, education, government, healthcare, manufacturing, and retail, and even any other sector. However, in many cases, datasets owned or available to specific data owners are incomplete or constrained in some fundamental way. Information sharing is one way to bridge the dataset gap, and data sharing has become an increasingly common practice. Sharing data has many benefits. However, there are also many concerns and obstacles. Summary of the Invention

[0002] Embodiments of this disclosure relate to techniques for linking, triggering, and / or enforcing rights in a constrained environment. At a high level, a constrained environment (e.g., a data trustee environment or a portion thereof) may be provided with protected assets that need to be present or enforced. For example, a constrained environment can enforce a right by restricting the output of the right (e.g., an aggregated dataset) to the constrained environment, identifying the restricted right output as an intermediate dataset, and performing downstream operations consistent with the right. Thus, a beneficiary can use the granted right as input to a data privacy pipeline without requiring the grantor's approval for each specific downstream operation.

[0003] A constrained environment can provide the flexibility to grant access to specific protected assets for unspecified downstream use in another way: by allowing authorized participants in a data privacy pipeline to construct intermediate datasets generated by intermediate steps in the data privacy pipeline. More specifically, rights can be granted on the intermediate datasets, beneficiaries can construct rights, and the constrained environment can enforce those rights by performing downstream operations consistent with those rights.

[0004] Typically, a constrained environment enforces rights by satisfying applicable constraints when accessing rights and applicable policies when performing downstream operations. Data, such as intermediate datasets, can be exported from a constrained environment when a specific data consumer seeking export has sufficient ownership rights or export permissions and any applicable policies have been satisfied. Therefore, a data privacy pipeline can be constructed by linking one or more rights to a pipeline of computational steps, and the pipeline with rights can be triggered and executed within the constrained environment.

[0005] Because the output of rights from a data privacy pipeline and intermediate datasets can be linked together, downstream rights can potentially be granted to a beneficiary who is not a party to a cooperative intelligence contract governing access to upstream protected assets (e.g., input datasets, data privacy pipelines). Therefore, when a data consumer requests to trigger a pipeline or other computation that relies on any rights (e.g., a data privacy pipeline built on rights, or a data privacy pipeline whose access to the pipeline itself has been delegated through rights), an implementation mechanism can operate to verify whether the data consumer's triggering of the requested pipeline or other computation satisfies a right before the requested pipeline or other computation is triggered.

[0006] More specifically, the rules engine can access all root entities that require rights to a pipeline, load all contracts and / or corresponding pipelines referencing one of the root entities, and search for valid access paths through the loaded contracts / pipelines. To achieve this, the rules engine can proceed step by step through each pipeline, validating any constraints and policies applicable to each step. If only one contract allows access to a specific root entity through a single access path, the rules engine can specify the access path to use. If multiple contracts and / or multiple access paths allow access to a specific root entity, various conflict rules can be configured to select the contract and access path to use. If all root entities have valid access paths, the constrained environment can use the identified access path for each root entity to execute the requested pipeline or computation.

[0007] Therefore, rights can be linked, triggered, and / or enforced within a constrained environment. Whether a right is granted for a dataset, the output of a data privacy pipeline, or an intermediate dataset generated by an intermediate step in a data privacy pipeline, the output of the right can be restricted to the constrained environment and assigned an identifier. Thus, subject to specific access constraints and policies applicable to downstream use, the owner of a protected asset can grant a beneficiary the right to use the protected asset within a constrained environment without exposing the protected asset and without requiring the grantor to explicitly authorize each downstream use. Therefore, an authorized beneficiary can construct pipelines and other computations utilizing any number of rights within a constrained environment without the rights grantor's involvement in the construction of downstream pipelines and other computations.

[0008] This summary is provided to introduce a selection of concepts in a simplified form, which will be further described in the detailed description below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone as an aid in determining the scope of the claimed subject matter. Attached Figure Description

[0009] The present invention will now be described in detail with reference to the accompanying drawings, wherein:

[0010] Figure 1This is a block diagram of an example multi-participant pipeline according to the embodiments described herein;

[0011] Figure 2 This is a block diagram of an example multi-participant pipeline implemented according to the embodiments described herein;

[0012] Figure 3 This is a block diagram of an example collaborative intelligence environment based on the embodiments described herein;

[0013] Figure 4 This is a block diagram of an example chain of rights according to the embodiments described herein;

[0014] Figure 5 This is a block diagram of an example data privacy pipeline including the claims according to the embodiments described herein;

[0015] Figure 6 This is a flowchart illustrating an example method for implementing the claims according to the embodiments described herein;

[0016] Figure 7 This is a flowchart illustrating an example method for implementing the claims according to the embodiments described herein;

[0017] Figure 8 This is a flowchart illustrating another example method of implementing the claims according to the embodiments described herein;

[0018] Figure 9 This is a flowchart illustrating another example method of implementing the claims according to the embodiments described herein;

[0019] Figure 10 This is a flowchart illustrating another example method of implementing the claims according to the embodiments described herein;

[0020] Figure 11 This is a block diagram of an example computing environment suitable for implementing the embodiments described herein; and

[0021] Figure 12 This is a block diagram of an example computing environment applicable to implementing the embodiments described herein. Detailed Implementation

[0022] Overview

[0023] Sharing data offers numerous benefits. For example, it typically leads to more complete datasets, encourages collaborative efforts, and generates better intelligence (e.g., understanding or knowledge of events or environments, or information, relationships, and facts about different types of entities). Researchers benefit from having more data available. Furthermore, sharing can stimulate interest in research and incentivize the production of higher-quality data. Generally, sharing can foster synergies and efficiencies in research and development.

[0024] However, many concerns and obstacles remain regarding data sharing. In reality, the ability and willingness to share data vary across different industries. Data privacy and confidentiality are critical for many sectors, such as healthcare and banking. In many cases, laws, regulations, and consumer requirements impose restrictions on the ability to share data. Furthermore, observing, collecting, exporting, and analyzing datasets is often costly and labor-intensive, and many fear that sharing data will lead to a loss of competitive advantage. Even with sufficient motivation to share data, issues of controlling and accessing shared data often act as barriers. Indeed, these barriers often hinder data sharing and the opportunities for its advancement. Therefore, fostering collaborative intelligence while ensuring data privacy and promoting control and access to shared data is a necessity for data-sharing technologies.

[0025] Therefore, embodiments of this disclosure relate to techniques for sharing and using protected assets required to exist or be executed within a data trustee environment. At a high level, a data trustee can operate a trustee environment configured to allow tenants, subject to configurable constraints, to export collaborative intelligence without exposing the underlying raw data provided by the tenants. By relying on trustee computation to perform data processing, tenants can export collaborative data from each other's data without compromising data privacy. To achieve this, the trustee environment may include one or more data privacy pipelines that need to be executed within the data trustee environment, through which data privacy pipelines can ingest, fuse, export, and / or sanitize data to generate collaborative data. Typically, collaborative data refers to data derived from input data from any number of sources (e.g., different users or tenants). Input data can be processed through any number of computational steps in a data privacy pipeline executed within the data trustee environment to generate collaborative data. Data privacy pipelines can be considered as templates or patterns that can be triggered and initiated by authorized participants within the data trustee environment. In this way, the data privacy pipeline can generate collaborative data from data provided by one or more tenants and provide agreed-upon access to the collaborative data without sharing the underlying raw data with all tenants.

[0026] Previous applications described how participants collaborate to construct a collaborative intelligence contract for a specified data privacy pipeline configuration. Unlike a data privacy pipeline requiring multiple participants to agree to the entire pipeline, equivalent computation can be implemented, authorized, and / or triggered in other ways. For example, the owner of a protected asset (e.g., a dataset, script) or other authorized participant can establish a collaborative intelligence contract that grants another participant the right to use the protected asset within a data trustee environment, subject to any specified rights constraints and / or policies. Thus, rights can be granted for access to a designated protected asset for use by unspecified downstream entities within the data trustee environment. For example, a data contributor may wish to provide access to their data (or certain other protected assets) but may not wish to participate in the approval and implementation of complex pipelines using their data. In this case, the data contributor can grant a specific beneficiary the right to access and / or use their data, subject to specified rights constraints and / or policies. Using the granted rights, the beneficiary can use this data within their own pipeline, subject to any rights constraints and / or policies specified by the data contributor.

[0027] This application describes techniques in which the output of a right can be linked, triggered, and / or enforced within a constrained environment. At a high level, a constrained environment (e.g., a data trustee environment or a portion thereof) can be provided, containing protected assets that need to be present or enforced. For example, a constrained environment can enforce a right by restricting the output of a right (e.g., an aggregated dataset) to the constrained environment, identifying the restricted right output as an intermediate dataset, and performing downstream operations consistent with the right. Thus, a beneficiary can use the granted right as input to a data privacy pipeline without requiring the grantor to approve each specific downstream operation. A constrained environment can provide the flexibility to grant access to a specific protected asset for unspecified downstream use in another way: allowing authorized participants in the data privacy pipeline to construct intermediate datasets generated by intermediate steps of the data privacy pipeline. More specifically, rights can be granted on intermediate datasets, beneficiaries can construct rights, and the constrained environment can enforce the rights by performing downstream operations consistent with the rights. Typically, a constrained environment can enforce rights by implementing applicable constraints when accessing the right and applicable policies when performing downstream operations. Data such as intermediate datasets (whether from authorized outputs or generated by a data privacy pipeline) can be exported from a constrained environment when a specific data consumer seeking export has sufficient ownership rights or export permissions and any applicable policies have been satisfied. Therefore, a data privacy pipeline can be constructed by linking one or more rights to a pipeline of computational steps, and pipelines with rights can be triggered and enforced within a constrained environment.

[0028] More specifically, the owner of a specific protected asset (e.g., a dataset, a computational script) or certain other authorized participants can construct collaborative intelligence contracts and / or rights that grant access to the protected asset for unspecified downstream use within a constrained environment, subject to defined constraints and / or policies. For example, a contractual agreement for shared data can specify a rights output or the output of an intermediate step in a data privacy pipeline as an intermediate dataset required to exist within the constrained environment. The intermediate dataset (e.g., with a unique ID) can be identified, ownership can be assigned or otherwise determined, and downstream use of the intermediate dataset within the constrained environment can be authorized, but subject to defined policies. New authorizations can be created to manage downstream use of the intermediate dataset, regardless of whether it originates from a rights output or an intermediate step in a data privacy pipeline, but are subject to defined constraints and / or policies. Therefore, intermediate data from various sources (e.g., rights and / or data privacy pipelines) can be linked in various ways to form more flexible pipelines (e.g., multi-participant pipelines, such as multi-tenant pipelines) that need to be executed within a constrained environment without requiring grantor approval for each downstream use.

[0029] Typically, an identified intermediate dataset can be exported from a constrained environment when a specific data consumer seeking export has sufficient ownership rights and any applicable policies have been implemented. Regarding ownership rights, participants in a collaborative intelligence contract can specify explicit ownership of the intermediate dataset (whether it's an authorized output or the output generated by an intermediate step in a data privacy pipeline) via policies associated with the contract. If no specific ownership rules are specified, the intermediate dataset generated by the data privacy pipeline can be considered owned and exportable by any participant in the pipeline (e.g., the party managing the contract). Therefore, if a specific data consumer with ownership rights to the intermediate dataset requests export of the intermediate dataset (or a portion thereof) from a constrained environment, and any applicable policies can be implemented, the constrained environment can implement the policies and export the intermediate dataset.

[0030] Because the rights output from the data privacy pipeline and the intermediate dataset can be linked together, downstream rights may be granted to a beneficiary who is not a party to a collaborative intelligence contract governing access to the intermediate dataset or other upstream protected assets. Therefore, a specific data consumer may be granted the right to use an intermediate dataset for which the data consumer does not have ownership rights. If the specific data consumer seeking to derive the data does not have sufficient ownership rights, or if there are applicable policies that cannot be met, the intermediate dataset may be restricted to a constrained environment, and the derive request may be rejected. However, in some embodiments, the intermediate dataset may still be usable within a constrained environment, but subject to any applicable policies. Thus, requested downstream use of the intermediate dataset within a constrained environment, determined to be consistent with the right to govern the use of the intermediate dataset, can be authorized, and any applicable policies can be enforced by the constrained environment.

[0031] Furthermore, because the rights output from a data privacy pipeline and intermediate datasets can be linked together, downstream rights can be granted to a beneficiary who is not a party to a collaborative intelligence contract governing access to upstream protected assets (e.g., input datasets, data privacy pipelines). For example, suppose Party A grants beneficiary B the right to trigger Party A's pipeline. However, Party A's pipeline may be constructed from many other multi-party owned protected assets (e.g., input datasets). For instance, Party A's pipeline could be a pipeline with multiple participants, each contributing data. In another example, Party A's pipeline could construct protected assets (e.g., input datasets) managed by rights granted to Party A by Party C. Therefore, beneficiary B may have been entrusted with access to various protected assets managed by agreements in which B is not a party and / or managed by rights not initially granted to B.

[0032] Therefore, when a data consumer requests to trigger a pipeline or other computation that depends on any right (e.g., a data privacy pipeline that constructs a right, or a data privacy pipeline that grants access to the pipeline itself via a right delegation), the implementation mechanism can operate to verify whether the data consumer's triggering of the requested pipeline or computation satisfies the right (i.e., the constraints / policies defined by the right) before triggering the requested pipeline or computation. More specifically, the rule engine can access all root entities that require the right for the pipeline, load all contracts and / or corresponding pipelines referencing one of the root entities, and search for a valid access path through the loaded contracts / pipelines. To achieve this, the rule engine can proceed through each step of the pipeline, verifying any constraints and policies applicable to each step. If only one contract allows access to a specific root entity through a single access path, the rule engine can specify the access path to be used. If multiple contracts and / or multiple access paths allow access to a specific root entity, various conflict rules can be configured to select the contract and access path to be used. If all root entities have valid access paths, the constrained environment can use the identified access path for each root entity to execute the requested pipeline or computation.

[0033] Therefore, rights can be linked, triggered, and / or enforced within a constrained environment. Regardless of whether a right is granted on a dataset, the output of a data privacy pipeline, or an intermediate dataset generated by an intermediate step in the data privacy pipeline, the output of the right can be restricted to the constrained environment and assigned an identifier. Thus, the owner of a protected asset can grant a beneficiary the right to use the protected asset within a constrained environment, subject to constraints on access and policies applied to downstream use, without exposing the protected asset and without the grantor explicitly authorizing each downstream use. Therefore, the authorized beneficiary can construct pipelines and other computations utilizing any number of rights within a constrained environment without the right grantor's involvement in constructing downstream pipelines and other computations. When a specific data consumer requests the triggering of a pipeline or other computation dependent on any right, the enforcement mechanism can operate to verify whether the data consumer's triggering of the requested pipeline or computation will satisfy the right. If the right constraints and policies can be satisfied, the requested pipeline or computation can be authorized, triggered, and executed within the constrained environment. Therefore, the techniques described herein provide enhancements to data privacy pipelines, allowing parties to come together and decide what to compute in a more flexible manner than existing technologies.

[0034] Rights and Example Operating Environment

[0035] Unlike the premise that all participants should agree on all computational steps performed when generating collaborative data, a constrained environment can be provided that allows owners (e.g., data owners, script owners) or other authorized participants to grant access to specific protected resources for unspecified downstream use within the constrained environment, subject to specified constraints on access and policies applied to downstream use. For example, a data contributor may wish to grant access to their data (or certain other protected assets) but may not want to be involved in the approval and implementation of complex pipelines for using their data. In this case, the data contributor can grant specific beneficiaries the right to access and / or use the data contributor's data, subject to specified rights constraints and / or policies. Parameters of the rights, including applicable constraints, policies, data ownership, and / or export permissions, can be defined by an associated collaborative intelligence contract that can specify and parameterize access to any number of protected assets (e.g., datasets, computational steps, pipelines, jobs, queries, audit events, etc.). Access to specific protected assets can be customized based on specific user accounts, user groups, roles, or other bases.

[0036] As an example, Figure 1 and Figure 2 Two different ways of generating collaborative data are shown. Figure 1 An example of a data privacy pipeline with three data contributors A, B, and C is shown. In this example, the three data contributors A, B, and C collaborate to build pipeline 100, which serves as the basis for a single contractual agreement between the three data participants. Therefore, data contributors A, B, and C are all participants in pipeline 100. In this simple example, each participant contributes data, and pipeline 100 is configured to merge and perform some computations on the data, storing the results in some queryable storage.

[0037] Now, as long as certain specific constraints are met, such as aggregation constraints (e.g., applying a certain aggregation script to any part of the data A used), A is not concerned with specific computations or the probability of different possible downstream queries. In some embodiments, A may grant other participants, such as B, the right to use A's data, but subject to defined rights constraints (applied when the data is accessed) and / or rights policies (enforced against downstream use), rather than requiring A to collaborate across the entire pipeline 100, which may require A to review and sign off on the entire pipeline. Figure 2An example is shown. Example pipeline 200 may involve calculations similar to pipeline 100. However, instead of having A participate in building pipeline 200, A grants B a right 210 to use A's data (or some other protected asset required to exist or be enforced in a constrained environment) based on aggregation constraint 220, but subject to the restrictions of aggregation constraint 220. Thus, B can use right 210 to build upon A's data. When B accesses and / or uses A's data based on the right, aggregation constraint 220 can be automatically applied to generate a right output 230, which can then be used in downstream operations. Therefore, B can use A's data in collaboration with C without A's participation in the collaboration.

[0038] Typically, when a protected asset managed by a right is accessed—for example, by ingesting or otherwise identifying a protected asset within a constrained environment—a right output is generated, and right constraints can be applied and enforced by the constrained environment. For instance, input datasets can be filtered and aggregated, and resulting datasets can be used as right outputs. Right policies can define rules and restrictions on how right outputs are used within the constrained environment and / or on downstream operations within that environment. Therefore, once a right is exercised and its constraints are met, applicable policies can be applied to downstream operations. For example, a data residency policy can be applied to ensure that specified data does not leave a specific geographic area.

[0039] In some embodiments, a constrained environment can be configured to export data only if the data consumer requesting the export has sufficient ownership of the data to be exported or otherwise has export permission. Ownership and / or export permission can be defined in an associated collaborative intelligence contract. If the data consumer does not have ownership of the rights output, the constrained environment can prohibit the export of the rights output and / or can prohibit associated computations (e.g., by rejecting requests to trigger an associated pipeline used to generate and export data that the data consumer is not authorized to export). In the absence of data ownership or export permission, the beneficiary of the rights can alternatively be authorized to use the rights output within the constrained environment, subject to defined policies. For example, the beneficiary could be authorized to use the rights output in a data privacy pipeline within the constrained environment, grant another right on an intermediate dataset, or in other scenarios. Thus, in some embodiments, the constrained environment can allow the data consumer to trigger computations involving the rights output within the constrained environment even when the data consumer does not have ownership of the rights output. More specifically, in the absence of data ownership or export permission, a rights output can be treated as an intermediate dataset required to exist within a constrained environment, and the requested computation (e.g., a pipeline dependent on the rights) can be permitted when the rights constraints and policies can be satisfied. For example, while a data consumer may not have the right to export a specific dataset, it may have the right to derive and export some collaborative data (e.g., statistics) from an intermediate dataset. More generally, a rights output can remain in a constrained environment as an intermediate dataset that can be used in various ways within the constrained environment without the grantor's approval for each downstream use.

[0040] A constrained environment can enforce any applicable constraints upon access and any available policies while performing the requested downstream computation. For example, if a right to use a particular dataset is accompanied by a policy requiring any downstream operation to run a specific script (e.g., an aggregation script) at the end, the constrained environment can allow and execute the downstream operation and run the script on the output. Generally, the policy of an upstream right can be carried downstream and enforced in the downstream computation. In some embodiments, other rights can be granted on intermediate datasets that will be generated by downstream computations in the constrained environment, in which case the policy regarding the upstream right can be carried downstream and applied as constraints and / or policies regarding the downstream right.

[0041] The specified rights constraints and policies can be defined using any of the various types of constraints described herein, including data access constraints, data processing constraints, data aggregation constraints, and / or data cleansing constraints. In some embodiments, one or more data management policies can be defined and applied. Example data management policies include data residency policies (e.g., data is not allowed to leave a specific geographic area), encryption policies (e.g., output must be encrypted), data retention policies, data tagging policies (e.g., data must be tagged as public, private, or confidential), etc. Generally, some constraints and policies may be able to be satisfied and eliminated at execution time. For example, a policy may require an aggregation script to run at a certain point in time. In this case, the policy can be eliminated after the script is executed, in which case the policy no longer needs to be tracked and carried forward. In other cases, policies may affect the output from a constrained environment. For example, a policy may require some kind of aggregation constraint on any data output from a constrained environment, such as requiring a minimum amount of aggregated output data (e.g., at least N rows or distinct field values). In this case, the policy can be tracked and carried forward, and operations that satisfy the policy can be permitted, while operations that do not satisfy the policy can be rejected.

[0042] Data ownership rights and / or permissions to export data can be specified and parameterized in the associated collaborative intelligence contract. For example, a collaborative intelligence contract defining a data privacy pipeline can specify ownership of data generated at any stage of the pipeline and / or permissions to export data generated at any stage of the pipeline, including intermediate datasets generated by intermediate steps and collaborative data generated by final or output steps. In some embodiments, ownership of intermediate datasets can be specified by granting rights on the intermediate dataset and specifying ownership or export permissions using an export strategy. Thus, an authorizing participant in a data privacy pipeline can grant rights to use intermediate data generated by the data privacy pipeline for beneficiaries, subject to an export strategy that prohibits data export but allows the generation and export of some derived data (e.g., statistical data). In some embodiments, when ownership / permission to export (e.g., generated by intermediate steps) intermediate datasets is not specified, the intermediate datasets can be considered owned by the management contract and can be claimed by any participant in the data privacy pipeline or a party to the contract. In these embodiments, when ownership / export permissions are not specified, any participant in the data privacy pipeline can be authorized to export intermediate datasets, grant rights on the intermediate datasets, and / or delegate export rights to beneficiaries. In some embodiments, when rights to data owned by the contract are granted, the contract may relinquish ownership, allowing downstream users to export the obtained data, provided all constraints and policies have been satisfied, unless some other ownership or export permission rules have been specified. Therefore, any participant with ownership rights or export permissions (e.g., rights granted to themselves or third parties) may grant rights to compute and / or export intermediate datasets to a constrained environment.

[0043] To facilitate downstream use of intermediate datasets, datasets derived from intermediate steps in a data privacy pipeline or rights outputs constrained by a restricted environment, intermediate datasets can be identifiable (e.g., assigned an ID), and ownership / derivative licenses can be assigned or otherwise determined (e.g., by a management contract or restricted environment). Since intermediate datasets are identifiable and ownership / derivative rights have been defined, new rights can be granted on the intermediate datasets. These granted new rights may be more or less binding than upstream constraints or policies (e.g., for upstream rights or data privacy pipelines). However, as described in more detail below, implementation mechanisms can operate to verify whether a data consumer's triggering of a requested pipeline or computation satisfies any invoked rights, including verifying any applicable constraints and / or policies at each relevant step.

[0044] Therefore, rights allow for the creation of more flexible multi-party pipelines than existing technologies. While certain situations may be well-suited for multi-party agreement on all steps of a data privacy pipeline, rights allow participants to contribute (e.g., data) without requiring each contributor to approve every downstream use. Rights can be used in a variety of applications. For example, collaborating parties can grant rights to each other, which does not have to be symmetrical. Policies can be specified to ensure that policy-bound rights achieve the same or similar effects as the data privacy pipeline agreed upon by both parties. In another example, rights can be linked to other rights and / or any number of data privacy pipelines to form a sequence, and policies can be specified such that the resulting sequence achieves the same or similar effects as a generally agreed-upon data privacy pipeline.

[0045] Turn now Figure 3 , Figure 3 A block diagram of an example collaborative intelligence environment 300 suitable for implementing embodiments of the present invention is shown. Typically, the collaborative intelligence environment 300 is suitable for generating collaborative intelligence and, among other things, can facilitate constraint computation and constraint querying. The collaborative intelligence environment 300, or a portion thereof (e.g., data trustee environment 310), may, but does not need to, be implemented in a distributed computing environment such as distributed computing environment 1100, as described below regarding... Figure 11 The collaborative intelligence environment 300 discussed herein can be implemented as any type of computing device or parts thereof. For example, in one embodiment, each of the data consumer devices 303a to 303n can be a computing device such as computing device 1200, as referenced below. Figure 12 As described. Furthermore, the data trustee environment 310 can be implemented using one or more such computing devices. In embodiments, these devices can be any combination of personal computers (PCs), laptops, workstations, servers, mobile computing devices, PDAs, cellular phones, etc. Components of the collaborative intelligence environment 300 can communicate with each other via one or more networks, which can include, but are not limited to, one or more local area networks (LANs) and / or wide area networks (WANs). Such network environments are common in offices, enterprise-wide computer networks, intranets, and the Internet.

[0046] At a high level, the collaborative intelligence environment 300 may include a constrained environment (e.g., a data trustee environment 310 or a portion thereof, such as constrained environment 350), in which designated protected assets are required to be present or function. Typically, data trustee environment 310 and / or constrained environment 350 may be able to derive collaborative data using protected assets (e.g., data, scripts, data privacy pipelines) provided by a data owner or other authorized provider (e.g., a tenant) subject to configurable constraints, without exposing the protected assets. Any number of tenants may import or otherwise configure any number of assets (e.g., assets 305a to 305n) into data trustee environment 310 and / or constrained environment 350, specifying one or more constraints and / or policies governing their use. Data trustee environment 310 and / or constrained environment 350 may obtain collaborative data (e.g., collaborative dataset 307) based on one or more constraints and / or policies.

[0047] As used herein, a constrained environment refers to a secure, executable environment operated by a trusted party of some type, in which specified protected assets can be accessed and / or used while specified constraints and policies are enforced. A constrained environment may be able to perform constrained computations to generate collaborative data using protected assets (e.g., data, scripts, data privacy pipelines) without exposing the protected assets, intermediate datasets, or other restricted data to unauthorized parties. For example, to avoid exposing restricted data, any tenant or data consumer may be unable to access the constrained environment (e.g., the constrained environment may lack network access). Any number of data consumers (e.g., operating data consumer devices among data consumer devices 103a to 103n) may issue requests to trigger access to and / or use of pipelines or other computations that are required to exist or be performed in the constrained environment. Before triggering the requested pipeline or other computation, an enforcement mechanism may operate (e.g., via access and enforcement component 340) to verify whether the data consumer's triggering of the requested pipeline or computation satisfies a right (i.e., a constraint / policy defined by the right). If approved, the constrained environment can perform the requested pipeline or computation. In some embodiments, the constrained environment may temporarily store protected assets, initiate triggered data privacy pipelines or other applicable computations, generate any applicable intermediate datasets (e.g., intermediate dataset 380), export collaborative data upon authorization, and / or shut down any initiated pipelines or other computations (e.g., by deleting cached data, such as intermediate datasets used to reach collaborative data, temporarily stored protected assets), and / or similarly. In some embodiments, the constrained environment may be provided as part of a data trustee environment (e.g., constrained environment 350 of the data trustee environment), but this is not necessary. Although embodiments have been described herein with respect to constrained environment 350, Figure 3 The configuration described herein is not intended to be limiting, and other configurations may be implemented within the scope of this disclosure.

[0048] exist Figure 3In the illustrated embodiments, the data trustee environment 310 can receive various requests to access protected assets managed by collaborative intelligence contracts (e.g., via interface 312). For example, a data consumer (e.g., an operating data consumer device among data consumer devices 303a to 303n) can issue a request to trigger a pipeline for using the protected asset, a request to access the protected asset through a management right, or some other type of request. In some embodiments, a tenant can store assets designated as protected assets in the data trustee environment 310 (e.g., in storage allocated to the tenant). When a protected asset is designated for use by a specific collaborative intelligence contract (e.g., a data privacy pipeline or right), the digitized record associated with the contract, pipeline, and / or right may include a reference to or otherwise identify the location of the protected asset. In this way, when a request to trigger a pipeline or computation is received, any relevant protected assets can be identified (e.g., by finding protected assets associated with the invoked contract 330, pipeline 332, and / or right 334 via constraint manager 315), and access and enforcement component 340 can determine whether to access each protected asset associated with the request. In embodiments where the requested protected asset is managed by a right (e.g., one of right 334), access and enforcement component 340 can trigger right access rule engine 345 to determine whether a valid access path to the protected asset exists via one of contract 330. If access to the protected asset is granted, access and enforcement component 340 can ingest the protected asset into a secure, constrained, and / or sandboxed portion of data trustee environment 310, such as constrained environment 350.

[0049] In some embodiments, digital representations of the collaborative intelligence contract 330, data privacy pipeline 332, and / or rights 334 can be maintained in a contract database 325 accessible to the constraint manager 315. For example, contractual agreements for shared data can be stored using one or more data structures to digitally represent, reference, or otherwise identify contracts (e.g., unique identifiers), authorized participants and data consumers, access rights, protected assets, computational steps, ownership / export licenses, and so on. Therefore, the digital collaborative intelligence contract 330 can specify and / or parameterize access to any number of protected assets that are used only within the constrained environment. Example protected assets include datasets, computational steps, pipelines, jobs, queries, audit events, etc.

[0050] In some cases, a digital contract 330 may identify an associated data privacy pipeline 332, and / or vice versa. In one example, a digital contract between participants may define an associated data privacy pipeline that has been agreed upon between the participants. In this case, the digital contract and the associated data privacy pipeline may be associated with each other. In another example, a first data privacy pipeline defined by a first contract may be constructed in some way (e.g., constructing an intermediate dataset generated by an intermediate step of the data privacy pipeline, constructing data generated by the final step or output step of the data privacy pipeline) and used in a second data privacy pipeline that uses protected assets managed by a second contract. Thus, some data privacy pipelines may be based on and traceable to multiple contracts. Therefore, each digital contract that manages access to protected assets used in a multi-contract pipeline may be associated with a multi-contract pipeline. Since pipelines can be created based on many contracts, it should be understood that in some embodiments, the digital contract and the data privacy pipeline may be different entities. The digital contract 330 and / or the associated pipeline 332 can digitally represent the authorized access path through the computational steps of the pipeline (e.g., via a graph with nodes and edges), and can digitally represent the associated constraints and indications of whether a particular constraint has been satisfied (e.g., via node or edge attributes).

[0051] In some cases, a digital contract 330 can identify an associated right 334 to a protected asset. In one example, a digital contract between participants can define an associated right from a grantor, granting the beneficiary access to the protected asset (e.g., a dataset or script owned by the grantor, a data privacy pipeline where the grantor is an authorized participant, or an intermediate dataset generated by an intermediate step of a data privacy pipeline where the grantor is an authorized participant). In some cases, a right defined by a particular contract can be constructed in a way that, for example, by using a right output in a pipeline that uses a protected asset managed by another contract, and / or by using a right output from another right managed by another contract in the pipeline. Thus, a particular pipeline can be based on multiple rights and / or multiple contracts, and any of these digital entities can be associated with and traceable to each other. For example, each digital contract managing a right to a protected asset can be associated with and traceable to any pipeline that uses that right or the protected asset. In another example, each right can be associated with and traceable to each digital contract managing access to the protected asset used by the right (e.g., a right to an intermediate dataset or a complete output from a multi-contract pipeline). Since rights can be granted on protected assets managed by multiple contracts, it should be understood that in some embodiments, digital contract 330 and digital right 334 may be different entities. In some embodiments, digital right 334 may identify an associated executable constraint to be applied when accessing the protected asset. Additionally or alternatively, digital right 334 may identify an associated executable policy that will be carried along with the right output and applied during downstream use. Some policies may be satisfied and eliminated at execution time (e.g., aggregate scripts), while others may be carried and applied downstream.

[0052] Typically, the digital contract 330, the associated right 334, and / or the associated pipeline 332 may be associated with a digital representation of the authorized access path through the right 334 or the associated pipeline 332 (e.g., through a graph with nodes and edges), or with a digital representation associated with constraints, policies, and / or an indication of whether a particular constraint or policy has been satisfied (e.g., through node or edge attributes).

[0053] exist Figure 3In the illustrated embodiment, when the data trustee environment receives a request to trigger a data privacy pipeline or some other computation (e.g., via interface 312), the access and enforcement component 340 can determine whether to grant access to each protected asset associated with the request. In some embodiments, any number of tenants (e.g., tenants of the data trustee environment 310) can designate any number of protected assets for any number of data privacy pipelines and / or rights. In some cases, assets designated by tenants as protected assets may be stored in a portion of the data trustee environment 310 allocated to tenants for their use. In some cases, assets designated by tenants as protected assets may be stored in a designated location outside the data trustee environment that is accessible to the data trustee environment. In any case, once a request for access to a protected asset is received (e.g., a request to trigger a data privacy pipeline using the protected asset, access to the protected asset by rights), the access and enforcement component 340 can evaluate the access request and determine whether to grant access, as explained in more detail below. Any suitable access control technology or tool (e.g., role-based access control, access control lists, data management tools) can be used to enable access to be assessed based on any suitable identity (e.g., user identity, role, group, certain other attributes). If access is authorized, the requested asset(s) can be ingested into the secure, constrained, and / or sandboxed portion of the data trustee environment 310, such as constrained environment 350, where the asset can be used as a protected asset.

[0054] Access and enforcement component 340 can determine whether to grant access to each protected asset associated with the request in any appropriate manner. For example, an incoming request to trigger a specific data privacy pipeline may include identifiers that can be used to look up relevant parameters in the contract database 325, including any associated contracts, rights, and / or other relevant data privacy pipelines (e.g., which may be part of the triggered pipeline), any of which can be used to find the relevant protected assets that will be needed to execute the requested pipeline. The decision to grant access to each protected asset may depend on whether the requested pipeline includes any rights. For example, if a participant in a data privacy pipeline without any rights requests to trigger the pipeline, access to any protected assets used in accessing the data privacy pipeline may have already been consented to by the participant. Thus, access and enforcement component 340 can determine that a participant in a data privacy pipeline without any rights is authorized to access the associated protected assets and derive the resulting dataset (e.g., collaborative dataset 307). In embodiments where the associated protected asset is managed by a right (e.g., one of rights 334), the access and enforcement component 340 may trigger the rights access rules engine 345 to determine whether a valid access path to the protected asset exists through one of contracts 330, as described in more detail below. Additionally or alternatively, the access and enforcement component 340 may determine whether any requested output that depends on or is otherwise derived from a right (e.g., a request to generate and export collaborative data from the constrained environment 350 and / or the data trustee environment 310) is consistent with any specified data ownership rights and / or export licenses. If the access and enforcement component 340 determines that the requesting data consumer is authorized to access the associated protected asset and export the requested dataset, the access and enforcement component 340 may trigger the constrained environment 350 to perform the requested pipeline or other computations.

[0055] If access is authorized, the access and enforcement component 340 can trigger the constrained environment 350 to ingest any associated protected assets 360 and / or generate any rights output 370. For example, the constrained environment 350 can access any assets associated with a request (e.g., from a tenant's account storage) and / or can ingest them and temporarily store them (or a requested portion thereof) in the constrained environment 350 as protected assets 360. In some scenarios, any protected asset 360 can be used as a rights output. Additionally or alternatively, in embodiments where the rights designation requires some additional processing (e.g., cleansing constraints), the constrained environment 350 can apply the rights constraints to generate a rights output 370 from the ingested protected assets 360, and / or can temporarily store it in the constrained environment 350. In this way, the constrained environment can initiate triggered data privacy pipelines (e.g., data privacy pipelines 320a and 320b) or other applicable computations, generate any applicable intermediate datasets (e.g., intermediate dataset 380), export collaborative data upon authorization (e.g., collaborative dataset 307), and / or shut down any rotated pipelines or other computations (e.g., by deleting cached data, such as intermediate datasets used to reach collaborative data, temporarily stored protected assets), and so on.

[0056] Turn now Figure 4 , Figure 4 This is a block diagram of an example rights chain 400 according to an embodiment described herein. In this example, assume that a specific company A is granted different rights to use input datasets from two other companies. These two rights are managed by different contracts and define corresponding rights outputs EO1 and EO2. Therefore, company A can create a data privacy pipeline that uses these input datasets to generate some computational results. Company A can then grant others the right to trigger their pipelines and use the computational results as rights output EO3, assuming there is no conflict with the upstream rights managing EO1 and EO2.

[0057] For example, suppose company A wants to collaborate with company B. Company A will provide its data privacy pipeline via a right granted to company B (defining right output EO3), and company B will provide its dataset accessible via a separate right from company C (defining right output EO4). In this scenario, companies A and B can collaborate to create a data privacy pipeline built upon right outputs EO3 and EO4. (Assumptions and Management EO) 1-4 If there is no conflict between the upstream rights or the agreement between Company A and Company B, then Company A or Company B may grant other companies the right to use the computation results of its data privacy pipeline as a right output EO5.

[0058] Figure 4The example right chain 400 shown is available in Figure 3 An example of a computational sequence digitally represented in the contract database 325. For example, the entire chain 400 can be represented as a single main pipeline and / or a collection of pipelines. Any or all management contracts, pipelines(s), and / or underlying rights can be associated with each other or otherwise identified by the contract database 325. Thus, when a specific data consumer (e.g., Company B) requests to trigger a specific pipeline (e.g., the entire chain 400), the access and enforcement component 340 can look up the management contract, pipelines(s), and underlying rights, and the rights access rules engine 345 can verify whether Company B's triggering of the pipeline satisfies all applicable constraints and rights policies.

[0059] At a high level, when a data consumer requests a trigger that depends on a pipeline or other computation that relies on any right, the right access rules engine 345 can operate an enforcement mechanism to verify whether the data consumer's triggering of the requested pipeline or computation satisfies the right (i.e., the constraint / policy defined by the right). Figure 5 This section illustrates some considerations when operating an enforcement mechanism on an example data privacy pipeline 500 that includes rights. In this example, the data privacy pipeline 500 includes input data provided by five different entities A through E. For example, entities A through E might be different hospitals, universities, and research institutions collaborating to attempt to identify treatments for cancer. It is assumed that entity E has negotiated rights to use data from entities A through D, as stipulated in rights contract K. A-D Management, and entity E uses different rights to design data privacy pipeline 500.

[0060] Initially, in order to access and retrieve data from points A and B, contract K... A and K B Any defined rights constraints must be satisfied (e.g., data can only be accessed within a specific time window). As long as these rights constraints are satisfied, entity E can ingest data from A and B and perform some computational steps 1 (e.g., a fusion operation) on the ingested data. In this example, this is achieved by contract K. A and K B Any defined strategy is carried downstream and executed on downstream operations. Therefore, in this example, the result of computation step 1 at point 530 must satisfy any strategy from A and B (e.g., only merge with data having the minimum number of rows or distinct field values). Similarly, to access and ingest data from C and D to point 540, contract K... C and K D Any defined rights constraints must be satisfied, and in order to perform calculation step 2 to generate the calculation result at point 550, contract K... C With K DAll defined strategies must be satisfied. To perform computation step 3 to generate the result at point 560, contract K... A K B K C and K D Any defined strategy must be satisfied, and so on. Therefore, to verify whether the triggering of the data consumer's request pipeline will satisfy any combined rights, an implementation mechanism (e.g.) is required. Figure 3 The rights access rules engine (345) can proceed through the requested computation steps and verify any applicable constraints and policies on each step.

[0061] Suppose that entity A grants entity E a right to a data privacy pipeline that itself relies on upstream rights granted to entity A to use data from some other entity FH. Now, pipeline 500 may need to ingest data from entities F through H, and if entity E wants to trigger pipeline 500, any constraints and policies defined by the use of data from entities F through H managed by these upstream rights may need to be looked up and evaluated. More generally, when a data user requests to trigger a pipeline or other computation, all root entities that must be accessed through rights can be identified. As used herein, a root entity can be defined as an input asset (e.g., an input dataset or script) that is provided to the requested pipeline or other computation without any preprocessing of the input asset. When a request to trigger a data privacy pipeline is received, the pipeline can be part of a larger main pipeline that includes any number of component pipelines and rights. Root entities can be considered as inputs to the main pipeline (e.g., datasets, scripts, etc.).

[0062] In this way and back to Figure 3 When a request to trigger a specific pipeline is received, the rights access rule engine 345 can access all root entities requiring the rights for the pipeline, load all contracts and / or corresponding pipelines referencing one of the root entities, and search for a valid access path through the loaded contracts / pipelines. To achieve this, the rights access rule engine 345 can proceed through each step of the pipeline, validating any applicable constraints and policies at each step. If only one contract allows access to a specific root entity through a single access path, the rights access rule engine 345 can specify the access path to use. If multiple contracts and / or multiple access paths allow access to a specific root entity, the rights access rule engine 345 can apply configured and / or predefined conflict rules to select which contract and access path to specify for use. If all root entities have valid access paths, the rights access rule engine 345 can authorize the request and trigger the constrained environment 350 to execute the requested pipeline using the identified access path for each root entity.

[0063] More specifically, now turn to Figure 6 , Figure 6 A flowchart illustrating an example method 600 for implementing the rights is provided. This method can be performed using the collaborative intelligence environment described herein. For example, in some embodiments, one or more computer storage media having computer-executable instructions thereon can cause one or more processors to execute the method in a collaborative intelligence environment when executed by one or more processors. In some embodiments, method 600 can be performed by… Figure 3 Access and enforcement components 340 and / or rights access rules engine 345 are executed.

[0064] Initially, in box 610, a request to trigger the data privacy pipeline is received. For example, a data consumer (e.g., Figure 3 One of the data consumer devices (303a to 303n) in operation can issue a request to trigger a specific data privacy pipeline via interface 312 of the data trustee environment 310. Typically, the software associated with the data trustee environment 310 (e.g., functionality associated with access and enforcement component 340 and / or rights access rules engine 345) can evaluate whether to execute the request by performing the following steps. To support such a configuration, in some embodiments, the request can be routed to the appropriate component for evaluation.

[0065] In box 620, retrieve the data privacy pipeline that was triggered by the request (e.g., from...). Figure 3 In the contract database (325), and in box 630, all root entities requiring access to the root entities of a data privacy pipeline are identified. Typically, a particular data privacy pipeline can run on any number of protected assets (e.g., datasets, scripts, etc.). In some scenarios, the protected assets used by the data privacy pipeline may have been agreed upon by all participants in the pipeline, making the protected assets not requiring rights. In other scenarios, the data privacy pipeline can run on one or more rights outputs that depend on access to certain upstream protected assets (e.g., rights or another data privacy pipeline) that require access to and / or generate the protected assets. In yet another scenario, the data consumer requesting to trigger the data privacy pipeline may not have participated in establishing the pipeline, but one of the participants granted the data consumer the right to trigger the pipeline. Any or all of these scenarios may apply. Therefore, various techniques can be applied to identify all root entities requiring rights for a data privacy pipeline.

[0066] For example, all root entities claiming rights for a triggered data privacy pipeline can be identified by obtaining a digital representation of the triggered data privacy pipeline, the associated contract, the associated pipeline, and / or the associated rights, and by identifying any root entities managed by one of the associated rights. In some embodiments, these root entities can be identified (e.g., before or after receiving a request to trigger the pipeline), and the identified root entities can be identified through the pipeline, the associated contract, the associated pipeline, and / or the associated rights (e.g., in...). Figure 3 The digital representation of the data privacy pipeline (325) can be associated with and retrieved. Thus, upon receiving a request to trigger a specific pipeline, the digital representation of the pipeline and / or any associated contracts, pipelines, and / or rights can be obtained, and the associated root entity can be located. As a non-limiting example, the root entity claiming the rights can be identified by a list, attributes, metadata, and / or other indications associated with the triggered data privacy pipeline (or associated contracts, pipelines, and / or rights). Additionally or alternatively, in some scenarios, the access path can be traced upstream from the data privacy pipeline until it reaches the root entity claiming the rights. For example, if the triggered data privacy pipeline operates on the output generated by an upstream data privacy pipeline, the access path can be traced backward from the output via the upstream data privacy pipeline. The access path can be traced through any number of upstream pipelines to identify all root entities claiming the rights.

[0067] After identifying all root entities that request access to the root entity, in box 640, all contracts managing access to the identified root entities can be loaded. Typically, any number of contracts can grant access to a specific root entity. For example, a specific script or dataset can be made available to any number of collaborating parties according to the terms managed and / or implemented by any number of corresponding contracts. For example, a specific contract can define or identify data privacy pipelines, rights, or certain other access paths referencing a specific root entity. Thus, the digital set of contracts (e.g., Figure 3 Contracts in the contract database 325 (contracts 330) can be searched to identify contracts that reference and / or grant access to the root entity.

[0068] In box 650, contracts unrelated to the requesting data consumer are filtered. For example, some contracts define, store, or otherwise identify access constraints based on the identity or account associated with the user triggering access. Access constraints for such contracts can be used to identify and filter contracts that do not grant access to the requesting data consumer, for example, based on the data consumer's identity (e.g., because the user is not on a specified whitelist, is part of an authorized account, etc.). Contracts that do not have access constraints based on the identity or account associated with the requesting user can be considered to have passed the access query.

[0069] In box 660, a method is executed for each identified root entity. More specifically, the methods shown in boxes 662 through 674 can be applied to each identified root entity. Taking a particular root entity as an example, in box 662, each loaded contract can be searched to find a valid access path to the root entity through the contract.

[0070] First, potential access paths and associated constraints and policies can be identified from the loaded contract. In some cases, a loaded contract may grant access to a root entity based on an associated data privacy pipeline agreed upon between the contract's participants, subject to defined constraints. In other cases, a loaded contract may grant access to a root entity based on associated rights, subject to defined constraints and / or policies. In either case, the contract may identify an access path and any associated constraints and / or policies through one or more computational steps (e.g., steps in a data privacy pipeline or processing policy). In some cases, some computational steps not on a direct route from the root entity may also need to be performed in order to access a specific root entity using a specific access path. Therefore, a potential access path through a contract may include computational steps that need to be performed to access a specific root entity. In any case, a potential access path through a contract, along with associated constraints and policies, can be digitally represented, associated with the contract, and looked up. In some cases, this can be considered as identifying a main pipeline that includes all computational steps required to trigger the requested pipeline (e.g., upstream computational steps via rights merging).

[0071] When a potential access path through a loaded contract includes a data privacy pipeline without rights (e.g., an upstream data privacy pipeline where access to and use of all protected assets has been approved by all participants in the upstream pipeline), each computational step of the pipeline can be evaluated to determine if the associated constraints are satisfied. Initially, computational steps can be evaluated without execution to determine if the associated constraints will be satisfied. For example, an associated constraint might restrict access to a specific time frame (e.g., only on Tuesdays or until a fixed end date), in which case the constraint can be evaluated based on the context of the request (e.g., the time associated with the request). If the constraints on the computational step will be satisfied, subsequent computational steps on the potential access path can be evaluated. If any potential access path is determined to be invalid because one of the constraints along that path will not be satisfied, that access path can be considered invalid. Similarly, if the only access path through a loaded contract is determined to be invalid, the contract may be considered invalid when used to fulfill the request. However, in some cases, a particular pipeline or contract may have multiple potential access paths (e.g., alternative paths through common computational steps), in which case each potential access path can be evaluated. If there is no permitted access path through the contract, the contract may be discarded. If a permitted access path exists through the contract, the contract can be marked as a candidate contract.

[0072] Typically, in the case of rights and / or associated management contracts, rights may have associated right constraints and / or policies. When a potential access path through a loaded contract includes the right to manage access to protected assets, any associated right constraints applicable when accessing the protected assets can be evaluated, and any policies that must be carried downstream can be evaluated in association with downstream computation steps. Associated right constraints can be evaluated to determine whether they are satisfied without generating a right output. If the right constraints will be satisfied, the applicable policies can be carried downstream and evaluated in association with downstream computation steps (e.g., in a downstream data privacy pipeline, such as a request-triggered pipeline), without performing steps to determine whether the applicable policies will be satisfied.

[0073] Typically, forward-carrying strategies for a particular computational step can be evaluated in a manner similar to the applicable constraints on that computational step. In some cases, applicable constraints or strategies can be evaluated across multiple computational steps (e.g., data is only available for a specific time). Furthermore, some applicable constraints may overlap, in which case only the more stringent constraint may need to be evaluated (e.g., when one constraint restricts usage to once a week while another restricts it to once a month, the more stringent monthly constraint can be identified and evaluated). Typically, applicable and / or satisfied constraints / strategies for each step can be tracked through the proposed computational steps. In some cases, applicable constraints or strategies may be revocable (e.g., the requirement to anonymize data before merging with another dataset, the requirement to run a specific script somewhere). In these scenarios, when it is determined that a particular computational step will satisfy a revocable constraint or strategy, in some cases, that constraint or strategy may no longer need to be evaluated in further downstream computational steps. Typically, computational steps via a contract through a potential access path can be evaluated sequentially, progressing to subsequent computational steps as the determination that applicable constraints / strategies for previous computational steps can be satisfied is made. If a particular computational step will not satisfy the applicable constraints / policies, in some cases, an alternative access path can be used to revisit the previous computational step.

[0074] In some cases where a potential access path uses an intermediate dataset generated by an intermediate step in a data privacy pipeline, subsequent computational steps in the data privacy pipeline downstream of the intermediate step (e.g., computational steps that are not part of the potential access path and are not required to trigger the requested pipeline) may not need to be evaluated. Such computational steps can be flagged to indicate that these steps should not be performed if the access path is considered valid and designated for use.

[0075] In some cases, a particular constraint or policy may only be validated when one or more computations are performed along a potential access path (e.g., during runtime). For example, a policy might allow access to a dataset provided that the dataset is merged into data with at least some entries (e.g., one million rows). To validate the policy, computational steps along the potential access path can be performed conditionally (e.g., for limited purposes of evaluating compliance), and the policy can be evaluated based on the results. Since performing computational steps conditionally can be computationally expensive, in some embodiments, a constraint or policy that is fully validated only during runtime may be considered conditionally satisfied and evaluated only at a later time, such as when it is determined that the potential access path is otherwise valid. Further, in such a scenario, a notification may be provided (e.g., provided to the requesting data consumer), and / or an interruption may be provided before continuation to request (e.g., from the requesting data consumer) confirmation.

[0076] Accordingly, the identified computational steps involving potential access paths through the loaded contract, including the rights, can be evaluated. If no permitted access path exists through the rights, the corresponding contract may be discarded. If a permitted access path exists through the rights, the corresponding contract can be marked as a candidate contract.

[0077] Return now Figure 6 The search process in box 662 may result in zero or more candidate contracts and zero or more potential access paths for a specific root entity. In box 664, a determination is made as to whether at least one valid access path has been identified. If not, in box 668, the request to trigger the data privacy pipeline is rejected. If at least one valid access path has been identified, in box 670, a determination is made as to whether a single access path has been identified. If a single contract with a single valid access path to the root entity is identified, in box 674, the access path and its management contract can be specified for accessing the root entity. On the other hand, if more than one valid access path has been identified, in box 672, conflict rules can be applied to identify a single contract and a single access path. That is, if a single contract is identified as having multiple valid access paths to the root entity, or if multiple candidate contracts are identified as having multiple valid access paths to the root entity, any number and type of conflict rules can be applied to select a single management contract and / or a single valid access path. For example, default and / or preferred options can be selected, interruptions can be provided to request the selection of specific options (e.g., from a requesting data consumer), and / or other metrics can be applied to select options (e.g., having the cheapest computational cost, the least computation, the minimum amount of data to be generated, the most or least restrictive options, options that allow access to smaller or larger categories of resources, and / or others).

[0078] A repeating box 660 can be used for each identified root entity. If a valid access path is identified for each root entity, the identified access path can be used to trigger the requested data privacy pipeline. For example, Figure 3 The access enforcement component 340 and / or the rights access rules engine 345 can trigger the constrained environment 350 to execute the requested data privacy pipeline using the identified access path.

[0079] Turn now Figure 7 A flowchart illustrating an example method 700 implementing the claims according to embodiments described herein is provided. This method can be performed using the collaborative intelligence environment described herein. For example, in some embodiments, one or more computer storage media containing computer-executable instructions thereon, when executed by one or more processors, can cause one or more processors to execute the method in a collaborative intelligence environment.

[0080] First, in Box 710, a digital representation of the stored rights is provided. This rights are granted by the grantor to the beneficiary for the use of the grantor's protected assets as the output of a right required to exist or be enforced within the data trustee environment, but subject to the rights constraints or rights policies stipulated by the rights. In Box 720, a configuration of the data privacy pipeline is received from the beneficiary. The received configuration includes specifications for: (i) the input dataset to the data privacy pipeline, (ii) one or more computational steps of the data privacy pipeline, (iii) the stipulated use of the rights output, and (iv) the output dataset of the data privacy pipeline. In Box 730, a configuration of the data privacy pipeline is deployed within the data trustee environment without exposing the protected assets, either by imposing rights constraints when accessing the protected assets according to the rights, or by implementing rights policies for downstream use of the rights output by the data privacy pipeline.

[0081] Turn now Figure 8 A flowchart illustrating an example method 800 implementing the claims according to embodiments described herein is provided. This method can be performed using the collaborative intelligence environment described herein. For example, in some embodiments, one or more computer storage media having computer-executable instructions thereon, when executed by one or more processors, can cause one or more processors to execute the method in a collaborative intelligence environment.

[0082] First, in box 810, the configuration of the right is received. The right is granted by the grantor to the beneficiary to use the output of an intermediate step of the first data privacy pipeline as an intermediate dataset required to exist in the data trustee environment, but subject to the right constraints or right policies specified by the right. In box 820, the configuration of the second data privacy pipeline is received from the beneficiary. The received configuration includes specifications for: (i) the input dataset to the second data privacy pipeline, (ii) one or more computational steps of the second data privacy pipeline, (iii) the specified use of the intermediate dataset according to the right, and (iv) the output dataset of the second data privacy pipeline. In box 830, the first and second data privacy pipelines are deployed in the data trustee environment without exposing the intermediate dataset by imposing right constraints on the intermediate dataset generated by the first data privacy pipeline or by imposing right policies on the downstream use of the intermediate dataset by the second data privacy pipeline.

[0083] Turn now Figure 9 A flowchart illustrating an example method 900 implementing the claims according to embodiments described herein is provided. This method can be performed using the collaborative intelligence environment described herein. For example, in some embodiments, one or more computer storage media having computer-executable instructions thereon can cause one or more processors to execute the method in a collaborative intelligence environment when executed by one or more processors.

[0084] First, in box 910, the configuration of the rights is received. The rights are granted by the grantor to the beneficiary to use the grantor's dataset in a constrained environment, but are subject to the rights constraints or rights policies specified by the rights. In box 920, the configuration of the data privacy pipeline is received from the beneficiary. The received configuration includes specifications for: (i) the input dataset to the data privacy pipeline, (ii) one or more computational steps of the data privacy pipeline, (iii) the specified use of the grantor's dataset, and (iv) the output dataset of the data privacy pipeline. At box 930, the data privacy pipeline is executed in the constrained environment without exposing the grantor's dataset by imposing rights constraints when accessing the grantor's dataset or by imposing rights policies for operations downstream of the grantor's dataset in the data privacy pipeline. At box 940, the output dataset is exported from the constrained environment.

[0085] Turn now Figure 10 A flowchart illustrating an example method 1000 implementing the claims according to embodiments described herein is provided. This method can be performed using the collaborative intelligence environment described herein. For example, in some embodiments, one or more computer storage media having computer-executable instructions thereon can cause one or more processors to execute the method in a collaborative intelligence environment when executed by one or more processors.

[0086] First, in box 1010, a request is received from the data consumer. This request is used to trigger a data privacy pipeline required to execute in a constrained environment inaccessible to the data consumer, and to derive data generated by the data privacy pipeline from the constrained environment. At box 1020, it is determined that executing the data privacy pipeline within the constrained environment will satisfy the associated rights of the root entity used for using the data privacy pipeline within the constrained environment. These associated rights include the constraints on access to the root entity within the constrained environment and the policies for downstream computations derived from the root entity within the constrained environment. At box 1030, it is determined that the data consumer has permission to derive data from the constrained environment. At box 1040, the data privacy pipeline is triggered to execute within the constrained environment using the root entity in accordance with the associated rights, without exposing the root entity.

[0087] Example Collaborative Intelligence Environment

[0088] Some embodiments relate to techniques for obtaining collaborative intelligence based on constraint computation and constraint querying. At a high level, a data trustee can operate a trustee environment configured to obtain collaborative intelligence for tenants under configurable constraints without exposing the underlying raw data provided by the tenants or collaborative data protected by the trustee environment. As used herein, collaborative data refers to data obtained from shared input data (e.g., data from different users). Shared input data can come from any number of sources (e.g., different users) and can be processed to generate intermediate data, which itself can be processed to generate collaborative data. Collaborative data can include publicly shareable portions and restricted portions that are not allowed to be shared. Although the restricted portions of collaborative data may not be shared, they can include operable portions that can be used to obtain shareable collaborative intelligence. In some embodiments, collaborative intelligence can be obtained from exposed data and / or restricted data, and collaborative intelligence can be provided without exposing the restricted data. For example, configurable constraints can programmatically manage restrictions on certain underlying data (e.g., personally identifiable information, certain other sensitive information, or any other specified information collected, stored, or used) (e.g., allowing some operations but disallowing others), and under what circumstances the underlying data can and cannot be accessed, used, stored, or displayed (or variations thereof). Furthermore, configurable constraints can programmatically support collaborative intelligence operations on accessible data (e.g., obtaining aggregate statistics) without displaying the individual data entries being operated on.

[0089] By relying on a trustee to perform data processing, tenants can gain collaborative intelligence from each other's data without compromising data privacy. To achieve this, the trustee environment can include one or more data privacy pipelines through which data can be acquired, fused, obtained, and / or sanitized to generate collaborative data. Data privacy pipelines can be provided as distributed computing or cloud computing services (cloud services) implemented within the trustee environment and can be enabled or disabled as needed. In some embodiments, tenants providing data to a data privacy pipeline cannot access the pipeline. Instead, the pipeline outputs collaborative data, but subject to constraints provided by one or more tenants. Depending on the specified constraints, collaborative data can be output from the trustee environment (e.g., because it has been sanitized according to the specified constraints) and / or can be stored in and protected by the trustee environment. Protected collaborative data can be queried to obtain collaborative intelligence, but subject to configurable constraints (e.g., not exposing protected collaborative data).

[0090] Typically, a data privacy pipeline can accept data from one or more tenants. Initially, the data privacy pipeline can determine whether the input data is federated data based on contracts with one or more tenants or other tenant agreements. Data determined to be federated data can be ingested, and information determined not to be federated data can be discarded. In this regard, federated data refers to any shared data specified for ingestion when generating collaborative data (e.g., specified in a tenant agreement with one or more tenants or otherwise identified as c). The ingested data can include data from multiple sources, and therefore the data privacy pipeline can fuse data from multiple sources based on computations and constraints specified in the tenant agreement. For example, constrained data fusion can implement one or more constraints to combine ingested data in any number of ways to form fused federated data. These ways include using one or more federated operations (e.g., left, right, inner, outer, inverse), custom federations (e.g., via imperative scripts), data appending, normalization operations, some combination thereof, etc.

[0091] In some embodiments, a data privacy pipeline can perform constrained computations to generate acquired federated data. Constrained computations can acquire data from a source (e.g., ingested data, fused federated data) and perform any number of prescribed computations (e.g., arithmetic operations, aggregation, summarization, filtering, sorting, binding). A simple example of constrained computation is calculating the average age of each city, where the calculation can only be performed on a city if the underlying dataset contains entries for at least five people in that city. Additionally or alternatively, a data privacy pipeline can run data cleaning to generate collaborative data that implements constraints on storage, access, accuracy, etc. For example, data cleaning can implement constraints specified in a tenant agreement that specify whether collaborative data should be protected (e.g., stored in a trustee environment), whether collaborative data can be exported, whether collaborative data should be restricted (e.g., not exporting emails, credit card numbers, or portions thereof), and so on. In this way, a data privacy pipeline can generate collaborative data from data provided by one or more tenants and provide consensual access to the collaborative data without sharing the underlying raw data with all tenants.

[0092] In some embodiments, to enable constraint computation and querying, the use and generation of collaborative data within the trustee environment can be monitored and coordinated by configurable constraints. At a high level, these constraints can be provided via a user interface, allowing tenants (e.g., customers, enterprises, users) to specify the expected computations and constraints for their use and access to their data within the trustee environment, including eligible data sources and how their data is processed or shared. Any number of various types of constraints can be implemented, including data access constraints, data processing constraints, data aggregation constraints, and data cleaning constraints.

[0093] For example, data access constraints can be specified to allow or deny access (e.g., access to specific users, accounts, or organizations). In some embodiments, the specified constraints can be general, such that they apply to all potential data consumers (e.g., allowing only access to average age regardless of the data consumer). In some embodiments, the specified constraints can be applied to specific users, accounts, organizations, etc. (e.g., disallowing group A from accessing salary data, but allowing group B to access salary data). Typically, tenants can specify constraints that define how tenant data can be merged with a specified dataset or a portion thereof, constraints that restrict the patterns of data read from tenant data (e.g., specifying horizontal filtering to be applied to tenant data), constraints that restrict the size of ingested data (e.g., specifying storage limits, subsampling of tenant data, vertical filtering applied to tenant data), constraints that restrict the patterns of collaborative data that can be output, constraints that define ownership of collaborative data, constraints that define whether collaborative data should be open, encrypted, or protected (e.g., stored in a trustee environment), and so on.

[0094] In some embodiments, various types of data processing constraints may be specified, such as constraints specifying which operations can be performed (e.g., for permitted and restricted computations, binary checks), constraints limiting comparison precision (e.g., for numerical data, geographic data, date and time data), constraints limiting accumulation precision (e.g., for geographic data, numerical data, date or time data), constraints limiting location boundary precision (e.g., limiting permitted geofencing determination to a specific grid, minimum geographic partition, such as neighborhood, county, city, state or country, etc.), and other precision and / or data processing requirements.

[0095] Additionally or alternatively, one or more data aggregation constraints can be specified. For example, constraints that require a minimum amount of aggregation (e.g., at least N rows or distinct field values), constraints that require certain statistical distribution conditions to be valid (e.g., minimum standard deviation), and constraints that define allowed aggregation functions (e.g., minimum, maximum, average, but not percentiles).

[0096] In some embodiments, one or more data cleansing constraints may be specified, such as constraints requiring the cleansing of personally identifiable information (e.g., removing emails, names, IDs, credit card numbers), constraints requiring lower precision cleansing (e.g., reducing the precision of numbers, data, and time and / or geolocation), constraints requiring the cleansing of values ​​from specific fields (this may require tracking transformations applied in a data privacy pipeline), constraints requiring custom cleansing (e.g., requiring the execution of one or more custom and / or third-party cleansing scripts), constraints requiring data masking (e.g., outputting some data, such as phone numbers, credit cards, dates, but masking some numbers), and so on.

[0097] Alternatively or additionally to the constraints listed above, one or more constraints may be specified to limit the number of queries and / or data accesses allowed per unit of time (e.g., minutes, hours, days). Such constraints can operate to mitigate the risk of brute-force reverse engineering attempts on protected data by asking a slightly different set of questions within a relatively small time window. Typically, one or more custom constraints may be specified, such as constraints requiring certain attributes to match certain conditions. These and other types of constraints are envisioned in this disclosure.

[0098] In some embodiments, a constraint manager can monitor and coordinate data flows, generation, and access, but subject to specified constraints. For example, the constraint manager can communicate with various components in the trustee environment (e.g., data privacy pipelines) to enforce constraints, which may be stored in a contract database accessible to the constraint manager. In some embodiments, a component can request permission from the constraint manager for other executable units used to perform specific commands, function calls, or logic. The constraint manager can evaluate the request and grant or deny permission. In some cases, permission can be granted, but subject to one or more conditions corresponding to one or more constraints within the constraint. As non-limiting examples, some possible conditions that can be implemented include: requiring operations that move, filter, or reshape data (e.g., applying comparison constraints, such as allowing merging only to a specific precision), requiring the replacement of one or more executable units of logic with one or more constrained executable units of logic (e.g., commands or operations) (e.g., replacing the average with a constrained average), and so on.

[0099] Typically, constraints can be checked, validated, or otherwise enforced at any time or step (e.g., constraint queries associated with any part of the data privacy pipeline). Accordingly, the corresponding functionality for enforcing constraints can be applied at any step or multiple steps. In some embodiments, the enforcement of some constraints can be assigned to certain parts of the data privacy pipeline (e.g., applying data access constraints during ingestion, applying processing and aggregation constraints during data fusion and / or constrained computation, and applying cleaning constraints during data cleaning). In another example, specific data access constraints (e.g., delivering data only to patients participating in at least five different studies) can be applied during data fusion. These are merely examples, and any appropriate constraint enforcement mechanism can be implemented within this disclosure.

[0100] Imposing constraints (e.g., precision or aggregation constraints) on specific executable units of logic (e.g., for a specified computation, a requested operation) can lead to any number of scenarios. In one example, a specific executable unit of logic can be completely denied. In another example, a specific executable unit of logic can be allowed, but the result is filtered (e.g., no value is returned for a specific row or data item). In yet another example, a specific executable unit of logic can be allowed, but the result is modified (e.g., precision is reduced, the question is answered false). These and other variations can be implemented.

[0101] When constraints are applied to generate collaborative data, any combination of schema, constraint, and / or attribute metadata can be associated with the collaborative data, intermediate data used to obtain the collaborative data, or other data. Constraints can typically be implemented across multiple steps and computations. Thus, in some embodiments, applicable and / or satisfied constraints for each step can be tracked and / or associated with the data generated by a given step. For example, with aggregation constraints, once an aggregation constraint has been satisfied during a particular step, subsequent steps no longer need to consider that constraint. In another example, where different constraints have been specified for different datasets to be merged, the merge operation may only require the application of more stringent constraints. Typically, as data flows through a data privacy pipeline, the appropriate allocation or combination of constraints can be applied and / or tracked. This tracking can facilitate verification that a particular constraint has been applied to a particular data. Accordingly, when constraints are applied and data is generated, the corresponding schema, applicable or satisfied constraints, and / or attribute metadata indicating ownership or provision can be associated with the dataset or the corresponding entry, row, field, or other data element. In some embodiments, any intermediate data used to obtain collaborative data (e.g., ingested data, fused joint data, acquired joint data) may be deleted, and the collaborative data may be stored in the trustee environment and / or provided as output, depending on applicable constraints.

[0102] In some embodiments, constraint queries can be applied to allow data consumers to query collaborative data within the trustee environment, subject to configurable constraints. At a higher level, constraint queries can function as a search engine, allowing data consumers to access collaborative data or derive collaborative intelligence from it without exposing the underlying raw data provided by the tenant or the collaborative data protected by the trustee environment. Constraints can be applied to responses to queries in various ways, including reformatting the query before execution, applying constraints after query execution, constraining queries that meet certain conditions for execution, applying access constraints before execution, and so on.

[0103] As a non-limiting example, by ensuring that the query contains at least one aggregate element and that the aggregate element(s) are consistent with the aggregate constraints, the issued query can be validated against the specified aggregate constraints. In another example, the execution plan corresponding to the issued query can be executed, and the results can be validated against the aggregate constraints and / or (or) aggregate elements(s) of the query (e.g., confirming that the results correspond to the requested number of distinct rows, fields, or statistical distributions). In some embodiments, constraints can be applied to corresponding elements of the query by modifying elements based on constraints (e.g., limiting the corresponding number of distinct rows, fields, or statistical distributions), by executing the modified elements before other elements of the query, combinations thereof, or other means.

[0104] For context, queries are typically not executable code. To execute a query, it is usually translated into an executable execution plan. In some embodiments, to impose constraints on a received query, the query can be parsed into a corresponding execution tree, which comprises a hierarchical arrangement of logical executable units that implement the query when executed. Applicable constraints can be accessed, and the logical executable units can be validated against the constraints. In some embodiments, if one or more logical executable units are not permitted, the query can be effectively reformatted by changing one or more of the logical executable units based on one or more constraints. More specifically, the execution tree corresponding to the query can be reformatted into a constrained execution tree by traversing the execution tree and replacing logical executable units that are inconsistent with the specific constraint with custom executable units that are consistent with the specific constraint. Additionally or alternatively, one or more logical executable units can be added to the constrained execution tree to impose constraints (e.g., precision constraints) on the output. These are merely examples; any suitable techniques for generating constrained execution trees can be implemented.

[0105] Typically, the executable units of an execution tree can be validated against a corresponding constraint context, which includes applicable accessed constraints and runtime information, such as information identifying the requesting data consumer issuing the query, information identifying applicable tenant agreements, information identifying the target collaborative data to be manipulated, and so on. Validation of the executable units can involve validating component commands or operations, one or more component parameters, and / or validating other parts of the execution tree. Validation of the executable units can produce several possible results. For example, the executable unit may be allowed (e.g., the operable unit may be copied into the constrained execution tree), the executable unit may be disallowed (e.g., the query may be rejected entirely), or the executable unit may be allowed but modified (e.g., the corresponding constrained executable unit may be copied into the constrained execution tree). In some embodiments, the resulting constrained execution tree is translated into a language used by the trustee environment. The resulting execution tree can be executed (e.g., by traversing and executing the hierarchy of the executable units of the tree), and the results can be returned to the requesting data consumer.

[0106] Therefore, using the implementation described herein, users can efficiently share data through a data trustee that allows them access to collaborative intelligence while ensuring data privacy and providing configurable control and access to the shared data. Related technologies are described in U.S. Patent Application No. 16 / 736,399, filed January 7, 2020, entitled “Multi-Participant and Cross-Environment Pipelines”; U.S. Patent Application No. 16 / 665,916, filed October 28, 2019, entitled “User Interface for Building a Data Privacy Pipeline and Contractual Agreement to Share Data”; and U.S. Patent Application No. 16 / 388,696, filed April 18, 2019, entitled “Data Privacy Pipeline Providing Collaborative Intelligence and Constraint Computing”, the contents of each of these U.S. patent applications are incorporated herein by reference in their entirety.

[0107] Example Distributed Computing Environment

[0108] Now for reference Figure 11 , Figure 11 An example distributed computing environment 1100 that can be implemented using this disclosure is shown. Specifically, Figure 11 A high-level architecture of an example cloud computing platform 1110, capable of hosting a collaborative intelligence environment or a portion thereof (e.g., a data trustee environment), is shown. It should be understood that this and other arrangements described herein are merely examples. For instance, many of the elements described above can be implemented as discrete or distributed components, or combined with other components, and implemented in any suitable combination and location. Other arrangements and elements (e.g., machine, interface, function, command, and function groupings) may be used in addition to those shown.

[0109] Data centers can support distributed computing environments 1100, including cloud computing platforms 1110, racks 1120, and nodes 1130 (e.g., computing devices, processing units, or blades) within the racks 1120. Collaborative intelligence environments and / or data trustee environments can be implemented through cloud computing platforms 1110 that run cloud services across different data centers and geographic regions. The cloud computing platform 1110 can implement infrastructure controller 1140 components for providing and managing the allocation, deployment, upgrades, and management of cloud services. Typically, the cloud computing platform 1110 is used to store data or run service applications in a distributed manner. The cloud computing infrastructure 1110 in the data center can be configured to host and support the operation of endpoints for specific service applications. The cloud computing infrastructure 1110 can be a public cloud, a private cloud, or a dedicated cloud.

[0110] Node 1130 may be equipped with a host 1150 (e.g., an operating system or runtime environment) on which a defined software stack runs. Node 1130 may also be configured to perform specialized functions (e.g., compute nodes or storage nodes) within cloud computing platform 1110. Node 1130 is assigned to run one or more portions of a tenant's service application. A tenant may refer to a customer using resources of cloud computing platform 1110. Service application components of cloud computing platform 1110 that support a particular tenant may be referred to as tenant infrastructure or a tenant. The terms service application, application, or service are used interchangeably herein and broadly refer to any software or software portion that runs on top of or accesses storage and compute equipment locations within a data center.

[0111] When node 1130 supports more than one individual service application, node 1130 can be partitioned into multiple virtual machines (e.g., virtual machine 1152 and virtual machine 1154). Physical machines can also run individual service applications simultaneously. Virtual machines or physical machines can be configured as personalized computing environments supported by resources 1160 (e.g., hardware and software resources) in the cloud computing platform 1110. It is conceivable to configure resources for specific service applications. Furthermore, each service application can be partitioned into functional parts, allowing each functional part to run on a separate virtual machine. In the cloud computing platform 1110, multiple servers can be used to run service applications and perform data storage operations in a cluster. Specifically, servers can perform data operations independently but are exposed as a single device called a cluster. Each server in the cluster can be implemented as a node.

[0112] Client device 1180 can be linked to service applications in cloud computing platform 1110. Client device 1180 can be any type of computing device, for example, it can correspond to a reference... Figure 12The computing device 1200 is described. Client device 1180 can be configured to issue commands to cloud computing platform 1110. In embodiments, client device 1180 can communicate with service applications via Virtual Internet Protocol (IP) and load balancers or other means of directing communication requests to designated endpoints within cloud computing platform 1110. Components of cloud computing platform 1110 can communicate with each other via a network (not shown), which may include, but is not limited to, one or more local area networks (LANs) and / or wide area networks (WANs).

[0113] Example operating environment

[0114] Following a brief overview of embodiments of the invention, example operating environments in which the various embodiments of the invention may be implemented are described below to provide a general context for the various aspects of the invention. First, special reference is made to… Figure 12 This illustration shows an example operating environment for implementing embodiments of the invention, and is generally designated as computing device 1200. Computing device 1200 is merely one example of a suitable computing environment and is not intended to impose any limitations on the scope or functionality of the invention. Computing device 1200 should also not be construed as having any dependency or requirement associated with any or a combination of the components shown.

[0115] This invention can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions (e.g., program modules) that are executed by a computer or other machine (e.g., a personal data assistant or other handheld device). Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This invention can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, and more specialized computing devices. This invention can also be practiced in distributed computing environments, where tasks are performed by remote processing devices linked via a communication network.

[0116] refer to Figure 12 The computing device 1200 includes a bus 1210 that is directly or indirectly coupled to the following devices: a memory 1212, one or more processors 1214, one or more presentation units 1216, an input / output port 1218, an input / output unit 1220, and an exemplary power supply 1222. The bus 1210 can represent one or more buses (e.g., an address bus, a data bus, or a combination thereof). For clarity of concept, Figure 12 The various boxes are shown with lines, and other arrangements of the described components and / or component functions are also envisioned. For example, presentation components such as those of a display device can be considered as I / O components. Furthermore, processors also have memory. We consider this common knowledge in the art, and we emphasize again... Figure 12The illustrations are merely illustrative examples of computing devices that can be used in conjunction with one or more embodiments of the present invention. There is no distinction between categories such as "workstation," "server," "laptop," and "handheld device," as all of these fall under this category. Figure 12 Within the scope of, and refer to “Computing Devices”.

[0117] Computing device 1200 typically includes a variety of computer-readable media. Computer-readable media can be any available medium that can be accessed by computing device 1200, and includes volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.

[0118] Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, cassette tape, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by the computing device 1200. Computer storage media itself does not include signals.

[0119] Communication media typically embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and include any information transmission medium. The term "modulated data signal" refers to a signal whose one or more characteristics are set or altered in a manner that encodes information within the signal. By way of example and not limitation, communication media include wired media such as wired networks or direct wired connections, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.

[0120] Memory 1212 includes computer storage media in the form of volatile and / or non-volatile memory. Memory can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. Computing device 1200 includes one or more processors that read data from various entities such as memory 1212 or I / O components 1220. Multiple presentation components 1216 present data indications to a user or other device. Exemplary presentation components include display devices, speakers, printing components, vibrating components, etc.

[0121] I / O port 1218 allows computing device 1200 to be logically coupled to other devices including I / O components 1220, some of which may be built-in. Exemplary components include microphones, joysticks, gamepads, satellite dish antennas, scanners, printers, wireless devices, etc.

[0122] Referring to the collaborative intelligence environment described herein, the embodiments described herein support constraint computation and / or constraint querying. Components of the collaborative intelligence environment may be integrated components, including a hardware architecture and software framework that support constraint computation and / or constraint querying functionality within the collaborative intelligence system. The hardware architecture refers to the physical components and their interrelationships, while the software framework refers to the software that provides functionality, which can be implemented using hardware embodied on the device.

[0123] End-to-end software-based systems can operate within system components to manipulate computer hardware to provide system functionality. At a low level, the hardware processor executes instructions selected from the given processor's machine language (also known as machine code or native) instruction set. The processor recognizes native instructions and performs corresponding low-level functions related to, for example, logic, control, and memory operations. Low-level software written in machine code can provide more complex functionality for higher-level software. As used herein, computer-executable instructions include any software, including low-level software written in machine code, high-level software such as application software, and any combination thereof. In this respect, system components can manage resources and provide services for system functionality. Any other variations and combinations thereof are contemplated for embodiments of the invention.

[0124] As an example, a collaborative intelligence system may include an API library containing specifications for routines, data structures, object classes, and variables that can support interaction between the device's hardware architecture and the collaborative intelligence system's software framework. These APIs include configuration specifications for the collaborative intelligence system, enabling different components within it to communicate with each other, as described herein.

[0125] Various components used herein have been identified, and it should be understood that any number of components and arrangements can be employed to achieve the desired functionality within the scope of this disclosure. For example, for clarity of concept, components in the embodiments depicted in the figures are shown with lines. Other arrangements of these and other components can also be implemented. For example, although some components are described as single components, many elements described herein can be implemented as discrete or distributed components, or combined with other components, and implemented in any suitable combination and location. Some components may be omitted entirely. Furthermore, the various functions described herein as being performed by one or more entities can be implemented by hardware, firmware, and / or software, as described below. For example, various functions can be performed by a processor executing instructions stored in memory. Therefore, other arrangements and elements (e.g., machines, interfaces, functions, commands, and function groups) can be used in addition to the arrangements and elements shown.

[0126] The embodiments described in the following paragraphs can be combined with one or more of the specifically described alternatives. In particular, the claimed embodiments may contain references to more than one other embodiment. The claimed embodiments may specify further limitations on the claimed subject matter.

[0127] The subject matter of embodiments of the present invention has been specifically described herein to satisfy legal requirements. However, the specification itself is not intended to limit the scope of this patent. Rather, the inventors have envisioned that the claimed subject matter may also be embodied in other ways to include different steps or combinations of steps similar to those described in this document, as well as other current or future techniques. Furthermore, although the terms “step” and / or “box” may be used herein to refer to different elements of the method employed, these terms should not be construed as implying any particular order between or between the various steps disclosed herein, unless the order of the various steps is explicitly described.

[0128] For the purposes of this disclosure, the word "comprising" has the same broad meaning as "including," and the word "access" includes "receiving," "referencing," or "retrieval." Furthermore, the word "communication" has the same broad meaning as "receiving" or "transmitting" implemented using a software- or hardware-based bus, receiver, or transmitter employing the communication medium described herein. Additionally, unless otherwise stated, words such as "a" and "an" include both plural and singular forms. Thus, for example, the constraint of "feature" is satisfied when one or more features are present. Furthermore, the term "or" includes conjunctions, disjuncts, and both (therefore, a or b includes both a or b and both a and b).

[0129] For the purposes of the detailed discussion above, embodiments of the invention are described with reference to a distributed computing environment; however, the distributed computing environment described herein is merely exemplary. Components may be configured to perform novel aspects of the embodiments, wherein the term "configured for" may mean "programmed to" perform a particular task or implement a particular abstract data type using code. Furthermore, while embodiments of the invention can generally be referred to with reference to the collaborative intelligence environment and schematic diagrams described herein, it should be understood that the described techniques can be extended to other implementation environments.

[0130] Embodiments of the invention have been described in conjunction with specific examples, which are intended in all respects to be illustrative and not restrictive. Alternative embodiments will become apparent to those skilled in the art without departing from the scope of the invention.

[0131] As can be seen from the foregoing, the present invention is well adapted to achieve all of the above-mentioned objectives and goals, and can also achieve other obvious and inherent structural advantages.

[0132] It should be understood that some features and sub-combinations are useful and can be used without reference to other features or sub-combinations. This is within the scope of the claims and is contemplated by the claims.

Claims

1. A data trustee system for implementing a data privacy pipeline, comprising: One or more computer storage media storing computer-usable instructions that, when used by one or more computing devices, cause the one or more computing devices to perform operations including: Receive requests from the data consumer to trigger a data privacy pipeline that needs to be executed within the data trustee's environment; Identify all root entities of the data privacy pipeline, all of which require rights from an grantor who is not a participant in the data privacy pipeline; Load multiple contracts that manage access to the root entity within the data trustee's environment; For each of the root entities, the plurality of contracts are searched to identify a valid access path based on one of the associated contracts. Utilizing this valid access path, the data privacy pipeline can use the root entity while simultaneously implementing constraints defined by the associated contract and applicable when accessing the root entity, and while implementing strategies defined by the associated contract and applicable to computations downstream of the root entity within the data privacy pipeline. Based on each root entity having an identified valid access path, the identified valid access path triggers the execution of the data privacy pipeline within the data trustee environment according to an identified associated contract among the plurality of contracts, using the identified valid access path and the identified associated contract to access each root entity without exposing the root entity; The search of the plurality of contracts to identify a valid access path for each root entity includes: evaluating the computation steps of the potential access path without performing the first set of computation steps in the computation steps of the potential access path, and conditionally performing the second set of computation steps to evaluate specific constraints or policies that can only be verified during runtime.

2. The data trustee system of claim 1, wherein all root entities identifying the claims of the data privacy pipeline include: Access the digital representation of the data privacy pipeline, which has an associated list, attributes, or metadata that identifies the root entity.

3. The data trustee system according to claim 1, further comprising: Before searching the multiple contracts to identify a valid access path for each root entity, a set of the contracts that do not grant access to the data consumer are filtered out based on the identity of the data consumer.

4. The data trustee system of claim 1, wherein searching the plurality of contracts for each root entity comprises: For each of the multiple contracts that manages access to the root entity: Identify potential access paths containing all computational steps that would require execution within the data trustee environment to trigger the data privacy pipeline within that environment, thereby accessing the root entity using the contract; and Determine whether the potential access path will implement the constraints defined by the contract and applicable when accessing the root entity, and whether it will implement the strategy defined by the contract and applicable to a set of computational steps downstream of the root entity.

5. The data trustee system of claim 1, wherein searching the plurality of contracts for each root entity comprises: For each of the multiple contracts that manages access to the root entity: Identify potential access paths containing all computational steps that would require execution within the data trustee environment to trigger the data privacy pipeline within that environment, thereby accessing the root entity using the contract; and The computational steps for verifying the potential access path will satisfy applicable constraints and policies, without performing the computational steps.

6. The data trustee system of claim 1, wherein searching the plurality of contracts to identify valid access paths for each root entity, identifying a plurality of candidate contracts or a plurality of valid access paths for at least a first root entity among the root entities, the operation further includes: Conflict rules are applied to select one of the plurality of valid access paths as the identified valid access, or to select one of the plurality of candidate contracts as the identified associated contract for the first root entity.

7. One or more computer storage media storing computer-usable instructions, which, when used by one or more computing devices, cause the one or more computing devices to perform operations including: Receive requests from the data consumer to trigger a data privacy pipeline that needs to be executed within the data trustee's environment; Identify all root entities of the data privacy pipeline, all of which require rights from an grantor who is not a participant in the data privacy pipeline; Identify a set of contracts that manage access to the root entity within the data trustee environment and define valid access paths for each root entity, so that the data privacy pipeline can use the root entity while implementing constraints and policies defined by the set of contracts, the constraints being applicable when accessing the root entity and the policies being applicable to computations of the data privacy pipeline downstream of the root entity; as well as The execution of the data privacy pipeline within the data trustee environment is triggered to access the root entity using the identified set of contracts without exposing the root entity; The set of contracts that identifies a valid access path for each root entity includes: evaluating the computation steps of the potential access path without performing the first set of computation steps in the computation steps of the potential access path, and conditionally performing the second set of computation steps to evaluate specific constraints or policies that can only be verified during runtime.

8. The computer storage medium of claim 7, wherein all root entities identifying the claims of the data privacy pipeline include: Access the digital representation of the data privacy pipeline, which has an associated list, attributes, or metadata that identifies the root entity.

9. The operation further comprises: Load multiple contracts that manage access to the root entity within the data trustee's environment; Based on the identity of the data consumer, a subset of the contracts that do not grant access to the data consumer are filtered out, leaving the remaining set of contracts; Search the remaining set of contracts to identify the set of contracts that manage access to the root entity.

10. The one or more computer storage media of claim 7, wherein the set of contracts identifying a valid access path for each root entity of the root entity comprises: For each root entity and each contract that manages access to said root entity: Identify potential access paths containing all computational steps that would require execution within the data trustee environment to trigger the data privacy pipeline within that environment, thereby accessing the root entity using the contract; and Determine whether the potential access path will implement the constraints defined by the contract and applicable when accessing the root entity, and whether it will implement the strategy defined by the contract and applicable to a set of computational steps downstream of the root entity.

11. The computer storage medium of claim 7, wherein the set of contracts identifying a valid access path for each root entity of the root entity comprises: For each root entity and each contract that manages access to said root entity: Identify potential access paths containing all computational steps that would require execution within the data trustee environment to trigger the data privacy pipeline within that environment, thereby accessing the root entity using the contract; and The computational steps for verifying the potential access path will satisfy applicable constraints and policies, without performing the computational steps.

12. The computer storage medium of claim 7, wherein the set of contracts identifying the valid access path includes at least a first root entity among the root entities: Identify multiple candidate contracts or multiple valid access paths for access to the first root entity; and Conflict rules are applied to identify a single contract and a single valid access path for the first root entity based on at least one of the plurality of candidate contracts or the plurality of valid access paths.

13. A method for performing a data privacy pipeline, comprising: Receive a request from the data consumer to trigger a data privacy pipeline that needs to be executed in a constrained environment inaccessible to the data consumer, and export the data generated by the data privacy pipeline from the constrained environment. It is determined that executing the data privacy pipeline within the constrained environment will satisfy the associated rights of the root entity using the data privacy pipeline within the constrained environment, the associated rights specifying the constraints on accessing the root entity within the constrained environment and the strategy for downstream computations derived from the root entity within the constrained environment; It is determined that the data consumer has permission to export the data from the constrained environment; as well as Trigger the execution of the data privacy pipeline within the constrained environment to use the root entity in accordance with the associated rights without exposing the root entity; The determination that executing the data privacy pipeline within the constrained environment will satisfy the associated rights using the root entity includes: for each of the multiple contracts managing access to the root entity within the constrained environment: Identify potential access paths containing all computational steps that would require execution within the constrained environment to trigger the data privacy pipeline within that environment, thereby accessing the root entity using the contract; and The computational steps that verify the potential access path will satisfy applicable constraints and policies, without performing the first set of computational steps, and conditionally performing the second set of computational steps to evaluate specific constraints or policies that can only be verified during runtime.

14. The method according to claim 13, further comprising: Identify all root entities of the data privacy pipeline, wherein all root entities require corresponding rights from the grantor, who is not a participant in the data privacy pipeline; Determining that executing the data privacy pipeline within the constrained environment will satisfy the associated rights using the root entity includes: searching for contracts that manage access to the root entity within the constrained environment to identify valid access paths for each root entity.

15. The method according to claim 13, further comprising: All root entities of the data privacy pipeline are identified by accessing a digital representation of an associated list, attribute, or metadata that identifies the root entity, and all root entities require corresponding rights from an grantor who is not a participant in the data privacy pipeline.

16. The method according to claim 13, further comprising: Identify all root entities of the data privacy pipeline, wherein all root entities require corresponding rights from the grantor, who is not a participant in the data privacy pipeline; Load multiple contracts that manage access to the root entity within the constrained environment; as well as Based on the identity of the data consumer, a subset of the multiple contracts that do not grant access to the data consumer are filtered out, leaving the remaining set of contracts; Determining that executing the data privacy pipeline within the constrained environment will satisfy the associated rights of the root entity includes: searching the remaining set of contracts to identify valid access paths to the root entity.

17. The method of claim 13, wherein determining that executing the data privacy pipeline within the constrained environment will satisfy the associated rights to use the root entity includes: For each of the multiple contracts that manage access to the root entity within the constrained environment: Identify potential access paths containing all computational steps that would require execution within the constrained environment to trigger the data privacy pipeline within that environment, thereby accessing the root entity using the contract; and Determine whether the potential access path will implement the constraints defined by the contract and applicable when accessing the root entity, and whether it will implement the strategy defined by the contract and applicable to a set of computational steps downstream of the root entity.

Citation Information

Patent Citations

  • Data privacy pipeline providing collaborative intelligence and constraint computing

    US20200334370A1

  • User interface for building a data privacy pipeline and contractual agreement to share data

    US20200334377A1

  • Multi-participant and cross-environment pipelines

    US20200336488A1

  • Sharable multi-tenant reference data utility and repository, including value enhancement and on-demand data delivery and methods of operation

    WO2006076520A2