Data management and governance system and method

JP2024536689A5Pending Publication Date: 2025-10-03INTERTRUST TECH CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024510622
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-10
Filing Date
2022-09-09
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Enterprises face challenges in securely managing and governing data across multiple departments and geographic locations, ensuring compliance with regulations like GDPR, while maintaining data integrity and preventing unauthorized access and duplication, especially in on-premises and cloud environments.

Method used

A system and method for data management and governance that includes Identity and Access Management (IAM), Data Virtualization (DV), Secure Execution Environment (SEE), and Time Series Database (TSDB) services, providing secure data access, governance, and interoperability without requiring system redesign, using APIs for integration and supporting compliance with regulations.

Benefits of technology

Enables secure, manageable, and governable data access across departments and locations, ensuring compliance with regulations, and facilitating collaboration while preventing data duplication and unauthorized access, maintaining data integrity in on-premises and cloud environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure relates, inter alia, to systems and methods for scalable data processing, storage, and / or management. Certain embodiments disclosed herein provide a data management architecture that allows for more secure storage of enterprise data, makes it more secure, usable, and / or interoperable, and facilitates data usage across information silos. Further embodiments provide comprehensive data access authentication and / or authorization capabilities between various services included in embodiments of the disclosed architecture.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Related Applications This application claims the benefit of priority under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 243,067, filed September 10, 2021, and entitled “DATA MANAGEMENT AND GOVERNANCE SYSTEMS AND METHODS,” the contents of which are incorporated herein by reference in their entirety.

[0002] Copyright Permission Portions of the disclosure of this patent document may contain material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction of either the patent document or the patent disclosure, as appearing in the U.S. Patent and Trademark Office patent file or records, but otherwise reserves all copyright rights whatsoever. Summary of the Invention

[0003] The present disclosure relates generally to systems and methods for securely managing data, and more particularly, but not exclusively, to systems and methods for managing data, enforcing rights and / or other governance terms associated with the data, and / or providing an execution environment for facilitating collaboration among entities that interact with the data.

[0004] Enterprises generate and store valuable data in many internal and / or external applications. To ensure that data consumers do not copy and / or publish their data and / or create governance and security risks, many enterprises may want to identify their valuable data and ensure it is accessed in a secure, manageable, and / or otherwise governable manner.

[0005] The systems and methods disclosed herein provide various mechanisms for addressing these challenges. In various embodiments, the disclosed systems and methods may be used to govern data residing on-premise and in the cloud without duplicating and / or migrating data, to perform audit data access to ensure compliance with government, jurisdictional and / or industry regulations, to provide a secure execution environment to facilitate collaboration with partners and service providers without exposing data, etc.

[0006] Various embodiments disclosed herein may be described in connection with one or more non-limiting examples. Some non-limiting examples may reference a fictitious company ACME to illustrate various aspects of the disclosed systems and methods. In various examples, ACME may be associated with data across multiple departments (e.g., sales, service, human resources, etc.) and / or geographic locations. Each department may use different tools and / or techniques to access, manage, and / or interact with the data based on business needs. The geographic diversity of the company may introduce certain challenges with respect to how data is maintained and / or managed (e.g., GDPR restrictions and / or the like). Various non-limiting examples related to ACME described herein may illustrate how aspects of the disclosed systems and methods may address various challenges related to data sovereignty and / or management and should be viewed as illustrative of various embodiments, not limiting. [Brief description of the drawings]

[0007] The work body of the present invention will be readily understood by reference to the following detailed description taken in conjunction with the accompanying drawings. [Figure 1] 1 illustrates a non-limiting example of a data management architecture and associated interactions consistent with certain embodiments disclosed herein. [Diagram 2] 1 illustrates a non-limiting example of management of a data set using a directory service consistent with certain embodiments disclosed herein. [Diagram 3] 1 illustrates a non-limiting example of querying a data set using a directory service consistent with certain embodiments disclosed herein. [Figure 4] 1 illustrates a non-limiting example of a time series database data management architecture consistent with certain embodiments disclosed herein. [Diagram 5] 1 illustrates a flowchart of a non-limiting example of a data query and access authentication process consistent with certain embodiments disclosed herein. [Figure 6] 1 illustrates non-limiting examples of systems that can be used to implement certain embodiments of the systems and methods of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0008] Detailed descriptions of systems and methods consistent with embodiments of the present disclosure are provided below. Although several embodiments are described, it should be understood that the present disclosure is not limited to any one embodiment, but instead encompasses numerous alternatives, modifications, and equivalents. Numerous specific details are set forth in the following description to provide a thorough understanding of the embodiments disclosed herein, although some embodiments may be practiced without some or all of these details. Moreover, for the sake of clarity, certain technical material that is known in the relevant art has not been described in detail to avoid unnecessarily obscuring the disclosure.

[0009] The embodiments of the present disclosure can be understood by reference to the drawings. The components of the disclosed embodiments can be arranged and designed in a wide variety of configurations, as generally described and illustrated in the figures herein. Thus, the following detailed description of embodiments of the systems and methods of the present disclosure is not intended to limit the scope of the disclosure as claimed, but is merely representative of possible embodiments of the disclosure. In addition, the steps of any method disclosed herein do not necessarily have to be performed in any particular order, nor sequentially, nor do steps have to be performed only once, unless otherwise specified.

[0010] Embodiments of the disclosed systems and methods provide data management services that facilitate secure data rights management and governance, interoperability, and / or analytical capabilities in a deployment-independent environment. In some embodiments, the data management services may be designed, at least in part, to serve the needs of an enterprise data workflow. The data management services may comprise multiple services focused on specific aspects of the data lifecycle, including, for example, but not limited to, ingestion, storage, analysis, processing, access, and / or distribution.

[0011] In certain embodiments, aspects of the disclosed data management services can make enterprise data more secure, usable, and / or interoperable, facilitating data usage across information silos, potentially without the need to redesign system architectures or move large amounts of data to new systems. Various embodiments can be integrated in conjunction with existing software ecosystems to provide security and governance solutions without significantly disrupting established workflows. In some embodiments, application programming interfaces ("APIs") may be used to integrate various developer applications. The data management services can further provide comprehensive authentication, authorization, and / or data access capabilities to applications via appropriate protocols.

[0012] Various embodiments of the disclosed data management services may provide a combination of services and / or applications, which may include, for example, but are not limited to, one or more of the following: · Identity and Access Management ("IAM") services, which may include security, directory and / or metadata, components and / or services. · Data Virtualization (“DV”) services, which may include catalog services, data services, etc. · Secure Execution Environment ("SEE") Services. · Time Series Database ("TSDB") Services. ·Audit services.

[0013] System Architecture and Example Interactions 1 illustrates a non-limiting example of a data management architecture 100 and associated interactions consistent with certain embodiments disclosed herein. Non-limiting examples of the various components and / or services of the illustrated architecture 100 and interactions between the illustrated components and / or services and other users, systems, and / or services are described herein below. It will be understood that several variations can be made to the architecture 100 and related relationships, examples, and / or interactions within the scope of the inventive body of work. For example, but not by way of limitation, certain illustrated and / or described components and / or services may be combined and / or distributed among multiple components, systems, and / or services.

[0014] 1 and described in detail below, the architecture 100 may include one or more data sources 102, a TSDB service 104 that may interact with one or more storage tiers 106, 108, an IAM service 110 that may include an associated directory 112 and / or interact in other ways, a DV service 114, and an SEE 116 that may be used to execute one or more applications 118 in a protected manner using a sandboxed execution environment 120. In some embodiments, by interacting with the DV service 114, one or more physical and / or virtual datasets 122-126 may be accessed and, in some implementations, may be associated with and / or mapped to data stored and / or managed by the TSDB service 104 in one or more service tiers.

[0015] Consistent with embodiments disclosed herein, one or more data sources 102, which may include, for example, without limitation, Internet of Things ("IOT") devices, wind turbine systems and / or associated sensors, nuclear reactors and / or other energy generating systems, manufacturing facility systems, vehicles and / or associated systems and / or sensors, and / or the like, may ingest data into the TSDB service 104. While shown providing data for ingestion directly to the TSDB service 104, it will be understood that one or more data sources 102 may provide data for ingestion to the TSDB service 104 via one or more intermediate systems and / or services. For example, in various embodiments, the data sources 102 may include the original source of the data (e.g., a source that generated the data) and / or a source that provides the data for ingestion to the service. As used herein, depending on the context, the term "data source" may also be used to describe a data source object to which a dataset is bound, where the data source refers to where the stored data in the dataset resides (e.g., SQL database, AWS / S3 parquet files, elastic search, etc.).

[0016] The TSDB service 104 may authenticate and / or perform data ingest access checks and / or any other related authorization checks using the IAM service 110 based on information included in a data ingest request issued by a data source 102 seeking to ingest data into the service. As described in more detail herein, the IAM service 110 may interact with the directory 112 in connection with the data access and / or ingest request and / or other authorization check processes.

[0017] If a data source 102 seeking to ingest data into the service is authorized to do so (e.g., based on interaction with the IAM service 110), the TSDB service 104 may store the ingested data in the hot storage tier 106. Consistent with embodiments disclosed herein, data stored in the hot storage tier 106 may be migrated to storage in the cold storage tier 108.

[0018] In at least one non-limiting example, the privileged user system 128 can query the TSDB service 104. The query can include, for example, but not limited to, information identifying the queried data and / or data set and / or credentials and / or other associated identifying information that can be used to authenticate the user system 128 and / or associated applications and / or users. The TSDB service 104 can query the IAM service 110 to determine whether the privileged user system 128 and / or associated users and / or applications are authenticated and / or authorized to access and / or otherwise use the queried data, potentially interacting with the directory 112 in connection with the authentication and / or access determination.

[0019] If the user system 128 is authenticated and access is permitted, the TSDB service 104 may retrieve the queried data from the hot storage tier 106 and / or the cold storage tier 108. The TSDB service 104 may then provide the retrieved data to the user system 128.

[0020] In another non-limiting example, an application 118 executing within the sandboxed execution environment 120 of the SEE 116 may query the DV service 114 for data and / or datasets managed by the TSDB service 104. In various embodiments, the DV service 114 may comprise one or more constituent services, which may include, for example, but are not limited to, a catalog service 138 and / or a data service 140. As described in more detail below, the catalog service 138 may manage datasets and / or DV objects. The data service 140 may query and / or perform certain other procedures in connection with datasets, which may include, for example, but are not limited to, one or more virtual datasets 122-126. As used herein, for purposes of clarity and explanation, certain functions of the component services of the DV service 114 (e.g., the catalog service 138, the data service 140, etc.) may be generally described as being performed by the DV service 114. However, in certain other examples, some DV service functions may be described as being performed by component catalog service 138 and / or data service 140. For example, in some embodiments, certain service calls that may include services of component DV service 114 may be described herein as being issued directly to catalog service 138 and / or data service 140.

[0021] The DV service 114 may authenticate the application 118 and / or associated user by querying the IAM service 110 to authenticate the application 118 and / or user and / or to determine whether the application 118 and / or user is authorized to access and / or otherwise use the queried data and / or data set. If the application 118 and / or user is authenticated and / or authorized to access and / or otherwise use the queried data and / or data set, the DV service 114 may query the TSDB service 104 for the data and / or data set.

[0022] The TSDB service 104 can authenticate the DV service 114 and / or determine associated access rights and / or permissions (e.g., by interacting with the IAM service 110). The TSDB service 104 can retrieve the queried data from the hot storage tier 106 and / or the cold storage tier 108 and return the queried data and / or data set to the DV service 114 for communication to the requesting application 118.

[0023] In a further non-limiting example, an application 118 executing within the sandboxed execution environment 120 of the SEE 116 can query the DV service 114 for access to and / or use of one or more virtual datasets 122-126 managed by the DV service 114. In certain embodiments, one or more virtual datasets 122-126 may be mapped to and / or associated with data and / or datasets managed by the TSDB service 104. In further embodiments, one or more virtual datasets 122-126 may be mapped to and / or associated with data and / or datasets not directly and / or indirectly managed by the TSDB service 104 (e.g., data outside the TSDB service 104). The DV service 114 may authenticate the application 118 and / or associated user by querying the IAM service 110 to authenticate the application 118 and / or user and / or to determine whether the application 118 and / or user is authorized to access and / or otherwise use the queried data and / or data set(s). If the application 118 and / or user is authenticated and / or authorized to access and / or otherwise use the queried data, the DV service 114 may retrieve the queried data and / or data set(s).

[0024] The application 118 may further interact with the DV service 114 in connection with storing data as one or more virtual datasets 122-126 managed by the DV service 114. For example, the application 118 may request that the DV service 114 ingest particular data into one or more virtual datasets and / or the TSDB service 104. The DV service 114 may interact with the IAM service 110 in connection with authenticating the application 118 and / or associated user and / or determining whether the application 118 and / or user is authorized to ingest data into the DV service 114. If authorized, the DV service 114 may ingest the data into one or more virtual datasets 122-126 managed by the DV service 114 and / or the TSDB service 104.

[0025] In some embodiments, a user may directly query the DV service 114 for data managed within one or more of the virtual datasets 122-126 and / or the TSDB service 104. For example, as shown, a user system 130 may query the DV service 118 to access data managed by the DV service 118 and / or the TSDB service 104. The DV service 114 may interact with the IAM service 110 in connection with authenticating the user system 130 and / or associated user and / or determining whether the user system 130 and / or user is authorized to access the queried data. If authenticated and / or otherwise authorized, the DV service 114 may provide the queried data and / or dataset to the requesting user system 130.

[0026] In particular embodiments, an application 118 executing within a sandbox environment 120 of a SEE service 116 may be permitted to query some data (e.g., access some data / functionality) by virtue of the fact that it is executing within the sandbox environment 120. However, when executed outside of the sandbox environment 120, the application 118 may not be permitted to successfully make a query, depending on applicable rules (if any).

[0027] Various aspects and / or details regarding the various components and / or services of the illustrated architecture 100, as well as further non-limiting examples of interactions between the illustrated components and / or services and other users, systems, and / or services, are described in more detail below.

[0028] Data Management Service Application Embodiments of the disclosed data management services can provide services to application developers that enable them to create secure applications and / or applications that interact with various systems and / or services of the disclosed architecture. In various implementations, the applications can include one or more of the following: Applications deployed along with data management services, for example, but not limited to, IAM applications, catalog applications, SEE applications, TSDB applications, audit applications. Applications running within the SEE 116 that may be deployed and monitored by the SEE Service. These applications may use the Data Management Service APIs or may be developed by users of the Data Management Service. Applications deployed outside the Data Management Service that may use APIs and / or services provided by the Data Management Service.

[0029] Data Management Services - API Data management service users can access the services through the service API. Data management service developers can call the service API in their applications. Data management service users and / or managers may use applications that can call the API on their behalf.

[0030] The Data Management Service API can provide a programmatic interface to Data Management Service functionality. For example, there can be APIs used to access data from a dataset, create organization, account, and / or group objects, and / or run workloads on a cluster (e.g., a Kubernetes cluster). The Data Management Service API can be exposed by core data management services, which can include, for example, but are not limited to, the IAM service 110, the SEE service 116, the catalog service, the DV service 114, and / or the TSDB service 104. For example, a Data Management Service application can authenticate a user by calling the IAM Service API.

[0031] Data Management Services - Application Development Data management service developers can use any programming language to build applications. The data management service can be referenced via a REST-based API, so in some embodiments, any language capable of making HTTP requests can be used. Developers can also build applications in one or more containers, reference them in components, reference the components in workloads, and / or run workloads under the SEE 116 in a cluster.

[0032] Developers can use different services for their applications, including, for example, but not limited to, one or more of the following: · IAM services 110 may include security, directory, and metadata services to perform identification, authentication, and authorization, and manage a large set of data management service entities (organizations, users, accounts, groups, sources, applications, clients, etc.). · SEE services 116 for managing workloads, clusters, and cluster instances and for supporting workloads running on the cluster instances. A catalog service(s) that may contain definitions and mappings of data sources, datasets, procedures, and / or workspaces. The catalog service may also perform create, read, update, delete ("CRUD") operations on them. In some embodiments, procedures may comprise a type of data management service object managed by the catalog service. Procedures may be used to both query and / or update datasets. Procedure objects, like other data management service objects, may be governed objects, and operations on such objects are governed based on privileges associated with those operations. · A DV service 114 that can govern access to data sources and data sets and can restrict access based on values ​​returned from access checks (including enforcement of restrictions).

[0033] Identity and Access Management The IAM service 110 can define and enforce security rules by managing associated entities and providing identity and authorization services (e.g., based on OAuth2 and OpenID Connect). The IAM service 116 can include a security component 132, a directory component 134, and a metadata component 136.

[0034] IAM Service - Access Token A data management service user can log into the system and receive access tokens. These tokens can be used to authenticate the user in data management service API calls. Access tokens can be obtained through the IAM service 110. To obtain an access token without going through a web application, the user can call the correct IAM service endpoint and specify the appropriate client credentials (e.g., client ID and client secret, user and / or API key credentials, etc.). In some embodiments, the user credentials may include an email address or a username and password, although it will be understood that various other user credentials may be used. The API key credentials may comprise an API key ID and / or an API key secret. When the user accesses the data management service web application, the user may be redirected to the IAM service 110 to log in. The IAM service 110 can then return the access token to the web application after successful authentication, which can then pass it to various data management services (e.g., via the data management service APIs).

[0035] The access token may be used, for example, but not limited to, to obtain one or more of an account ID, a tenant ID, a client ID, an application ID, and / or an organization ID associated with the logged-in session.

[0036] In certain embodiments, the access token may have a limited and / or configurable lifespan. In some implementations, security may be improved by limiting the lifespan of the access token to a potentially relatively short duration (e.g., one hour). However, for user convenience reasons, a longer lifespan may be used (e.g., one week). Upon expiration, the user and / or application that obtained the access token may obtain a new access token to continue making API calls to the data management service. In certain embodiments, there may be multiple ways to re-obtain an access token. For example, user and / or API key credentials may be provided in the same and / or similar manner that the initial access token is obtained.

[0037] In further embodiments, a refresh token may be used, which may be optionally issued when an access token is issued. In some embodiments, the refresh token may be a single-use token. In some embodiments, the refresh token may have a longer validity period than the access token. It may be provided to the IAM service 110 to obtain new access and refresh tokens. The expiration intervals for the access and refresh tokens may be specified at the deployment level and / or tenant level (e.g., per organization).

[0038] IAM Service - Security Components The security component 132 of the IAM service 110 can manage rule sets, which can include rules that describe privileges granted to data management service subjects to access specified data management service objects. Rule subjects can include accounts, groups, and / or organizations, and rule objects can be any data management service entity in the directory.

[0039] In some embodiments, the security component 132 can manage role authorizations that can provide role-based access control services. Role authorizations can grant a subject a particular role. A role can then name a policy that can define rules similar to the rules in a rule set. In some embodiments, subjects can be excluded because subjects can be specified in the role authorization.

[0040] The APIs provided by the security component 132 can be used to, among other things: · Perform identification, authentication, and / or authorization of data management service subjects. · Perform access checks on built-in and custom entities. Enumerate accessible objects for an account. · Determine which subjects have access to a specified object including the privileges that those subjects have been granted.

[0041] In various embodiments, the governed entities may include entities managed by the catalog service (e.g., datasets, data sources, procedures, and / or workspaces), entities managed by the SEE service 116 (e.g., components, workloads, deployments, and / or namespaces, etc.), entities managed by the TSDB service 104 (e.g., tables and / or namespaces, etc.), and entities managed by the IAM service 110 itself (e.g., organizations, accounts, groups, folders, etc.). The governed entities may also be custom entities that may be defined by a customer and / or may represent application-specific entities.

[0042] Subject Identification, Authentication, and Authorization A data management service subject may first be identified and / or otherwise recognized and authenticated by the system and then authorized to use the system. Identification may involve maintaining an account ID, and a user account may be associated with a unique ID among other data management service entities. Authentication may include the data management service subject presenting credentials so that the IAM service 110 can verify the credentials. In some embodiments, authentication may be delegated to an external third-party identity provider, such as one using Security Association Markup Language ("SAML") or OpenID Connect, and / or Lightweight Directory Access Protocol ("LDAP"). As described in more detail below, subjects may include, but are not limited to, accounts, for example.

[0043] In certain embodiments, human and / or non-human users (e.g., programs, applications, and / or devices) may be bound to accounts that may be authenticated by the IAM service 110. In this manner, a subject may be associated with an account. However, in an access rule, a subject may be any of an account, an account proxy, an organization, a service, and / or a group. When a user and / or non-human user authenticates, it may be bound to an account (or, in the case of an API key, to an account proxy and then to an account). As described in more detail below, if the account is a member of an organization or group specified in the access rule, the subject may be granted or denied rights associated with the organization, group, and / or account.

[0044] A human user may authenticate themselves using a combination of an email address and password, a username and password, etc. In some embodiments, a non-human actor, such as a program and / or script, may authenticate themselves using a combination of an API key ID and an API key secret. Authorization may include determining what privileges a data management service object has and / or verifying whether the data management service object is permitted to perform a requested action.

[0045] IAM Service - Directory Component In some embodiments, the directory component 134 can maintain a directory database 112, which may be generally referred to as a "directory" in some examples herein, and can support an API for manipulating, managing, and / or querying the directory 112. For example, the directory component 134 can maintain a directed acyclic graph of data management service objects (which may be referred to as data management service entities in certain instances herein) in the directory 112, such as, for example, but not limited to, organizations, groups, accounts, folders, datasets, clusters, etc. Data management service objects can be associated with a "type." There may be built-in types of the IAM service 110, and objects of these types may be managed by the IAM service 110. In some embodiments, there may also be custom types that are managed by other data management service services, such as the catalog service and the SEE service 110.

[0046] To support custom types, the directory 112 can maintain a type registry. Data management services and applications, as well as customer written applications, can register entity types. In some embodiments, entities of the registered types can be created in the directory.

[0047] The catalog service may manage data sources, data sets, procedures, and / or workspaces (e.g., folders used to group virtual data sets managed by the DV service 114), which may be added to the directory 112 by the catalog service using the IAM service API. Once in the directory 112, these entities may be queried and access checked using the IAM service API.

[0048] 2 illustrates a non-limiting example of management of datasets using directory 112 consistent with certain embodiments disclosed herein. As illustrated, a user associated with an account "Harry" may issue a request to a catalog service 138, which may be a component of the DV service, that a new dataset "Dataset 2" be created. The request may comprise a variety of information, including, for example, but not limited to, identification information associated with the dataset, user, and / or request, a name and / or other identifier associated with the dataset, a type associated with the request and / or associated dataset, an indication of a parent and / or other organizational relationship associated with the user and / or account (e.g., the ACME organization, etc.), an identification of the user and / or account issuing the request, and / or any data included in the dataset.

[0049] The catalog service 138 can interact with the IAM service 110 and issue a custom entity creation request. If authenticated and / or authorized by the IAM service 110, the data set can be added to the directory 112 as a new data set object as a result of the successful custom entity creation request. As shown in the illustrated example, the root parent object of the new data set can be the ACME organization. The IAM service 110 can return a response to the custom entity creation request to the catalog service 138. The response can include, for example, but not limited to, a status of the response (e.g., “granted,” “created,” etc.), identity information associated with the data set, user, and / or request, and / or any other information related to the creation of the entity by the IAM service 110 in the directory 112.

[0050] FIG. 3 illustrates a non-limiting example of a query for a dataset using directory 112 consistent with certain embodiments disclosed herein. As illustrated, a user associated with account “Sally” can issue a query to data service 140 for dataset “Dataset2”. A catalog service can be used to create, update, and / or delete objects in directory 112. Data service 140 can be used to query datasets and / or invoke procedures. For example, as illustrated, data service 140 can interact with IAM service 110 to issue an access check request to IAM service 110. The access check request can include a variety of information including, for example, but not limited to, an identity of the subject associated with the request (e.g., “Account Sally”), an identity of the object of the request (e.g., “Dataset2”), and / or an identity of the requested access privileges (“query data”).

[0051] Directory 112 may include at least three objects: dataset 2, Mary's account, and rule set 1. Rule set "rule set 1" may specify account targets, dataset 2 objects, and "query-data" privileges for Mary, as well as a restriction that specifies that Mary is not allowed to view the column "email" in the dataset. IAM service 110 may manage objects in directory 112. In response to a query, IAM service 110 may return an access check response based on the query and the objects managed in directory 112. For example, IAM service 110 may return an access check response indicating whether Mary is allowed to query dataset 2 and / or any restrictions on such access (e.g., restricted from viewing values ​​in the "email" column of dataset 2), etc.

[0052] 1, other services of the data management services, such as the SEE service 116 and the TSDB service 104, can also manage objects having custom entity types in the directory 112. For example, but not limited to, components, workspaces, deployments, and / or vault objects of the SEE service 116, tables and / or workspaces of the TSDB, etc. can be managed by the SEE service 116 and / or the TSDB service 104. External applications can also define new entity types and create entities of those types. The IAM service 110 can manage the directory location of these entities, just as it does for internal entities.

[0053] Entity types defined by the IAM service's directory 112 may include, for example, but are not limited to, accounts, organizations, groups, applications, clients, privileges, privilege sets, rule sets, role authorizations, roles, policies, etc. Privileges may define what actions are possible on an entity. While there may be many privileges (e.g., predefined privileges) built into the data management service, custom applications may define their own privilege sets and privileges. These custom privileges may be used in rule sets and policies. Applications and / or services may then call the IAM service's 110 APIs to perform access checks on data management service objects, which then enforce these privileges and any associated restrictions. While other components and applications may also define privileges, in some embodiments, the directory 112 may specifically implement system-defined privileges. Data management service objects may be associated with a type and a unique ID, and may also have a set of attributes and metadata fields.

[0054] In some embodiments, governance objects, which may comprise rule sets and / or role authorizations, may be stored along with other data management service objects in the data management service directory 112. This, among other things, allows the data management service to associate governance with the governance objects themselves, providing the ability to control who can add, remove, and / or update governance objects.

[0055] IAM Service - Metadata Component The metadata component 136 can support attaching metadata to any data management service entities stored in the directory 112 and can provide APIs that allow for querying the metadata as well as searching for entities in the directory 112 by metadata. Object attributes can also be used in restrictions used in rule sets and / or role authorizations to enable rich access checks of datasets or other entities.

[0056] Data Virtualization The DV service 114 may enable fine-grained data access and governance for a diverse set of data sources using a common interface. In some embodiments, supported types of data sources may include, for example, but not limited to, SQL databases with JDBC drivers (e.g., MySQL, PostgreSQL, MS SQL Server, Oracle, Redshift, AWS Athena), TSDB, InCountry, Parquet files (e.g., stored in AWS / S3), etc. The DV service 114 may support joining (as in a SQL JOIN) heterogeneous data sources (e.g., SQL data sources and flat file data sources). It may also support joining data sources from different geographic locations. Supported data access interfaces may include, for example, but not limited to, SQL via JDBC and REST APIs. In some embodiments, the DV service 114 may also support ANSI SQL queries with a set of geospatial query functions. In various embodiments, the DV service 114 may include a catalog service component and a data service component.

[0057] DV-Catalog Service The catalog service can manage data sources, datasets, procedures, and / or workspaces, which can be governed using the IAM service 110 mechanism. Objects defined in the catalog service can be registered in the directory 112 and / or can call the IAM service 110 to perform access checks on these objects. The catalog service can facilitate connections to physical data stores and manage information about them. Users can use the catalog service to define different objects, including, for example, but not limited to, data source objects, physical dataset objects, virtual dataset objects, procedure objects, and / or workspace objects. Users can use rule sets and role authorizations from the IAM service 110 to grant privileges and specify restrictions on these objects.

[0058] The virtual datasets 122-126 may be derived from physical datasets stored in the hot and / or cold storage tiers 106-108 and / or other virtual and / or physical datasets. The virtual datasets 122-126 may query information from multiple datasets, which may be physical and / or virtual. In some embodiments, the virtual datasets 122-126 may include physical datasets from heterogeneous data sources. For example, the virtual datasets 122-126 may map to physical storage in the hot storage tier 106 and / or the cold storage tier 108. However, it will be understood that the virtual datasets 122-126 may alternatively and / or additionally be other virtual and / or physical datasets, including datasets (e.g., SQL database tables) that may not be associated with the TSDB service 104.

[0059] DV-Data Service Data services may be a component of DV services 114 that may run in conjunction with IAM services 110, SEE services 116, catalog services, and / or TSDB services 104. Data services may also run in conjunction with SEE workload access to data. In various embodiments, data services may perform, among other things: Implementing queries of the dataset and / or enforcing privileges and restrictions defined in the Catalog Service. · It provides a set of endpoints that allow users to query the data they have access to. · Accepts SQL queries via JDBC protocol. Supports ANSI SQL read queries and enforces them across data stores: For SQL queries that run against a single SQL-based data store, data services can, in certain implementations, provide security guarantees while adding relatively low latency. Supports a limited set of data modification commands for data stores that allow writes. Expose a set of REST API endpoints that allow querying of all datasets in the system. Responses may be formatted in JSON and may be streamed or fully rendered.

[0060] When resolving queries against datasets derived from other datasets, in some embodiments the data service may operate to ensure that the final query does not reveal further information from the results that would have been returned by the user who created the related dataset.

[0061] A particular virtual dataset may be queried using a "run as" feature that may enable sharing of a virtual dataset 122-126 with the ability to run queries on behalf of the grantor without exposing details of the underlying physical datasets. For example, User A (e.g., a dataset admin) may want to share a virtual dataset with User B (e.g., a business analyst) and give them permission to run queries on his / her behalf without actually sharing the physical datasets that were built. User B may not initially have access to any datasets, but may rely on User A to give them access. To do so, User A may specify a "run as" field using a catalog application or catalog service API. In this case, User B may not need to have privileges to access the underlying physical datasets on which the virtual datasets are built. They may, however, use the "run as" User A feature and run queries on the virtual datasets shared with them.

[0062] DV - Query Pushdown When a dataset query is executed, the DV service 114 may attempt to optimize query performance by pushing down the query to the underlying physical data sources. If the query cannot be pushed down, the query may be retried within the DV service 114. In some embodiments, query pushdown may be the default query execution behavior, but this feature can be configured within the catalog application user interface for individual datasets.

[0063] In general, pushdown queries can execute faster due to the ability of the underlying data sources to remove unnecessary data, especially if appropriate indexing is configured on the data sources. If the query cannot be pushed down to a single data source (e.g., joining data from tables across two or more data sources), the DV service 114 can facilitate the query execution itself, which may involve loading all the required data into memory and processing it. To accomplish this, the DV service 114 can execute parallel requests to retrieve the required data from the data sources as multiple partitions.

[0064] Secure Execution Environment The SEE service 116 may enable authorized data analysts / scientists / application developers to build and / or run models and / or applications (e.g., application 118) in an isolated sandbox environment 120. Data owners may analyze permitted data set(s) and maintain granular control over the data. The data management service SEE 116 may enable enterprises to use and confidently share such data with preferred partners and analytical experts to derive actionable business insights.

[0065] The set of applications that run in the SEE environment may not be limited, and there may be a set of defined frameworks and interfaces that can make it easier to create these applications. Analytical experts can deploy existing workloads from a container image sharing service and create new models and algorithms. The SEE service 116 can provide a managed service that insulates / shields most users from having to learn the details and / or specific requirements for implementing container-based workflows.

[0066] The SEE service 116 provides resources to a secure network-protected environment (e.g., sandboxed environment 120) and may be scalable for complex resource (e.g., memory, compute) intensive data processing. Network policies may control inputs, outputs, and / or any form of communication between deployments. In some embodiments, network policies may block or allow such communication. The SEE service 116 may be designed to facilitate collaboration between data scientists and data analysts to create and / or run models and / or algorithms in a secure environment and analyze authorized datasets.

[0067] Deployments running in the Data Management Service SEE environment can access the Data Services, and the SEE service 116 can be the gateway to the Data Management Service secure data. The Data Management Service can authenticate and authorize access by using the Data Management Service core API. This allows data owners to allow internal / external analysts and scientists to work with the data while controlling who can access what data.

[0068] Consistent with embodiments disclosed herein, the SEE service 116 may enable one or more of the following functions: · Setting access rules for queries of DV data such that access is allowed when operating within the sandbox environment 120 and not allowed outside the environment. A successfully authenticated and authorized data access request can gain access to the permitted data in the SEE service 116 . Privileged users may be allowed to create namespaces with the necessary resources. In certain embodiments, a namespace may include a data management service object that may represent an underlying namespace in a cluster. A namespace can be network protected. Jobs running within it can be protected from inputs and outputs. Easily deploy container images from public or private container registries (e.g. Docker Hub). · Running a server (e.g., a Jupyter notebook server) in the SEE 116 and building and importing programs, models, code files (notebooks), data files, and libraries that run on the server. Privileged users can give other users access to their programs (e.g. notebooks), components, and container images to work collaboratively. · Users can view their workload logs and monitor their program and model execution in the SEE application 118. SEE objects can modularize the deployment process and simplify replay.

[0069] SEE - Secure Data Access Data access may be granted via the Data Management Service following successful authentication and authorization by the Data Management Service Core Service.

[0070] SEE-governance The SEE service 116 can provide governed access control for different SEE objects. In some embodiments, this governance can follow the same and / or similar model as other data management service objects (e.g., using rules and targets). In some embodiments, containers under data management service management may be deployed in least privilege mode by default. This can help ensure that explicitly granted access rights are effective.

[0071] The governed objects in SEE 116 may include, for example, but are not limited to, one or more of the following: Namespaces. A namespace may include an object in the data management service directory that corresponds to a namespace in the underlying container orchestration environment (e.g., Kubernetes). Namespaces can be governed to control which users can run workloads within the namespace. A component (e.g., a container). A component may contain a SEE object that may encapsulate a container image, container parameters, and / or inbound / outbound connections. It may be a managed object, and thus its visibility, modifiability, etc., and its "use" may be governed. If a user does not have "use" privileges on a component, the user may in some implementations not be able to use it in a workload and therefore may not deploy a workload that uses this component. · Workload (e.g., a collection of components). A workload may include a SEE object that groups a set of components and may be a deployable unit. A user with appropriate privileges may "deploy" a workload, in which case the deployment may be created and successfully initiated. A deployment may be a governed object that allows a user with appropriate privileges to terminate, start, connect (e.g., over a network), or view the logs of components in the deployment. A user may, in some permitted circumstances (e.g., if the subject has edit privileges on the workload object), override parameters of a component when deploying a workload that references the component, in which case these overridden parameters may be present in the deployment object. Deployment (e.g., a running or already executed workload). When a workload is deployed, it can result in a deployment that can override parameters associated with the workload. Authorized users may be able to view logs associated with running and / or executed deployments. In certain embodiments, a deployment object may persist even after its associated workload has finished execution. In this manner, a user can query the log of a deployment to determine how it was executed and / or what it logged. · Vault. Credentials for container image registry access that may be required by the data management service to pull and run containers may be protected in a vault. Vaults may also be used for "sensitive" parameters to components. The SEE service 116 may retrieve these parameters from the vault and pass them to Kubernetes (e.g., as environment variables) so that the component (e.g., container) can use the parameter but not necessarily know the associated value. In some embodiments, there may be two "types" of vaults. One may have a well-defined schema so that the SEE service 116 can find the username / password for the container registry. The other may be free-form in the sense that there may be a less well-defined schema. The first kind may (optionally) be associated with an ImageSource, which itself may encapsulate the association between a docker registry URL and the vault where the credentials to that registry are stored. The second type may be used in "confidential" parameters for a vault, so that users cannot directly access and / or retrieve the value (e.g., unless they have full viewing privileges on the vault).

[0072] In certain embodiments, an entity may have knowledge of how to interpret entries in a vault if a less well-defined scheme is used. For example, in the case of a username / password in a container register, the SEE service 116 may have knowledge of how to find the username / password and how to provide them to the container register. In this case, the value may use a well-defined schema. In customer components that may contain sensitive information, any secret may be stored in the value, in which case the scheme for an external component (e.g., the SEE service 116) to identify may not be well-defined, as long as the component itself has knowledge of how to find the items in the value.

[0073] SEE-Isolation Environment The SEE service 116 can provide separate, isolated environments to run analytical models, external processes, and applications as workloads using separate managed clusters. Workloads may be isolated at the network level, and network policies may safely constrain both input and output. In some embodiments, the SEE service 116 can provide an "isolated by default" environment for running containers and networking using an API. There may be an abstract interface realized by the Kubernetes implementation, so users may not have direct access to Kubernetes and may therefore experience the SEE service API with their governance to manipulate the corresponding Kubernetes resources. This interface can allow users to create sandboxes 120 in which they can maintain specific behaviors / computations.

[0074] A user may be able to, for example and without limitation, enable the following: · Components that access DV services 114 but do not communicate with the outside world. · Controlled inbound access to these components, including rate limits on output, output volume limits, and historical cumulative access controls to ensure that DV Service 114 access to data does not allow export of raw data. Components with access to DV data (eg, data sets 122-126) generate outputs that are isolated from the outside world and restricted to being inputs to other components running in the SEE service 116 deployment.

[0075] In some embodiments, giving users direct access to the underlying Kubernetes cluster may allow customers to avoid "sandboxing." The SEE service 116 may control access to inbound and / or outbound connections and / or the creation of services that may provide load balancing and inbound access to components. The SEE service API may generate audit records that provide security accounting of operations. Through the SEE service 116, the data management service may allow users to write code and then run that code in an environment that may manipulate DV governance data, but not necessarily export that data outside of the container (without authorization).

[0076] Referring to the ACME enterprise example, large enterprises may engage with third-party service providers that specialize in building applications that address specific business needs, and may desire an environment that allows for relatively easy integration of these specialized applications in the context of their existing enterprise architecture.

[0077] ACME may want to protect their data assets by sharing them securely, and third-party service providers may want to secure their proprietary application code. In a non-limiting example, User A from ACME may be the data owner who controls access to the input data. Any other user (e.g., User B from the third-party service provider) cannot access the data unless User A provides explicit access via an API key. User B from the third-party service provider may create a component B that accesses the third-party application code via a container register. User B may create a component object with the correct configuration and parameter values ​​and protect the credentials provided for the service to access the component code. User A may not have the access credentials and therefore cannot access the proprietary code in component B, thereby protecting the intellectual property of the service provider. User B may, however, give User A the ability to "use" component B. User A may then include this component B along with his own components to perform arbitrarily complex data manipulations.

[0078] In the above example, user B's code and data can be protected. User A, who has limited access to the cluster (e.g., a Kubernetes cluster), may not inspect user B's component B, its data, any attached volumes, its logs, its images, its parameters, etc. In other words, user A may not have access to a third party's (e.g., user B's) code, data, and / or artifacts. On the other hand, if user A had full access to the cluster, user A could inspect the images, data, keys, parameters, etc. of user B's code. However, user A may use component B's functionality (e.g., a machine learning algorithm) to generate some "results" (e.g., output data). User A can control which "input data" component B could access. On the other hand, user B cannot access any data unless he is given permission to access it. Thus, untrusted code can run in an environment with less concern about data output from the system.

[0079] Time Series Database The TSDB 104 may include a modern, cloud-based, efficient, compressed, and / or scalable database. The store may be multi-tiered supporting fast, low-latency access to recent data and inexpensive storage of older data. Low-latency data may be stored in a key-value store (e.g., a Cassandra database). Long-term data may be stored as compressed data files in an object storage system (e.g., AWS S3, etc.).

[0080] In some embodiments, data in the TSDB 104 may be stored in a compressed and chunked format, and an index of these data chunks is maintained, thus allowing fine-grained access to the data. Data may be organized into namespaces, which may contain multiple datasets (logical representations), and each dataset may have multiple projections (physical representations of the data). In some embodiments, namespace objects may be used in connection with both the SEE service 116 and the TSDB service 104. In some implementations, a namespace may be conceptualized as a schema in a traditional relational database domain, a dataset may be conceptualized as a table, and a projection may be conceptualized as an index. Each dataset may contain at least one projection, called the primary, although more projections may be defined for a dataset.

[0081] TSDB - Storage Tier In particular embodiments, storage may be associated with two tiers: a hot storage tier 106 and a cold storage tier 108. In some embodiments, the hot storage tier 106 may store data in Apache Cassandra tables. The hot storage tier 106 may make ingested data available with minimal delay. In some implementations, a relatively small amount of recent data may be stored in the hot storage tier 106.

[0082] The cold storage tier 108 can store data in Apache Parquet files stored in AWS S3. In some implementations, it can provide low-cost and highly scalable storage, potentially at the expense of longer delays in available ingested data (and slightly slower data retrieval due to S3 read latency). Data can be added periodically to the cold storage tier 106 according to a defined time aggregation period. Data compression can also be configured to reduce data fragmentation. The compression process can merge separate data files of the same time aggregation period into one.

[0083] TSDB-behavior The TSDB service 104 can support data insertion, update, and deletion. Data may be ingested in any time order, but data belonging to the same partition may be ingested in time order to avoid data segmentation. In some embodiments, data can be ingested by sending individual records of a complete data file to its REST API. For governed data access, the data management service DV service 114 can be used. Additionally, the TSDB service 104 can provide its own data access REST API.

[0084] 4 illustrates a non-limiting example of a TSDB data management architecture 400 consistent with certain embodiments disclosed herein. As shown, architecture 400 may comprise systems, services, and / or components associated with hot and cold storage tiers. Architecture 400 may further include systems, services, and / or components shared between the hot and cold storage tiers, as well as systems, services, and / or components associated with canonical storage.

[0085] Data may be ingested through one or more ingestion layer 402 components, which may comprise, for example, but not limited to, a Kafka client, a REST API, and / or a bulk import module and / or interface. Data ingested into the data storage and management platform may be published to one or more partitioned topics, which in some implementations may include partitioned Kafka topics. In some embodiments, each message published to a topic may have a sequence number within an associated partition. For example, each message published to a Kafka topic may have an offset within a given Kafka topic partition, which may serve as a sequence number and / or indicator for various data management operations consistent with embodiments disclosed herein. In some embodiments, the data storage and management platform may expose a REST API that may enable external systems and / or services to insert data records into the platform.

[0086] The hot storage tier may comprise a streaming writer 404 and a hot data store 406. From each topic, data may be consumed by the streaming writer 404. In certain embodiments, the streaming writer 404 may be configured to detect which data partition an incoming data record belongs to and may store the record in the appropriate data partition key in the hot data store 406, which in some implementations may comprise a Cassandra key-value database. The streaming writer 404 may further detect new data partitions from the ingested data records, potentially redistribute the ingested data as needed (e.g., based on information included in the definition metastore 402), add data portions (as needed) to a data partition index 410 that may be shared between the hot storage tier and the cold storage tier, and then store the record with the new data partition key in the hot data store 406.

[0087] Definition metastore 408 can provide definitions related to namespaces, which can enable different users to operate on and / or process data in a particular table while operating in different namespaces. In some embodiments, definition metastore 408 can provide definitions regarding storage level and / or tier information for data. For example, definitions can be provided regarding whether and / or which data should be stored in a hot storage tier, a cold storage tier, both storage tiers, etc., retention periods for stored data that may vary depending on the tier in some implementations, update information for hot and / or cold storage tiers, criteria for data compression operations, etc. In this manner, information included in definition metastore 402 can help define the logical structure of data, how it should be partitioned by architecture 400, how it should be written to platform storage, etc. Management API 412, which can include a REST API, can be used to interact with and / or otherwise manage definition metastore 408.

[0088] The canonical storage tier may include a canonical store writer 414, a canonical store 416, and a canonical segment index. Data ingested into the data storage and management may be provided to the canonical store writer 414. The canonical store writer 414 may consume the received topic record data, process the data, and / or store the data in the canonical store 416. The canonical store 416, in some embodiments, may include a cloud-based storage service, such as, for example, but not limited to, AWS S3. Files written to the canonical store 416 may be associated with records added to the canonical segment index, which may provide index information for the records stored in the canonical store 416. Data stored in the canonical store 416 may be used in connection with various cold tier storage operations, partitioning and / or repartitioning operations, data backup operations, etc., as described in more detail below.

[0089] In some embodiments, the cold storage tier can include a canonical store crawler, a segment extraction service, a segment compression service, a cold data segment store 420, a data segment indexer, and a data segment index. Consistent with various disclosed embodiments, data stored in the canonical store 416 and / or index information included in the canonical segment index may be used to construct data records in the cold storage tier. For example, without limitation, the canonical store crawler and / or associated segment extraction service may interact with the canonical store 416 and / or the canonical segment index to access increments of data from the canonical store 416, potentially process the data (e.g., using a segment compression service), and store the data in the cold data segment store 420. As data is stored in the cold data segment store 420, the segment extraction service may interact with a data segment indexer service to generate one or more records in the data segment index 418 associated with the data stored in the cold data segment store 420.

[0090] In particular embodiments, the definition metastore 408 may contain information used by various systems, services, and / or components of the disclosed platform to determine which ingested topics should be recorded by the hot data storage tier and the canonical store (and, by extension, the cold data storage tier). For example, in some embodiments, the streaming writer 404 and the canonical store writer 414 may use information contained in the definition metastore 408 to determine which ingested data should be recorded to the hot data store 406 and / or the canonical store 416.

[0091] In various embodiments, the segment extraction service can store data in the cold data segment store 420 based at least in part on information included in the definition metastore 408. For example, the definition metastore 408 can include information regarding cold data storage tier data storage and / or update scheduling, which can include information regarding update periods, update frequencies, update data volume thresholds, etc. This information can be used by the segment extraction service to schedule data recording actions and / or updates from the canonical store 416 to the cold data segment store 420.

[0092] In various embodiments, the use of a canonical storage tier in conjunction with a cold storage tier consistent with certain aspects of the disclosed systems and methods can enable certain optimized data, processing, management, search, and / or query capabilities. For example, and without limitation, the canonical store 416 can store record data in a compressed format, but the partitioning and / or division of data and the use of time buckets in conjunction with the cold data segment store 420 can provide certain data processing, search, management, and / or query efficiencies not directly achieved by the canonical storage tier. Data stored in the canonical store 416 may be further used in connection with data restoration and / or backup operations and / or data repartitioning operations. For example, if data is deleted from the hot and / or cold storage tiers but remains stored in the canonical store 416, the data may be restored from the canonical store 416 to the hot and / or cold storage tiers.

[0093] The data read tier 422, which may comprise a read REST API, an adapter (e.g., a Calcite adapter), and / or a Spark data source engine, may interact with the streaming read API. When retrieving data from the platform, the streaming read API may be queried with relevant query information (e.g., identifying a data partition and / or a time period). The streaming read API may query the hot and cold storage tiers based on the identified data partition and / or time period. In some embodiments, the low-level data search component may apply filters to the fetched data. Records from different data partitions may be merged into a single result, and optional post-processing such as sorting or aggregation may be performed.

[0094] TSDB - Concept Overview Referring again to FIG. 1, the various components and / or elements associated with the TSDB service 104 may include, for example, but not limited to, one or more of the following: · Table − A set of data elements organized as rows of variable values. Variable - an element of a table identified by a name and with a determined data type (e.g. int, double, Boolean, string, etc.) that may operationally define a semantic interpretation for the value (e.g. used to extract a canonical timestamp value from a time axis variable) and optional metadata. Selector - a special non-null variable used in row keys. Row keys can have multiple elements, which may be called selectors. In general, selectors can be chosen to create reasonably large clusters of data, balancing the throughput of queries and the latency of filtering the required data. Data Partitioning Scheme - A set of selectors that are defined to determine how the table rows are organized and therefore the optimal data access pattern. Data partition - a subset of rows that have the same set of selector values. For example, a partition in a table for device data may be a single device, and the selector defined for the table is the device identifier. Projection - A replicated physical representation of data with different data partitioning schemes. Multiple projections can be defined for a table to accommodate different data access patterns at the expense of storage redundancy. Segment - a physical file stored in object storage that contains adjacent rows of data that belong to the same data partition. Rows within a segment may be ordered by time. Data within each data partition may be stored in multiple segments.

[0095] A user can define multiple tables in the TSDB service 104. In some embodiments, additional projections can be defined for each table for different data access patterns, which can enable efficient data access.

[0096] IAM Service - Managed Data Management Service Entities Data management service entities, in some embodiments, may belong to one or more of the following non-limiting categories: Subjects - Subjects may include actors using the data management service, such as accounts, organizations, and / or groups. Accounts may represent human users and / or programs or scripts that run independently of a user login session. Human users may be identified by user credentials, which in some embodiments may comprise an email address and password, or a username and password pair. Non-human users may be identified by API key credentials, which may include an API key ID and API key secret pair. Objects - Objects may include entities that a subject can act upon. Depending on the data management service, different entities may become objects. The data management service may expose objects. These services may include a catalog service (which may be, for example, a component of the DV service 114), a SEE service 116 (e.g., components, workloads, deployments, and / or namespaces), an IAM service 110 (e.g., organizations, accounts, folders, groups, privilege sets, privileges, rule sets, role authorizations, applications, clients, etc.), and a TSDB service 104 (e.g., tables and / or namespaces). · Privileges - Privileges can be associated with fine-grained operations that can be access checked. An operation can be any functionality that a service or application wants to govern. A single operation can be governed by either a single privilege or multiple privileges. Privileges can be organized into privilege sets, which can define a set of related privileges that can be used to govern a class of objects. · Groups - Groups can contain soft links to their member accounts, groups, and organizations. Folders - Folders can contain either soft-linked or hard-linked entities. Hard-linked entities can derive their access control from their parent folder, while soft-linked entities can derive it via hard links to other containers.

[0097] Data Management Services Entities - Objects A data management service object is the entity on which a data management service object operates. Depending on the data management service, different entities can be objects. For example, but not limited to, a data management service may expose the following objects: · Catalog Service 138: Data sources, datasets, procedures, and / or workspaces. · SEE Services 116: Components, Workloads, Deployments, Namespaces, and more. · IAM Service 110: Organizations, accounts, groups, folders, privilege sets, privileges, rule sets, applications, clients, etc. Some entities, such as accounts, organizations, and / or groups, can be both subjects and / or objects.

[0098] Data Management Service Entities - Target A particular data management service object, i.e., an account, can perform operations on other entities, called objects. Organizations and / or groups that can aggregate (e.g., directly and / or indirectly) other data management service objects, i.e., accounts, may be used in rules. Objects may include, for example, but are not limited to, one or more of the following: Account - An account can represent either a human or non-human actor in the system. Account subjects can be authenticated (e.g., by an authentication service). Once authenticated, they can be granted access to objects based on the evaluation of one or more access control rules. Groups - Groups may consist of a list of accounts, other groups, and / or organizations. Members of any named groups and / or organizations may be considered members of the top-level group. · Organization - An organization can contain members (e.g., accounts) and other sub-organizations. An account can be considered a member of an organization if it is a direct member of the organization or a member of any of its descendant sub-organizations.

[0099] As mentioned above, in some embodiments, groups and organizations may not be true subjects in the sense that groups and / or organizations may not be able to log in or authenticate themselves, but they may be considered valid subjects within a rule. In other words, it is possible to define rules that specify a set of accounts by naming a group and / or organization. When a group is used in a rule set, it may apply to all accounts that are direct or indirect members of that group. When an organization is used in a rule set or policy, it may apply to all accounts that are direct or indirect members of that organization.

[0100] Data Management Service Entities – Account A data management service administrator, which may be an organization administrator, may create an account when adding a user to an organization. An account may be a data management service entity that ties together a user and an organization and represents the user's membership in a particular organization. A user may log into the data management service with user credentials, which in some embodiments may include either the user's email address and / or password, or a username and / or password. For example, but not limited to, non-human actors, such as scripts, applications, programs, Jupyter notebooks, etc., may log into the data management service using an API key for authentication. Based on the provided credentials and / or key, an account may be located to associate with the authentication that may be the subject of any API calls made. In certain embodiments, when a non-human actor authenticates with an API key, the session may be tied to an account proxy, which may be tied to an account. When a human user authenticates, they may be tied directly to an account.

[0101] The IAM service 110 may generate access tokens that may be provided to applications or users when they log in. A data management service user may have multiple accounts, each in a different organization. When a user with multiple accounts wants to authenticate with the data management service, the user may specify the organization they are logging in with. Multiple accounts may have separate passwords, so authentication with an email or a particular username and password pair may uniquely identify an organization. In some embodiments, this may eliminate the need to specify an organization during login.

[0102] In various embodiments, accounts may be objects that define particular users and associate them with organizations. Accounts may in some implementations belong to a single organization (although in other embodiments an account may belong to multiple organizations), may belong to one or more groups, and / or may be granted privileges on objects. Accounts may be stored in the data management service directory as child objects of organizations.

[0103] Data Management Service Entity – Account Proxy One or more account proxies may be associated with an account. Each account proxy may inherit all privileges granted to the owning account or may be set up to have a subset of those privileges. Each account proxy may have one or more API keys tied to it. The API key credentials associated with an API key allow a non-human actor (e.g., a program, script, etc.) to authenticate itself and obtain the privileges associated with the account proxy's privileges. In some embodiments, an account proxy may be stored in the data management service as a child object of the associated account.

[0104] Data Management Service Entity – API Key An API key can represent credentials used by non-human data management service actors, such as programs and scripts, to authenticate themselves to the data management service. When such a program or script authenticates using the API key credentials, the resulting access token can represent the account proxy to which the API key is bound. That account proxy can then result in privileges being granted to the non-human actor. API keys can be stored in the data management service directory as child objects of the associated account proxy.

[0105] Data Management Services Entities – Organizations An organization may include basic data management service entities used to organize accounts, groups, applications, datasets, deployments, and / or other entities such as TSDB tables. An organization may have multiple child organizations, which may be referred to as sub-organizations in some instances herein, but may have a single parent organization. In various embodiments, a parent organization may manage and govern its sub-organizations.

[0106] An organization can include accounts, groups, and other organizations (e.g., internal divisions, subsidiaries, vendors, or business partners). As used herein, an organization does not necessarily represent a single company, but rather may be an abstract term that represents a group of entities with which data and applications are shared. Organizations can be nested in a hierarchy, so that child organizations are sometimes referred to as suborganizations. Members of a suborganization may be members of a parent organization in some implementations.

[0107] For example, within the ACME organization, there may be sub-organizations such as Sales, Marketing, Purchasing, HR, Finance, etc., and within the Purchasing sub-organization, there may be sub-organizations corresponding to the EMEA and APAC regions. Each top-level organization may be treated as a tenant in a particular deployment. Login-related configurations such as, but not limited to, password policies, multi-factor authentication ("MFA") requirements, external identity providers, and / or the like, may be restricted to the tenant organization.

[0108] An organization stored as a direct child of a root object in a data management service directory may be referred to as a tenant.

[0109] Data Management Service Entities - Identity Providers An identity provider may represent an external service that acts as an authentication service. These may use protocols such as, for example, but not limited to, OpenID Connect, OAuth2, SAML, Active Directory, etc. to authenticate users. A tenant organization may define one or more identity providers that can be used by members of the organization to authenticate themselves to the data management service.

[0110] Data Management Service Entities - Applications A data management service application entity may be created to represent a data management service application that hosts a client. The client may represent a web application that supports logging into the data management service. The application entity may also be useful to support defining application specific privileges for feature authorization checks.

[0111] Data Management Service Entities – Clients The data management service client may enable support for web applications to obtain access tokens from the IAM service 110. When a client logs in, it can authenticate itself with the IAM service 110 using an ID and secret. In some embodiments, the data management service client may represent a client in the OAuth2 and / or OpenID Connect protocols.

[0112] Data Management Service Entities - Groups Data management service groups can contain groupings of accounts, other groups, and / or organizations. Data management service groups can reference their members using soft links. When a group is added to a group, the members of the child group can be considered members of the parent group. When an organization is added to a group, the members of the organization and any of its suborganizations can be considered members of the group. Groups can be referenced in privilege authorization to provide privileges to many accounts at once. When a group (or organization) is used as a target in a rule set or role authorization, changes to those governance objects (rule set or role authorization) can be applied to the members of the group (or organization).

[0113] Data Management Service Entities - Privileges Privileges may be associated with fine-grained actions that may be access checked. An action may be governed by one or more privileges. Although embodiments of the disclosed data management service may define a number of governed actions, such as delete, list, query, execute, or view, application developers may also be free to define custom actions (e.g., by defining associated custom privileges).

[0114] Different privileges may be used to manage data management service entities, for example but not limited to: · To modify attributes of an entity, such as the name or image of an organization, it may be necessary to have modify privileges. · To add a member to an organization, you may need to have the add-child privilege.

[0115] A privilege may be identified by a unique ID and may be associated with one or more actions to which an application wants to govern access. For example, an application may be a web application with two pages, user and admin. If you want to allow only a subset of users to access the functionality on the admin page, you can create a privilege called "admin" and grant that privilege to administrators.

[0116] Data Management Service Entities - Privilege Sets A privilege may appear in the directory 112 as a child of a privilege set, which may be a container that may have any number of privileges and may reside anywhere within the directory 112.

[0117] Some privilege sets may be standard in the disclosed data management services. If none of the standard privileges are appropriate for an application, a user may create custom privileges. If a user requires custom privileges, a new privilege set may be created and the new privileges may be defined and included within the privilege set.

[0118] In some implementations, the system privilege set may have an ID prefix. For example, but not limited to, the system privilege set may include, but is not limited to, "sys:directory", "sys:governance", "sys:application", "sys:security", "sys:audit", "sys:apikey", "sys:account-proxy", "sys:custom-entity", "sys:catalog", "sys:executor", and / or "sys:storage".

[0119] The system privilege set "sys:directory" may include, for example, one or more of the following non-limiting examples of privileges: list, view, modify, delete, add-children, delete-children, and / or modify-read-only attributes.

[0120] These privileges may be used by the IAM service 110 directory operations when a user manipulates the data management service directory hierarchy. For example, to modify an organization or application object, a user may need to have the "sys:directory:modify" privilege on the object.

[0121] Other data management service components may use other privileges. For example, a data service may use privileges from its "sys:catalog" privilege set: write-data, call-mutation, inspect-resource, describe-execute, privilege-read-access, query-data, query-datasource, call-query, full-view, delete-datasource, mutate-data, insert-data, manage-query, delete-data, manage-execute, mutate-datasource, update-data, read-data, inspect-execute, call-procedure, and / or privilege-write-access.

[0122] Applications can define their own privileges (and privilege sets) and then perform access checks by calling the / security / checkIAM service endpoint.

[0123] A custom privilege set can have a unique ID that is a DCE / RPC UUID ("GUID"). A non-limiting example of a custom privilege set unique ID may be in the format: 3475f778-0d3e-4dcf-8237-3f1b022deffe. A privilege within a privilege set can have a short string as a name, such as, for example, but not limited to, admin.

[0124] The data management service ID of the privilege may include a concatenation of the privilege set ID and the privilege name. In the above example, the ID of the new privilege may be 3475f778-0d3e-4dcf-8237-3f1b022deffe:admin.

[0125] Data Management Service Entities - Policies A policy can contain an IAM service object, which is a named list of partial rules. A partial rule can specify the binding of privileges to objects, depths, and / or optional restrictions and restriction combinators. The privileges in a partial rule, like those in a rule, can specify for each privilege whether the privilege is allowed or denied.

[0126] Data Management Service Entities – Roles A role may be a named list of policies. Roles can exist as separate objects from policies to allow fine-grained specification of policies and to allow combining them in different ways in different roles. A policy can appear in multiple roles.

[0127] Data Management Service Entity – Role Authorization A role authorization may include a named binding between a subject and a role, attached to a particular object in the directory. A role authorization attached to object A in the directory may govern object A, and potentially objects below A in the directory, according to a depth field in the rules constructed from policies in the role referenced by the role authorization. As one non-limiting example, a depth of 0 may indicate "just this object," while a depth of 1 may mean "this object and its immediate children."

[0128] Data Management Service Entity-Covered Governance Rule Set Subject governance rule sets may be similar to rule sets, but may be bound to account proxies and used to provide a subset of privileges to those account proxies. They may be used in subject governance where privileges may be limited to a subset of privileges granted to the account that owns the account proxy.

[0129] Data Management Service Entity - Role Authorization Subject role authorizations may be similar to role authorizations, but may be tied to account proxies and used to provide a subset of privileges to those account proxies. They may be used in subject governance, where privileges may be limited to a subset of the privileges granted to the account that owns the account proxy.

[0130] DV Service Entity - Data Source In some embodiments, a DV data source may represent a source of data that may enable access to data from, for example, but not limited to, a SQL database (e.g., MySQL, PostgreSQL, MS SQL Server, Oracle, Redshift, AWS Athena), a TSDB table, InCountry, Parquet files, etc.

[0131] DV Service Entity - Data Set In some embodiments, DV datasets (e.g., virtual datasets 122-126) may include either physical datasets or virtual datasets. A physical dataset may represent, for example, an underlying table in a SQL data source, a TSDB table, a Parquet file, etc. A virtual dataset 122-126 may represent data from multiple physical datasets and other virtual datasets. Data from virtual datasets 122-126 may combine data from multiple datasets (e.g., as in a SQL JOIN).

[0132] DV Service Entity - Procedures In some embodiments, DV procedures may represent functionality similar to SQL stored procedures. DV procedures may allow for querying, updating, and / or deleting data from multiple datasets and / or multiple data sources.

[0133] DV Service Entity - Workspace In some embodiments, a DV workspace may represent a grouping construct, such as a folder in a file system or a folder in the IAM service 110, in which data sources, data sets, and procedures may be stored. Governance applied to a workspace may apply to all entities contained within the workspace.

[0134] SEE Service Entity - Namespace A namespace can contain an environment that allocates the actual resources (e.g., compute and memory) configured in its definition. Users can run models and container images in a namespace and can be charged for the resources reserved and used.

[0135] Privileged users can create, terminate, and / or delete namespaces. If there are jobs running in the namespace during termination, the user can be warned, but users with special privileges can, in some implementations, force the termination of running jobs in the namespace.

[0136] A privileged user can delete a namespace. Once deleted, the namespace may not be recovered, but may be maintained in system records as long as there is at least one object that depends on it. Unreferenced deleted namespaces may be purged periodically.

[0137] SEE Service Entity - Vault A vault can be designed to securely store sensitive information, known as assets, in key-value pair format. Privileged users can use and / or modify assets and may require special privileges (e.g., full view) to see asset values. A vault can be shared with other users to let them use the assets but keep the values ​​confidential.

[0138] A vault can be used to store authentication information (e.g., username, password, etc.) for a private container registry of containers. Such a vault can be used with an image source to access the container registry. A vault can also be used to store sensitive user inputs (e.g., parameters) for programs and / or models. Users with usage privileges on the vault can use these parameters but do not necessarily know their values.

[0139] In some embodiments, a privileged user may delete a vault. Once deleted, the namespace may not be recovered, but will be maintained in the system records as long as there is at least one object that depends on it. In certain embodiments, deleted vaults that are not referenced may be periodically purged. Objects that depend on a deleted vault may continue to run until they need to access / use the deleted vault. However, a user with special privileges may force the destruction of such objects when deleting a vault.

[0140] In certain implementations, a vault can be shared with other users without revealing its values. Thus, key-value pairs can make these sensitive values ​​available to components that need them without exposing them to users of those components. For example, a vault can contain a username and password for a private container registry that holds images needed by a component. Users of that component can source the required container image from the private registry and deploy it without exposing the registry credentials to those users. In some embodiments, this may be optional, but may be recommended if sensitive information is used within the component. In such cases, vault values ​​can be specified as component parameters, exposing the vault values ​​as environment variables that can be used by programs that run within the component.

[0141] SEE Service Entity - Image Source An image source can be a connection source for a container registry. It can include a URL, an image name, and / or a user email. It can include a vault (that stores secret credentials) for connecting to the container registry. Image sources can be shared with other users as needed.

[0142] A privileged user may delete an image source. Once deleted, in certain embodiments, the namespace may not be recovered, but may be maintained in the system records while at least one object that depends on it exists. Unreferenced image sources may be periodically purged. Objects that depend on a deleted image source may continue to run until they need to access / use the deleted image source. However, a user with special privileges may force such objects to be destroyed when deleting an image source. A user may use an image source to store connection details for a container registry.

[0143] Container images can be deployed as components within workloads in the SEE service 116. A user can store container images associated with components in an access-controlled container registry or can be interested in deploying images from a public repository. A user can save container image repository connection details in an image source and keep the credentials secret in a vault. Once defined, these details can be used multiple times to connect and pull container images whenever needed.

[0144] SEE Service Entity - Components A component can represent a program, a model, and / or a container image. A component can use an image source (e.g., container registry connection details) to pull a container image from a registry. Because the container registry connection details can be separated in the image source object, the image source object can be reused by multiple components. If a container image requires sensitive user input for execution, these parameters can be retrieved from the vault.

[0145] Components may be deployed within a workload and may include configurable parameters that allow pre-deployment customization. Components may be governed objects and users may be granted privileges to view, modify, or use components in their workloads.

[0146] A privileged user can mount a volume on a component to store intermediate and / or final results of data processing and / or read programs, code files, data files, and / or libraries stored on the volume. A component can be linked with two or more volumes, and similarly, a volume can be linked with two or more components.

[0147] A component may require an outgoing connection to a different service. A privileged user can add an outgoing connection configuration. A component may require one or more incoming connections. A privileged user can add an incoming connection configuration. A SEE workload can hold one or more components together, components may represent the same container image, use different input parameters or different container images that the user wants to run together.

[0148] A privileged user may delete a component. Once deleted, the namespace may not necessarily be restored, but it may be retained in system records while there is at least one object that depends on it. Unreferenced components may be purged periodically.

[0149] Objects that depend on a deleted component may continue to run until they need to access / use the deleted component, however a user with special privileges may force such objects to be destroyed when deleting a dependent component.

[0150] SEE Service Entity - Workloads According to embodiments disclosed herein, a workload may include a unit of work that represents one or more components (e.g., programs, container images, etc.) bundled together. Components in a workload may access data via data management services. Component parameters may be overridden in a workload without affecting the original value. Privileged users may create, deploy, modify, and / or delete workloads. Once deleted, in some embodiments, the namespace may not be recovered, but may be maintained in system records as long as there exists at least one object that depends on it. Unreferenced workloads may be periodically purged. Objects that depend on a deleted workload may continue to run until they need to access and / or use the deleted workload. However, a user with special privileges may force the termination of execution of any deployment that uses this workload.

[0151] SEE Service Entity - Deploy A deployment may comprise a workload that is deployed in the SEE environment. A privileged user may create, modify, start (e.g., run), terminate (e.g., stop), and / or delete a deployment. When deleted, a namespace may not be restored, but may be maintained in system records as long as there is at least one object that depends on it. Unreferenced deployments may be purged periodically. Deployment logs may be available for troubleshooting purposes, and the execution of deployments may be monitored in the SEE application.

[0152] SEE Service Entity - Volume In various embodiments, a volume may represent persistent storage governed by a data management service. It may be used to store programs, code files, data files, libraries, and / or program output. Volumes may be attached to components, workloads, and deployments deployed within a namespace and may persist beyond the destruction of a namespace. A privileged user may delete a volume. Once deleted, a namespace may not be recovered, but may be maintained in system records as long as there exists at least one object that depends on it. Unreferenced deleted volumes may be periodically purged. Objects that depend on a deleted volume may continue to run until they need to access and / or use the deleted volume. However, in some embodiments, a user with special privileges may force the destruction of such objects upon deleting a dependent volume.

[0153] TSDB Service Entities - Tables In various embodiments, a TSDB table can represent a set of data elements organized as rows of variable values. The TSDB table can be used as a DV data set. The TSDB table stores data that can include a field that represents a time.

[0154] TSDB Service Entity - Namespace In various embodiments, a TSDB namespace can provide an organization of TSDB tables. A namespace can include multiple data sets and can be conceptualized as a schema.

[0155] Data Management Services Governance Consistent with certain embodiments disclosed herein, the IAM service 110 may support an advanced access control system that allows rule sets and / or role authorizations to be attached to data management service objects. These rule sets and / or role authorizations may specify rules that govern the data management service objects named within the rules.

[0156] Governance as used herein may include the enforcement of rules and may include determining whether a subject has a given privilege on a given object. In some embodiments, a rule set or role authorization attached to a data management service object may govern only objects "lower" in the directory hierarchy.

[0157] In various embodiments of the disclosed data management services, the access checks may include the following specifications: ·subject. Objects. ·Privileges.

[0158] For example, a subject may be a user represented by an account that has an access token from an application or service, and an object may be any data management service entity.

[0159] In some embodiments, within a data management service access check may be an object. A user (or more generally an authenticated subject) may attempt to perform an operation on any data management service entity designated as an object. However, to perform an operation on a particular object, the user (and / or associated subject and / or account) may need the appropriate rights and / or permissions to do so. The IAM service 110 may provide an API that allows the data management service or any application to perform an access check to determine whether a specified subject is authorized to perform an operation (specified by a privilege) on a specified object.

[0160] An access check can return either "true" or "false" to indicate whether the subject has the privilege associated with the specified object, along with an optional list of restrictions that can be used to restrict aspects of the operation. For example, a restriction can inform the data service that the subject can only see a subset of columns in a dataset.

[0161] To attach a governance object (e.g., a rule set and / or role authorization) to an object in a data management service directory, a user may require the "sys:governance:add-child" privilege on the object. Without such privileges, a user may be denied the ability to govern the object. The attached governance object may govern any object below its attachment point in the directory, according to the value of the governance object's "depth" attribute. For example, a depth of 0 may indicate that the governance object governs only the object to which it is attached. A depth of 1 may indicate that governance applies to the attached object and its direct children. A depth of -1 may indicate that governance applies to any descendant objects in the directory. Thus, depth may represent the scope of governance of the attached governance object.

[0162] For example, if a user is an organizational administrator, they may be able to govern the organizational objects that represent their organization. They can also govern any objects created "under" those organizations. Thus, a user can create data management service objects so that the user can govern these same objects.

[0163] Any object can be specified in an access check. For example, an account may have custom "admin" privileges on a custom application that represents their (web) application. If so, you can use the application object as the object of the access check.

[0164] Using application objects may be appropriate for so-called "functional" privileges or operations where the particular object does not matter in the access check. This may enable or disable a particular feature of the program in question. The opposite may be true for privileges on data sets (e.g., data management service entities exposed by a catalog service), where there are many data sets and they may be governed separately.

[0165] In at least one non-limiting example, a data service may perform an access check to answer the question, "Does this account have the 'sys:catalog:query-data' privilege for this particular dataset?"

[0166] If a user requires specific object-level access control, the user can create a data management service entity. The user can use existing data management service entity types, attach rule sets and / or role authorizations to these objects, and perform access checks by specifying the objects in the access check API calls.

[0167] As a non-limiting example, consider that a user is an administrator for an application called MyApp. There may be a data management service entity called "MyApp" with a data management service ID, which may be a unique ID. The administrator may define a rule set or role authorization for this application, specifying rules such as allowing accounts and / or groups of accounts the above-referenced "admin" privileges on application objects. Given these rules, the user's application may use the IAM service " / security / check" API to check the access of a given account to perform "admin" operations on the application.

[0168] Many data management services entities (e.g., organizations, groups, folders, etc.) can contain other entities. The relationship between a parent container and a child entity may be referred to as a link. In some embodiments, links may be of two types: soft and hard. A child entity linked to its parent via a hard link may be governed by the rule set attached to the child and the rule set of its parent. This may be true all the way up the hierarchy to the root entity. The root entity may be the object located at the top of the hierarchy in the data management services directory. A child entity linked to its parent via a soft link may be considered a child of the parent, but governance may not consider any parents reached via a soft link.

[0169] Governance in Data Management Services An object inserted into the directory may be a governed object. A governed object may be an object where rules define the actions that can be performed on the object. Governance is a collaborative effort between services that manage objects. For example, data sets and / or data sources may be managed by a catalog service. These objects may be inserted into the directory 112 (e.g., hierarchy) as children of a parent object. For example, data sources may be hierarchically located under an organization object, a folder object, and / or any number of other locations. Due to the hierarchical nature of the directory 112, there may be a unique path from the root of the directory to any object in the directory 112. Governance may be "applied" to an object in the directory 112 or any parent along the path to the root. Any governance that exists outside of the direct path from the object to the root may not have any effect on the governance of that object.

[0170] In some embodiments, governance may be applied using one or more objects, which may include rule sets and / or role authorizations. Regardless of which mechanism is used, the effect may be the same: rules may be created, each granting one or more privileges from a single privilege set to a single subject, single object. Privileges within a rule may be allowed or denied. If an object hierarchy exists, privileges may be inherited by objects below the object referenced in the rule.

[0171] The granting of a privilege may be accompanied by restrictions. Restrictions may impose additional constraints on the access granted by a given privilege, and these restrictions may be enforced by the service that manages the object. Thus, for a data set, it may be a data service (which in some embodiments may be a component of the DV service 114) that enforces the restrictions. A rule may also have a depth, which may specify how far down in the directory hierarchy, starting from the object specified in the rule. A depth of 0 may specify that the rule grants or denies privileges only to the effective object. A depth of 1 may specify that the rule applies to the effective object and its immediate child objects. A depth of 2 may affect privileges for the effective object, its children, and its children's direct children. A depth of -1 may mean that the privilege applies to descendant objects starting from the effective object.

[0172] The object specified in the rule may or may not be the same as the effective object. The location in the directory 112 where the rule is attached may be called the attachment point of the rule. The attachment point may be the same as the parent of the rule object (e.g., the rule set and / or role authorization).

[0173] If a rule is attached directly to the object referenced in the rule (as is the case for rules applied by IAM and catalog applications), the effective object may be equal to the object. In that case, the attachment point may be the same as the object specified in the rule. However, if the rule specifies an object above the attachment point, the effective object may be narrowed to the attachment point.

[0174] Multiple governance objects (e.g., rule sets and / or role authorizations) may be attached to the same object, each specifying a separate priority. A governance object with a lower numerical value for its priority can take precedence over governance objects attached to the same object with a higher numerical value. For example, if an object has two child (attached) governance objects, RuleSet1 and RoleGrant1, with priorities 0 and 1, respectively, the rules in RuleSet1 can take precedence over the rules in RoleGrant1.

[0175] As detailed above, governance objects can be attached directly to an object and / or to any parents on the path to the root of the directory. These governance objects can affect governance over an object. Governance objects higher in the directory hierarchy can take precedence over governance objects lower in the directory hierarchy. This can reflect the ability of an organization's administrators to impose rules on subordinates within the organization. Governance objects can be implicitly associated with a level, which is a level in the directory hierarchy. The root can have a level 0. The immediate child of the root can have a level 1. Governance with a numerically smaller level can take precedence over governance with a numerically higher value.

[0176] rules Consistent with embodiments disclosed herein, a rule may specify one or more of the following: Target ID. Object ID. · Privileged IDs (multiple but within the same privilege set). · Allow / Deny flag. Depth. ·limit. · Restriction Combinators - Used by services and applications to combine restrictions.

[0177] Rules may be specified as fields within a rule set, rather than as data management service entities stored separately within a directory.

[0178] Partial Rules Partial rules may provide essentially the same functionality as rules and may be composed of the same fields as rules, except that they may not include a subject. Partial rules may be specified in policies, which may be named in roles. Roles may be specified in role authorizations along with subjects. For this reason, subjects may not be required in partial rules, since they are specified in role authorizations that indirectly grant the privileges in the partial rules of the associated policy.

[0179] permission Various embodiments of the disclosed data management services may provide a unified data governance, data virtualization, and / or secure computing environment. Using data management service objects, administrators can build a governance layer to enable organizations to securely share data with internal and external stakeholders. Data management service objects may include accounts, organizations, groups, and / or applications.

[0180] The data management service can manage hierarchical role-based relationships between individual user accounts and account groups. It can govern access to resources through the application of rule sets, role authorizations, privileges, and / or restrictions. The data management service can also be an application for defining custom objects and privileges. These custom objects can be stored and managed within the data management service, but their enforcement can be the responsibility of the application itself.

[0181] In particular embodiments, the data management service may provide a chain of trust between users and endpoints within the enterprise and partner networks. The data management service's IAM service 110 can govern access based on privileges and can include security, directory, and metadata services.

[0182] Permissions may include, for example, but are not limited to: · A set of security rules configured by a data management service administrator, embodied in a rule set. Roles, which include view, edit, manage, and admin, and are embodied by role authorization.

[0183] The following table lists non-limiting examples of roles and their associated meanings.

[0184] [Table 1] TIFF2024536689000002.tif69170

[0185] Roles can be granted to any subject within the data management service, including accounts, groups, and organizations. In various embodiments, roles may be applied to subjects in different ways, including, for example, but not limited to, the following: A role granted to an account may, in some embodiments, apply only to that account. A role granted to a group is applicable to all members of that group. A role granted to an organization can be applied to all accounts within that organization and any of its descendant suborganizations.

[0186] If a user has been granted privileges on an object through multiple authorizations, such as directly on the user's account and through membership in a group, the resulting permissions on that object may be the union of the granted permissions.

[0187] For example, assume Account A is a member of the Administrators group. If Account A is granted the view role on Object Z, but Administrators' group is granted the manage role on Object Z, then Account A may have the privileges allowed by the admin role on Object Z.

[0188] In some embodiments, separate roles may govern one or more of the following non-limiting examples of behaviors: View names and other non-confidential information about data sources. · Redact names and other non-confidential information about data sources. Delete the data source. View and edit information, including data source access credentials. · Add / remove physical datasets from a data source. ·Manage permissions to data sources.

[0189] Governance Objects and Subjects Embodiments of the disclosed systems and methods can support two primary forms of governance: access control lists ("ACLs") and role-based access control ("RBAC").

[0190] An embodiment of the disclosed data management service may implement ACL with rule sets and RBAC with role authorization. Rule sets may explicitly specify a list of rules. Role authorizations, on the other hand, may reference roles, which in turn may reference policies. Policies may encode partial rules. The advantage of using rule sets may be that they are simple, all-in-one, and suitable for "one-time" governance. Role authorizations, on the other hand, may involve planning, since appropriate policies and roles may need to be defined. However, the advantage of role authorizations is that there may be one place (the policy) where partial rules are defined, and all role authorizations for a particular role indirectly reference these partial rules. This may allow for modification of the policy, and all role authorizations that reference the policy (indirectly through the role) will automatically have the effective governance updated.

[0191] Audit Services In some embodiments, the audit service may capture security audit records from various data management services and store them securely in a data store, which in some embodiments may comprise an append-only data store. These audit records may capture the identification, authentication, authorization, and / or security checks performed by the IAM service 110. These audit records may also capture the subject, object, privileges, and any restrictions from any access checks, whether performed internally by the IAM service 110 or on behalf of an external data management service (e.g., the DV service 114, the SEE service 116, and / or the TSDB service 104). In some embodiments, the audit service may capture audit records for both successful and failed operations (e.g., authentication, authorization, and / or access checks).

[0192] In some embodiments, data management services (e.g., the IAM service 110, the catalog service, the data service, the SEE service 116, and / or the TSDB service 104) can securely transmit audit records to the audit service. In particular embodiments, such audit records can be tagged with a "component name" that identifies the service from which the audit record was obtained.

[0193] In some embodiments, the audit records capture sufficient security-context information, which may include the subject of successful or unsuccessful actions and the identity of the object. A set of audit records associated with a user's actions can provide a security administrator with insight into the actions performed by a given user and / or against a particular object. Authentication and authorization decisions performed by the system may be logged to an audit service.

[0194] Audit log entries can include a transaction ID that describes the sequence of operations and isolates the action or query that was performed. In some embodiments, the audit service can ensure that these log entries are not modified.

[0195] In certain embodiments, administrators can view the logs in an audit application. In some embodiments, because viewing the audit log can be a relatively privileged operation, the data management service may grant the privilege to perform this operation to a limited set of users. The audit log viewer interface may allow for searching of audit log entries by, for example, but not limited to, time, scope, object, and user.

[0196] Examples of governing data from external sources In certain non-limiting examples, some embodiments of the disclosed systems and methods may be leveraged to provide a data governance solution for data ingested into a data management service from one or more external sources. For example, but not by way of limitation, embodiments of the disclosed systems and methods may enable the ingested data to be manipulated and / or otherwise transformed in a secure and / or sandboxed environment, enabling how the derived (e.g., transformed) data generated can be protected and / or governed. The resulting data may be queried by authorized parties according to the governance rules established for the derived data.

[0197] The external data may be generated by a variety of active sources, including, for example, but not limited to, wind turbines, nuclear reactors, factories, automobiles, IoT devices, and / or the like. The data may include time series data having records generated by the data sources associated with timestamps. The data may be ingested into the TSDB service 104 and segmented by the TSDB service 104 by time and / or other data attributes. The data may be initially available in the hot storage tier 106 and eventually migrated to the cold storage tier 108. It may be queried by authorized users via APIs associated with the TSDB service 104 and may be subject to governance provided by the data governance components of the data management service on the TSDB tables and / or via the DV service APIs.

[0198] Programs and / or applications can be executed within the sandbox 120 associated with the SEE 116 and can gain direct access to the TSDB ingested data, which can be subject to governance rules. In some embodiments, the sandbox 120 can prevent this data from being exported outside of the sandbox 116. For example, but not limited to, machine learning algorithms and / or programs leveraging proprietary models can be executed within the SEE 116 and / or sandbox environment 120 to derive new datasets from the ingested data. The derived datasets can be stored using APIs associated with the DV service APIs to create new datasets in the data management service. Governance over these datasets consistent with various aspects of the disclosed embodiments can ensure that only appropriate subjects can query the derived data. APIs associated with the DV service 114 can be used by authorized users to query the derived data, subject to governance over the datasets, which can include restrictions on row and / or column values ​​returned.

[0199] Users manipulating and / or querying data from the TSDB 104 and / or DV service 114 may be identified and / or authenticated via the IAM service 110. Access checks performed by the DV, SEE, and / or TSDB services 114, 116, 104 may be performed using an access check API provided by the IAM service 110. Allowed and denied access requests may be audited by an audit service. A security administrator with sufficient privileges may inspect audit records generated and / or maintained by the audit service. A security administrator with sufficient privileges may inspect audit records maintained by the audit service.

[0200] In various embodiments disclosed herein, the integration of separate components and / or services within the disclosed data management services may provide a combined governance model used by the integrated components and / or services. The object directory 112 may hold objects, objects, and / or rules and may provide a consistent and / or secure model for manipulating these objects. The SEE sandbox 116 may ensure that data manipulated within it is not exported outside of the environment. Ingestion entities and / or objects associated with systems (e.g., wind turbines), software that manipulates data within the SEE sandbox 116, and / or users querying the data set may be represented in the shared directory 112. Governance rules used to protect data and restrict access to the data may be stored in the directory as well. The result may provide a consistent and secure representation of entities providing a secure data storage and / or management solution.

[0201] Examples of Application and / or Client-Based Access Control In various embodiments, access control may depend at least in part on the application and / or client making the query request. As explained above, within the data management service, an application object may represent an application (e.g., a web application) that may call the data management service API after obtaining an access token from the IAM service 110. A client object may represent a client (e.g., an OAuth2 client) that is used to authenticate a user to the application. In this manner, an application may be associated with multiple clients that may not match the client device used to access the application. This client may be used during the authentication process, and a client ID and / or client secret may be used to authenticate the endpoint itself (e.g., an OAuth2 client) and are precursors to authenticating a user wishing to utilize the application.

[0202] When a user authenticates with an application, for example by providing client credentials (e.g., client ID and / or secret) of a client associated with the application, the user may receive an access token. The access token may be used with any API call (to any data management service). However, a particular data management service may not check which application was used for authentication. For example, and without limitation, client "A" may be authenticated to access application "A" and may receive an access token. The access token may then be used from application "B" and application "C", and the access checks made against that access token may be the same. This may be used in a single sign-on ("SSO") implementation where it may be desirable for a user to authenticate using one application (e.g., a web application) and then seamlessly use another application (e.g., another web application) without re-authenticating.

[0203] In particular embodiments, a user may be interested in allowing access to a resource (e.g., a data set) by a single application and / or a subset of applications, and / or may wish to deny access to a resource from one or more applications. In particular embodiments, restrictions may be used in a syntax that allows access to an ID and allows the restriction service to allow or deny access based on the value of the ID. Non-limiting examples of restrictions may include the following: ·QueryRows($subject.application.id='78e6c24b-c1fe-422f-8fb5-09c9fb45f0ae')

[0204] This restriction can allow a row in a table to be queried if the target application's ID matches the given ID. Denial can be done using "!=". ·QueryRows($subject.application.id in('78e6c24b-c1fe-422f-8fb5-09c9fb45f0ae','bf7fe945-1865-4aa3-8686-3f8b502dae22'))

[0205] This restriction can allow a row in the table to be queried if the target application's ID is among those listed. Exclusions can use "not in". ·QueryRows($subject.client.id= '6a5b99a3-7561-4180-ad63-91b8538f3738')

[0206] This restriction can allow rows of a table to be queried if the ID of the target client matches the given ID.

[0207] It will be understood that various constraints may be used in connection with various aspects of the disclosed embodiments, and that any constraint (including constraints using multiple Boolean expressions) may be used in connection with the disclosed systems and methods.

[0208] 5 illustrates a flowchart of a non-limiting example of a data query process 500 consistent with certain embodiments disclosed herein. The illustrated method 500 may be implemented in a variety of ways, including using software, firmware, hardware, and / or any combination thereof. In an embodiment, various aspects of the process 500 and / or its constituent steps may be performed by one or more systems and / or services, including systems and / or services that may implement a data management architecture as described herein.

[0209] At 502, a first data query request may be received by a DV service executing on a data management service system from a requesting application. In some embodiments, the requesting application may include an application executing on a different user system (e.g., a requesting user system, etc.) than the data management service system. In further embodiments, the requesting application may include an application executing within a protected sandbox of a secure execution environment, which may be a secure execution environment provided by the data management service system that provides the DV service. The first identification information included in the first data query request may include information associated with the requesting application, identification information associated with a user system, identification information associated with a requesting user, etc.

[0210] The DV service may determine whether the first data query request should be granted. For example, an authentication query may be issued by the DV service to the IAM service at 504. The authentication query may include, for example, but not limited to, a first identity, an indication of the requested data set (and potentially an indication of requested access rights and / or intended use of the data set), and / or a second identity that in some embodiments may be associated with the DV service issuing the authentication request. At 506, an indication that the first data request should be granted by the DV service may be received from the IAM service in response to the authentication query.

[0211] At 508, a second data query request may be generated and issued from the DV service to the TSDB service. The second data query request may include, for example, but not limited to, one or more of the first identity, an indication of the requested data set (and potentially an indication of requested access rights and / or intended use of the data set), and / or a second identity that may be associated with the DV service issuing the authentication request. In an embodiment, information included in the second data query request may be used by the TSDB service to authenticate the DV service, the requesting application, the requesting user system, and / or the requesting user with an IAM service and / or another authentication service.

[0212] The DV service may receive 510 TSDB service data included in a requested dataset stored in one or more data stores managed by the TSDB service. In some embodiments, the requested dataset may include a virtual dataset managed by the DV service, and the data included in the virtual dataset may be associated with and / or otherwise mapped to data stored in one or more data stores managed by the TSDB service. Certain embodiments may include portions of data associated with a dataset that may be included in multiple data stores and / or data storage tiers. For example, in some embodiments, a first portion of the virtual dataset may be associated with data stored in a cold data store managed by the TSDB service, and a second portion of the virtual dataset may be associated with data stored in a hot data store managed by the TSDB service. At least the subset data received from the TSDB service may be communicated from the DV service to a requesting application, user, and / or system at 512.

[0213] Consistent with certain embodiments disclosed herein, the dataset may be governed by one or more rules that may clarify and / or otherwise define one or more restrictions and / or other rights associated with the dataset. In some embodiments, the restrictions and / or access rights may be tied to an account identity. For example, in certain embodiments, at least one restriction defined in at least one rule associated with the requested dataset may be identified. In certain embodiments, the at least one restriction may be associated with at least one identity. The data received from the TSDB service may be filtered based on the at least one identified restriction. At least a subset of the data communicated from the DV service to the requesting application, user, and / or system at 512 may include data filtered based on the at least one identified restriction.

[0214] 6 illustrates a non-limiting example of a system 600 that may be used to implement certain embodiments of the systems and methods of the present disclosure. The various systems, services and / or devices used in connection with aspects of the disclosed embodiments may be communicatively coupled using a variety of networks and / or network connections (e.g., network 608). In certain embodiments, network 608 may include a variety of network communication devices and / or channels and may utilize any suitable communication protocols and / or standards that facilitate communication between the systems and / or devices.

[0215] Network 608 may include the Internet, a local area network, a virtual private network, and / or any other communications network utilizing one or more electronic communications technologies and / or standards (e.g., Ethernet, etc.). In some embodiments, network 608 may include a wireless carrier system, such as a personal communications system ("PCS"), and / or other suitable communications system incorporating any suitable communications standard and / or protocol. In further embodiments, network 708 may include an analog mobile communications network utilizing, for example, Code Division Multiple Access ("CDMA"), Global System for Mobile Communications or Group Special Mobile ("GSM"), Frequency Division Multiple Access ("FDMA"), and / or Time Division Multiple Access ("TDMA") standards, and / or a digital mobile communications network. In certain embodiments, network 708 may incorporate one or more satellite communications links. In still further embodiments, network 708 may utilize the IEEE 802.11 standard, Bluetooth, Ultra Wideband ("UWB") Zigbee, and / or any other suitable standard or standards.

[0216] The various systems and / or devices used in connection with aspects of the disclosed embodiments may comprise a variety of computing devices and / or systems, including any computing system or systems suitable for implementing the systems and methods disclosed herein. For example, the connected devices and / or systems may include a variety of computing devices and systems, including laptop computer systems, desktop computer systems, server computer systems, distributed computer systems, smartphones, tablet computers, etc.

[0217] In certain embodiments, the systems and / or devices may comprise at least one processor system configured to execute instructions stored on an associated non-transitory computer-readable storage medium. As discussed in more detail below, systems used in connection with implementing various aspects of the disclosed embodiments may further comprise a secure processing unit ("SPU") configured to perform sensitive operations such as trusted credential and / or key management, cryptographic operations, secure policy management, and / or other aspects of the systems and methods disclosed herein. The systems and / or devices may further comprise software and / or hardware configured to enable electronic communication of information between the devices and / or systems over a network using any suitable communication technology and / or standard.

[0218] As illustrated in FIG. 6 , an exemplary system 600 may include a processing unit 602; system memory 604, which may include high-speed random access memory ("RAM"), non-volatile memory ("ROM"), and / or one or more bulk non-volatile non-transitory computer-readable storage media (e.g., hard disk, flash memory, etc.) for storing programs and other data for use and execution by the processing unit 602; a port 614 for interfacing with removable memory 616, which may include one or more diskettes, optical storage media (e.g., flash memory, thumb drives, USB dongles, compact discs, DVDs, etc.) and / or other non-transitory computer-readable storage media; a network interface 606 for communicating with other systems over one or more network connections and / or networks 608 using one or more communications technologies; a user interface 612, which may include a display and / or one or more input / output devices, such as, for example, a touch screen, keyboard, mouse, track pad, and the like; and one or more buses 618 for communicatively coupling the elements of the system.

[0219] In some embodiments, system 600 may alternatively or additionally include an SPU 610 that is protected from tampering by users of system 600 or other entities by utilizing secure physical and / or virtual security techniques. SPU 610 may help enforce security of sensitive operations such as personal information management, trusted credentials and / or key management, privacy and policy management, and other aspects of the systems and methods disclosed herein. In certain embodiments, SPU 610 may be configured to operate in a logically secure processing domain and to protect and manipulate sensitive information as described herein. In some embodiments, SPU 610 may include an internal memory that stores executable instructions or programs configured to enable SPU 610 to perform secure operations as described herein.

[0220] The operation of the system 600 may generally be controlled by the processing unit 602 and / or SPU 610, which operate by executing software instructions and programs stored in the system memory 604 (and / or other computer-readable media, such as removable memory 616). The system memory 604 may store various executable programs or modules for controlling the operation of the system 600. For example, the system memory may comprise an operating system ("OS") 620, which at least partially manages and coordinates system hardware resources and provides common services for the execution of various applications, and a trust and privacy management system 622 for implementing trust and privacy management functions, including protection and / or management of secure data through management and / or enforcement of associated policies. The system memory 604 may further include, without limitation, communications software 624 configured to partially enable communication with the system 600, one or more applications, data management services 626 configured to implement various aspects of the disclosed systems and / or methods, and / or any other information and / or applications configured to implement embodiments and / or aspects of the systems and methods disclosed herein.

[0221] The systems and methods disclosed herein are not inherently related to any particular computer, electronic control unit, or other apparatus, but may be implemented by any suitable combination of hardware, software, and / or firmware. A software implementation may include one or more computer programs that include executable code / instructions that, when executed by a processor, cause the processor to perform a method defined at least in part by the executable instructions. The computer programs may be written in any type of programming language, including compiled or interpreted languages, and may be deployed in any form, such as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. Furthermore, the computer programs may be deployed to be executed on one computer or on multiple computers at one site or distributed multiple sites, interconnected by a communication network.

[0222] The software embodiment may be implemented as a computer program product including a computer program and a non-transitory storage medium configured to store instructions that, when executed by a processor, are configured to cause the processor to perform a method in accordance with the instructions. In certain embodiments, the non-transitory storage medium may take any form capable of storing processor-readable instructions on the non-transitory storage medium. The non-transitory storage medium may be embodied by a compact disc, a digital video disk, a magnetic disk, a flash memory, an integrated circuit, or any other non-transitory digital processing apparatus memory device.

[0223] Although the foregoing has been described in some detail for clarity, it will be apparent that certain changes and modifications can be made without departing from the principles of the present invention. For example, it will be understood that several modifications can be made to the various embodiments, systems, services, and / or components presented in connection with the drawings and / or associated description within the scope of the body of work of the present invention, and that the examples presented in the drawings and described herein are provided for purposes of illustration and description, not limitation. It is further noted that there are many alternative ways of implementing both the systems and methods described herein. Thus, the present embodiment should be considered as illustrative and not limiting, and the embodiments of the present invention should not be limited to the details given herein, but may be modified within the scope of the appended claims and equivalents.

Claims

1. 1. A method of managing data executed by a data management service system, the method comprising: generating, by a data virtualization service executing on the data management service system, a virtual dataset that references data stored in one or more data stores managed by a database service and is generated based on at least one restriction defined in at least one rule associated with the data stored in the one or more data stores; receiving a first data query request at the data virtualization service from a requesting application, the first data query request including first identification information and an indication of the virtual data set; determining, by the data virtualization service, whether to allow the first data query request, issuing, by the data virtualization service, an authentication query to an identity and access management service, the authentication query including the first identification information and the indication of the virtual data set; receiving an indication from the identity and access management service in response to the authentication query authorizing the first data query request; generating a second data query request in response to receiving the indication of granting the first data query request, the second data query including the indication of the virtual data set; sending the second data query request to the database service; receiving, from the database service, the data associated with the virtual dataset stored in the one or more data stores; sending at least a subset of the data received from the database service to the requesting application; A method comprising:

2. The method of claim 1 , wherein the requesting application comprises an application running on a user system different from the data management service system.

3. The method of claim 2 , wherein the first identification information comprises an identification information associated with the user system.

4. The method of claim 2 , wherein the first identification information comprises identification information associated with a user of the user system.

5. The method of claim 1 , wherein the first identification information comprises identification information associated with the requesting application.

6. The method of claim 1 , wherein the requesting application comprises an application executing within a protected sandbox of a secure execution environment.

7. The method of claim 6 , wherein the secure execution environment comprises an execution environment of the data management service system.

8. The method of claim 1 , wherein the second data query request includes second identification information.

9. The method of claim 8 , wherein the second identification information includes information identifying the data virtualization service.

10. The method of claim 1 , wherein the virtual dataset is associated with data stored in multiple data stores managed by the database service.

11. 11. The method of claim 10, wherein a first portion of the virtual data set is associated with data stored in a cold data store and a second portion of the virtual data set is associated with data stored in a hot data store.

12. The method of claim 1 , wherein the virtual data set includes data contained in another virtual data set.

13. The method of claim 1 , wherein the virtual dataset comprises a time series dataset.

14. The method of claim 1 , wherein the identity and access management service is configured to query a directory managed by the identity and access management service based on the authentication query.

15. 15. The method of claim 14, wherein the directory includes a plurality of managed objects, the plurality of managed objects including at least a first managed object associated with the first identification and a second managed object associated with the indication of the virtual data set.

16. 15. The method of claim 14, wherein the first identification information includes an indication of a managed account in the directory, and the indication of the virtual dataset includes an indication of a managed virtual dataset in the directory, and the managed account and the managed virtual dataset are objects contained in the directory.

17. 10. The method of claim 1, further comprising filtering the data received from the database service based on the at least one restriction to generate filtered data, wherein the at least a subset of the data sent to the requesting application comprises the filtered data.