Central data governance and access control for enterprise data
The central data governance and access control method addresses platform-specific limitations by automating access requests and policies, enhancing compatibility and simplifying user access across diverse ecosystems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-04-17
- Publication Date
- 2026-03-26
AI Technical Summary
Existing data access control systems are tightly coupled to specific data platforms, leading to compatibility issues and complex access processes, and lack a unified layer across ecosystems, complicating user access management.
A method for central data governance and access control that includes receiving information identifying data of interest, generating access requests, automatically forwarding approval requests to access control entities, receiving responses, creating policies, and providing access based on these policies, with support for multiple storage technologies and automated communication with data owners and governance officers.
Provides a unified and simplified data access solution that reduces user effort and time by automating access requests and ensuring centralized auditing, enabling seamless integration with various platforms and technologies.
Smart Images

Figure 0007836473000001 
Figure 0007836473000002 
Figure 0007836473000003
Abstract
Description
Technical Field
[0001] This description relates to central data governance and access control for enterprise data and how to use it.
Background Art
[0002] There are many teams in an organization that generate data. There are numerous components or products that enable, monitor, and manage data security across the Hadoop platform, such as Apache Ranger. Similar products include Atlan and Datahub.
[0003] To make this data available, users must be granted access to the data. Data access management, or data access governance, provides secure coordination, visibility, and control of access to files and information. Data access governance is implemented across the organization, including business users and Information Technology (IT) teams. Data access governance includes maintaining data security, protecting Personal Identifiable Information (PII), providing access to data assets, managing permissions, and many other data security functions. A data access control system enables a company to securely protect confidential information, define data ownership, and enforce managed access control.
[0004] However, existing systems present problems in managing data access. For example, such systems are tightly coupled to data platforms and therefore cannot be extended to new data platforms that can be implemented. Existing systems also present compatibility issues with different data platforms and use complex data access processes that are cumbersome for users to use to obtain access. Existing access control layers provide functionality within ecosystems such as AWS or Azure, but do not provide a unified layer. Users often request permission from data owners to obtain access, but identifying who the data owners are, how to contact them, and other issues complicates the process. [Overview of the project] [Means for solving the problem]
[0005] In at least one embodiment, a method for providing central data governance and access control for enterprise data includes: receiving information identifying data of interest in a storage system provided by a data platform; generating an access request for the data of interest in the storage system; automatically forwarding an approval request to one or more access control entities based on the access request; receiving an approval response from one or more access control entities; creating a policy for accessing the data of interest in the storage system; and providing access to the data of interest in the storage system based on the policy.
[0006] In at least one embodiment, the data platform includes a memory for storing computer-readable instructions and a processor attached to the memory, the processor being configured to execute computer-readable instructions to receive information identifying data of interest in a storage system, generate an access request for the data of interest in the storage system, automatically forward an authorization request to one or more access control entities based on the access request, receive an authorization response from one or more access control entities, create a policy for accessing the data of interest in the storage system, and perform actions to provide access to the data of interest in the storage system based on the policy.
[0007] In at least one embodiment, a non-temporary computer-readable medium having computer-readable instructions stored thereon, when executed by a processor, causes the processor to perform operations including: receiving information identifying data of interest in a storage system provided by a data platform; generating an access request for the data of interest in the storage system; automatically forwarding an authorization request to one or more access control entities based on the access request; receiving an authorization response from one or more access control entities; creating a policy for accessing the data of interest in the storage system; and providing access to the data of interest in the storage system based on the policy.
[0008] The aspects of this disclosure will be best understood by reading the following detailed description in conjunction with the attached figures. Note that, in accordance with standard industry practice, various features are not depicted to scale. In fact, it is possible to enlarge or reduce the dimensions of various features for clarity in the description. [Brief explanation of the drawing]
[0009] [Figure 1] This is a diagram of a data discovery platform according to at least one embodiment. [Figure 2] This is a novel data access request entry interface according to at least one embodiment. [Figure 3] This is a diagram of the display of a new data access request according to at least one embodiment. [Figure 4] This is a diagram of a system for providing policy synchronization according to at least one embodiment. [Figure 5] This is a flowchart of a method for providing central data governance and access control for enterprise data by at least one embodiment. [Figure 6] This is a high-level functional block diagram of a processor-based system according to at least one embodiment. [Modes for carrying out the invention]
[0010] The embodiments described herein illustrate examples of implementing various features of the subject matter provided. For the sake of simplicity in disclosing the invention, examples of components, values, operations, materials, arrangements, etc., are described below. Naturally, these are examples and are not intended to be limiting. Other components, values, operations, materials, arrangements, etc., are contemplated. For example, the formation of a first feature above or on a second feature in the following description includes embodiments in which the first and second features are formed in direct contact, and also includes embodiments in which an additional feature is formed between the first and second features so that the first and second features cannot be in direct contact. In addition, this disclosure repeats reference numerals and / or letters in various examples. This repetition is for the sake of brevity and clarity and does not indicate relationships and / or configurations between the various embodiments described.
[0011] Furthermore, spatially relative terms such as “beneath,” “below,” “lower,” “above,” and “upper” are used herein to facilitate descriptions of the relationship between one element or feature and another element or feature, as shown in the figures. Spatially relative terms are intended to encompass different orientations of the device in use or operation, in addition to the orientation shown in the figures. The device may be oriented in other directions (rotated 90 degrees or other directions), and the spatially relative descriptors used herein will be interpreted accordingly.
[0012] Terms such as “User equipment,” “Mobile station,” “Mobile,” “Mobile device,” “Subscriber station,” “Subscriber equipment,” “Access terminal,” “Terminal,” and “Handset,” and similar terms, refer to wireless devices used by subscribers or users of wireless communication services to receive or transmit data, control, voice, video, sound, games, data streaming, or signaling streaming. The aforementioned terms are interchangeable in this specification and the associated drawings. Terms such as “Access point,” “Base station,” “Node B,” “Evolved Node B (eNode B),” “Next Generation Node B (gNB),” “Enhanced gNB (en-gNB),” “Home Node B (HNB),” and “Home Access Point (HAP)” refer to components or devices of a wireless network that transmit data, control, voice, video, sound, games, or any data streaming or signaling streaming to and receive from the UE.
[0013] In at least one embodiment, a method for providing central data governance and access control for enterprise data includes: receiving information identifying data of interest in a storage system provided by a data platform; generating an access request for the data of interest in the storage system; automatically forwarding an authorization request to one or more access control entities based on the access request; receiving an authorization response from one or more access control entities; creating a policy for accessing the data of interest in the storage system; and providing access to the data of interest in the storage system based on the policy.
[0014] According to at least one embodiment, the data platform includes a data discovery and catalog interface, a data governance and access control layer, and a data storage layer. The data governance and access control layer includes a data access layer and a data governance layer. The data governance layer includes a policy engine and an audit interface. The data storage layer can support different storage technologies such as MinIO, Yugabyte, MySQL®, and Kafka. Plugins can be added to the data governance and access control layer to support different storage technologies.
[0015] The data platform also implements data query and data computation interfaces. Information identifying data of interest in the storage system is received by the data discovery and catalog interfaces, and access requests for the data of interest in the storage system are generated. The data access layer automatically forwards authorization requests to one or more access control entities based on the access requests and receives authorization responses from one or more access control entities. The policy engine creates policies for accessing the data of interest in the storage system based on the authorization responses, and the data storage layer provides access to the data of interest in the storage system based on the policies. The data governance layer provides access verification for accessing the data of interest in the storage system based on the policies.
[0016] The data query interface receives input for querying the data storage layer, generates queries, forms search inputs for metadata and data definitions from the data discovery catalog to identify datasets based on the inputs, and generates query requests to access datasets in the storage system in response to the metadata and data definitions. Access rights are determined by the policy engine based on policies, and data access in the storage system is granted according to the determined access rights. Inputs are received in the data calculation interface to create a first-level report for analyzing data usage in the storage system. The data storage layer parses events, extracts logs associated with data access in the storage system, and sends the logs to the audit interface to audit data usage based on a deterministic process.
[0017] Embodiments described herein provide methods that offer one or more advantages. For example, a method for providing central data governance and access control for enterprise data provides a one-stop solution for data governance and access-related services. This reduces time and effort for end users because they do not need to create tickets to gain access to specific data and do not need to identify the data owner or data governance office for specific data. Communication with data owners and data governance officers is handled automatically. Trust and control are also provided through centralized and optimized auditing.
[0018] Figure 1 shows a data discovery platform 100 according to at least one embodiment.
[0019] In Figure 1, the data discovery platform 100 includes an authentication and middleware interface 110. Information for identifying data of interest is provided to the data discovery platform 100 by one or more sources, such as users 120, applications 121, analysts 122, data scientists 123, organizations 124, and administrators 125. Information for identifying data of interest is provided to the onboarding interface 130, a web user interface (UI) 132, or other access interfaces. The authentication and middleware interface 110 provides a centralized interface for controlling access across different platforms such as AWS, Kubernetes, and VMs, for end users to request access to data and for data owners to manage access requests.
[0020] The data discovery and catalog interface 140 enables a user to search for and discover 141 data of interest in the data discovery platform 100. The user can also tag the data and create classifications and categorizations 142 via the data discovery and catalog 140. Whenever the user discovers data of interest, the data discovery and catalog interface 140 generates an access request 144 that is provided to the data access layer 152 of the data governance and access control layer 150. The data governance and access control layer 150 provides control for access to the data storage layer 160.
[0021] The access request 144 can include detailed information such as a list of columns, masking details for personally identifiable information (PII), the duration of the request, and the reason for the request. The data governance and access control layer 150 is highly integrated with the data discovery platform 100 rather than being tightly coupled with any particular storage platform.
[0022] The data governance and access control layer 150 provides access control and auditing for other storage technologies such as MiniIO 161 and Yugabyte 162, as well as MySQL 163 and Kafka 164. The data is highly available with authentication, authorization, and accounting (AAA) provisions. For example, Apache Ranger can make data available in a shared state on a stretched Yugabyte cluster 162 on Kubernetes.
[0023] The data discovery and catalog interface 140 provides an access request 144 to the data access layer 152 of the data governance and access control layer 150, and an approval request 153 is sent to the data owner 170 and the data governance responsible person 172. Those skilled in the art will understand that these are provided as examples and that other data approval entities can be configured to grant approval to the approval request 153. The transfer of the approval request 153 to the data approval entity can be configured for automatic provisioning to the data approval entity. For example, the approval request 153 can be provided to the data owner 170 via a communication channel such as email. When the data access layer 152 receives an approval response 154 from the data owner 170, the data access layer 152 sends an approval request 156 to the data governance responsible person 172 for a second level of approval. The data access layer 152 waits for an approval response 157 from the data governance responsible person 172.
[0024] The data governance and access control layer 150 enables the data owner 170 or the data governance responsible person 172 to determine who can access a specific type of data for reading or writing, set a time period for access, enable access to data or specific columns based on predetermined limits, etc. The data governance and access control layer 150 also enables the data owner 170 or the data governance responsible person 172 to edit or delete permissions.
[0025] The data governance and access control layer 150 tracks approval requests 153 sent to the data owner 170 and approval responses 154 received from the data owner 170. It can also track approval requests 156 sent to the data governance officer 172 and approval responses 157 received from the data governance officer 172. Therefore, users do not need to create tickets to gain access to specific data, nor do the data owner 170 or the data governance office 172 need to find out who they are regarding specific data. The data governance and access control layer 150 automatically handles communication with the data owner 170 and the data governance officer 172.
[0026] The data access layer 152 receives an approval response from the data governance officer 172. After receiving approval responses 154 and 157 in the data access layer 152, the final approval response 158 is sent to the policy engine 180 in the data governance layer 180, and a policy 184 for accessing the requested data is created and applied in the data storage layer 160. Subsequently, access to the requested data is provided to the user in accordance with policy 184.
[0027] The data governance and access control layer 150 is not limited to a few components such as the Hadoop Distributed File System (HDFS) and edge-based systems. The data governance and access control layer 150 allows for the integration of new technologies by adding plug-ins to the data governance and access control layer 150 for new technologies. Furthermore, the data governance and access control layer 150 provides a simpler process that takes less time for users to gain access to data in the data storage layer 160.
[0028] Approval responses 154 from the data owner 170 and 157 from the data governance officer 172 are returned to the data access layer 152. The data access layer 152 creates a policy 184 using the approved access 158. The policy engine 182 creates the policy 184 and applies it to the data storage layer 160.
[0029] The data image is maintained by the data storage layer 160. For example, the data storage layer 160 can maintain data using MinIO 161, Yugabyte 162, MySQL 163, Apache Kafka 164, etc. MinIO 161 is a cloud-compatible object storage. Yugabyte 162 is a distributed SQL database for cloud-native applications. MySQL 163 is an open-source relational database management system. Apache Kafka 164 is a distributed event store and stream processing platform.
[0030] The policy engine 182 is configured according to a specific database and data type. For example, a new policy 184 can be created for MinIO 161 or Yugabyte 162. The data storage layer 160 communicates with the policy engine 182 to perform access verification 185. The audit interface 186 of the data governance layer 180 is provided to track what users are doing with data by obtaining logs or actions that users perform on specific data. Thus, users use the data discovery platform 100 to search a catalog of data useful for their purposes and obtain authorization to access data that satisfies their access request 144. The policies 184 created based on the authorization responses 153, 157 are then used in the data storage layer 160 to create policies 184 that are applied to the relevant data sources.
[0031] The user can also provide inputs to generate queries 191 in the data query interface 190. Through queries 191, the user can create views 192 on existing data in the data storage layer 160. Queries 191 can also be provided to search multiple data stores in the data storage layer 160. The data query interface 190 forms search inputs for metadata and data definitions 193 from the data discovery and catalog interface 140 to identify datasets. Using the metadata and data definitions 193, the data query interface 190 sends read or write access requests 194 to the policy engine 182.
[0032] The policy engine 182 checks for access rights. In response that the policy engine 182 already has an approved policy 184 associated with the access request, the access request 194 is sent to the data storage layer 160, which can perform access verification 185 of the access request 194. Once access rights are verified, the data storage layer 160 returns a response 195.
[0033] The data calculation interface 196 receives inputs 197 for generating a first-level report for analysis. After data access is granted, the user can generate query inputs to access the data using the query layer or from any user-defined process. The data storage layer 160 provides policy synchronization and auditing 198 from the data storage layer 160 to the data governance and access control layer 150. Read / write access requests 144 are validated against policies 184 created for the user. The data governance and access control layer 150 provides a synchronization process 199 to the data governance layer 180 so that policies 184 are kept up-to-date.
[0034] Therefore, the Data Discovery Platform 100 provides a solution for data governance and access-related services, reducing end-user time and effort, and providing reliability and control through centralized and optimized auditing. Current systems have limitations in providing data governance and access control for enterprise data. For example, existing products such as Apache Ranger, Atlan, and Datahub provide a framework for enabling, monitoring, and managing data security across the entire Hadoop platform. However, in such systems, data governance is tightly coupled to a specific data platform. Furthermore, the data access process for generating and managing data access requests to data users and owners is complex.
[0035] Ticketing services or manual intervention are used to request access, which is then manually handled by the data owner 170 and data governance officer 172 for services outside the data storage layer 160. The ability to provide data access and data governance does not enable access control outside of its own system. Apache Ranger also does not provide access control and auditing for MinIO and Yugabyte. Apache Ranger is also not Kubernetes compatible.
[0036] Figure 2 shows a novel data access request entry interface 200 according to at least one embodiment.
[0037] In Figure 2, a user 210 and a user contact 212 are provided. User 210 can provide an LDAP (Lightweight Directory Access Protocol) account ID 214 and an LDAP password 216. Having an LDAP account is a prerequisite for being able to request access. In response to user 210 not having an LDAP account ID 214, user 210 can create a new LDAP account 220. User 210 enters information to identify why the user needs the data 230. User 210 identifies where the data will be used 240. User 210 can choose whether the data will be used within the organization 242 or outside the organization 244.
[0038] Next, user 210 enters the data access start date 250 and the data access end date 252. The user enters information for the access type selection 260. The search column window 262 is provided to identify data columns such as All Columns 270, EnforcerID 271, PersiodStartTime 272, PeriodEndTime 273, ClientIP 274, ServerIP 275, and ServiceID 276. User 210 can choose to cancel 280 or submit 282.
[0039] Figure 3 shows a display of a new data access request 300 according to at least one embodiment.
[0040] Figure 3 displays information on total requests 310, pending requests 312, approved requests 314, rejected requests 316, and expired requests 318. Total requests 310 are shown as "1". Pending requests 312 are shown as "1". Approved requests 314 are shown as "0". Rejected requests 316 are shown as "0". Expired requests 318 are shown as "0".
[0041] The data source 320 includes the selection of List All 322, the selection of MinIO 324, the selection of Yugabyte (YB) 326, and the selection of data bus 328. The selection of MinIO 324 is for selecting cloud-compatible object storage. The selection of Yugabyte 326 is for selecting a distributed SQL database for cloud-native applications. The selection of data bus 328 is for selecting stream data that can be pushed by any application. However, those skilled in the art will recognize that other data sources can be used without departing from the embodiments described herein.
[0042] In Figure 3, a new data access request 300 includes the identification of one result 330 display 1. The result 330 contains information in columns identified using different headers. Figure 3 shows rows of header 340 for catalog 342, data classification 343, status 344, requester 345, usage period 346, and request date 347. However, those skilled in the art will recognize that other columns and headers can be used without departing from the embodiments described herein.
[0043] In Figure 3, row 350 is displayed in result 330. The header catalog 342 is identified as Allot HDR 352, the data classification 343 is identified as classification 353, the status 344 is identified as created request 354, the requester 345 is identified as User_1@abc.com 355, the usage period 346 is identified as January 1, 2023 to January 3, 2023 356, and the request date 347 is identified as January 1, 2023 357.
[0044] The user can also select a dropdown menu using the dropdown menu selector (Caret) 360. The dropdown menu selector 360 indicates that a dropdown menu can be displayed. Selecting the dropdown menu selector 360 opens a dropdown menu option when selected, and the dropdown menu can display a panel of additional selections from which the user can further choose an action to take. Additional actions can be selected using the Kebob selector (three vertical dots) 370. The Kebob selector 370 opens a smaller inline menu with additional options.
[0045] Figure 4 shows a system 400 for providing policy synchronization according to at least one embodiment.
[0046] In Figure 4, the data governance layer 410 provides the ability to create policies for any number of data lakes or databases. The data governance layer 410 includes a policy engine 412 and an audit interface 416. The policy engine 412 can create new policies 420, 422 according to the type of database and data. For example, it can create a new policy 420 for MinIO 430 or a new policy 422 for Yugabyte 440. However, the functionality can be extended to systems other than MinIO 430 and Yugabyte 440 to enable cross-platform access control. New policies 420, 422 are created according to the database and data. The policy engine 412 sends the policy details 413 to the backend process 414, and a new policy 420 is created by converting the policy details 413 to, for example, a MinIO 430 type, or a new policy 422 is created by converting the policy details 413 to, for example, a Yugabyte 430 type.
[0047] The policy engine 412 and audit interface 416 keep the policy synchronized with the central data governance layer 410. As shown in Figure 4, the new MinIO policy 430 is sent to the MinIO policy engine 432 of the MinIO system 430. The new Yugabyte policy 422 is sent to the Yugabyte access control 442 of the Yugabyte system 440.
[0048] The MinIO policy engine 432 applies the new MinIO policy 420, and the Yugabyte access control 442 applies the new Yugabyte policy 422. The MinIO event notification 434 of the MinIO system 430 parses the events and extracts the MinIO audit events and logs 450. The MinIO audit events and logs 450 are sent to the audit interface 416 of the data governance layer 410. The Yugabyte access log 444 parses the Yugabyte audit log 452. The Yugabyte audit log 452 is sent to the audit interface 416 of the data governance layer 410.
[0049] To perform large-scale analysis of MinIO events 450 or Yugabyte audit logs 452, the data governance layer 410 executes a backend spark job 417 that reads and parses the MinIO events 450 or Yugabyte audit logs 452 and loads the relevant access audit details 418 into the audit interface 416. Synchronization policies 460 and 462 are used to synchronize the MinIO audit events and logs 450 and the Yugabyte audit logs 452, respectively. The data governance layer allows the use of plug-ins to extend the data governance layer 410 to other databases with enterprise security.
[0050] The audit interface 416 provides audit feedback using a deterministic approach. A deterministic approach identifies results based on causes, and the results do not change unless the cause is corrected. As shown in Figure 4, different types of databases 430, 440 and different types of access logs 450, 452 can be provided. The deterministic approach provided by the audit interface 416 simplifies the processing of access logs 450, 452 and how users track different databases 430, 440. Although the audit interface 416 is described as using a deterministic approach, the embodiments described herein can be implemented in other ways, such as by applying artificial intelligence (AI) and machine learning (ML) to generate audit results.
[0051] Figure 5 is a flowchart 500 of a method for providing central data governance and access control for enterprise data according to at least one embodiment.
[0052] In Figure 5, the process starts at S502, and information identifying data of interest in the storage system is received from the data platform at S510. Referring to Figure 1, information for identifying data of interest is provided to the data discovery platform 100 by one or more sources, such as users 120, applications 121, analysts 122, data scientists 123, organizations 124, and administrators 125. Information for identifying data of interest is provided to the onboarding interface 130, the web user interface (UI) 132, or other access interfaces. The authentication and middleware interface 110 provides a centralized interface for controlling access across different platforms such as AWS, Kubernetes, and VMs, for end users to request access to data and for data owners to manage access requests.
[0053] S514 generates an access request for data of interest in the storage system. Referring to Figure 1, whenever a user discovers data of interest, the data discovery and catalog interface 140 generates an access request 144 which is provided to the data access layer 152 of the data governance and access control layer 150. The user can also provide input to generate a query 191 in the data query interface 190. Through the query 191, the user can create a view 192 on existing data in the data storage layer 160. The query 191 can also be provided to search multiple data stores in the data storage layer 160. The data query interface 190 forms search inputs for metadata and data definitions 193 from the data discovery and catalog interface 140 to identify the dataset. Using the metadata and data definitions 193, the data query interface 190 sends a read or write access request 194 to the policy engine 182. The policy engine 182 checks the access rights. In response that the policy engine 182 already has an approved policy 184 associated with the access request, the access request 194 is sent to the data storage layer 160, which can perform access verification 185 of the access request 194. Once access rights are verified, the data storage layer 160 returns a response 195.
[0054] S518. An approval request is automatically forwarded to one or more access control entities based on an access request. Referring to Figure 1, the data discovery and catalog interface 140 provides an access request 144 to the data access layer 152 of the data governance and access control layer 150, and the approval request 153 is sent to the data owner 170 and the data governance officer 172. Those skilled in the art will understand that these are provided as examples and other data approval entities can be configured to grant approval to the approval request 153. The forwarding of the approval request 153 to the data approval entities can be configured for automatic provisioning to the data approval entities. For example, the approval request 153 can be provided to the data owner 170 via a communication channel such as email. When the data access layer 152 receives an approval response 154 from the data owner 170, the data access layer 152 sends an approval request 156 to the data governance officer 172 for a second level of approval. The data access layer 152 awaits an approval response 157 from the data governance officer 172.
[0055] S522 receives an approval response from one or more access control entities. Referring to Figure 1, when the data access layer 152 receives an approval response 154 from the data owner 170, the data access layer 152 sends an approval request 156 to the data governance officer 172 for a second level of approval. The data access layer 152 awaits an approval response 157 from the data governance officer 172. The data access layer 152 receives an approval response from the data governance officer 172. The data governance and access control layer 150 enables the data owner 170 or the data governance officer 172 to manage the data, including determining who can access certain types of data for reading or writing, setting time periods for access, and enabling access to data or specific columns based on predetermined restrictions. The data governance and access control layer 150 also enables the data owner 170 or the data governance officer 172 to edit or delete permissions.
[0056] S526 is when a policy is created to access the data of interest in the storage system. Referring to Figure 1, after approval responses 154 and 157 are received in the data access layer 152, a final approval response 158 is sent to the policy engine 180 in the data governance layer 180, and a policy 184 for accessing the requested data is created and applied in the data storage layer 160.
[0057] S530 provides access to data of interest in the storage system based on policy. Referring to Figure 1, access to the requested data is then provided to the user according to policy 184.
[0058] Access verification is performed by the storage system based on policies to access data of interest in the storage system (S534). Referring to Figure 1, the data storage layer 160 communicates with the policy engine 182 to perform access verification 185. The data governance and access control layer 150 provides access control and auditing for MiniIO 161 and Yugabyte 162, as well as other storage technologies such as MySQL 163 and Kafka 164.
[0059] In S538, the event is analyzed and logs associated with access to data of interest in the storage system are extracted. Referring to Figure 4, the MinIO policy engine 432 applies the new MinIO policy 420, and the Yugabyte access control 442 applies the new Yugabyte policy 422. The MinIO event notification 434 of the MinIO system 430 analyzes the event and extracts the MinIO audit event and log 450. The Yugabyte access log 444 analyzes the Yugabyte audit log 452.
[0060] S542 sends logs to audit the use of data of interest in the storage system. Referring to Figure 4, the MinIO audit event and log 450 are sent to the audit interface 416 of the data governance layer 410. The Yugabyte audit log 452 is sent to the audit interface 416 of the data governance layer 410.
[0061] The use of data of interest in a storage system is audited based on a deterministic process S546. Referring to Figure 4, the audit interface 416 provides audit feedback using a deterministic method. The deterministic method identifies results based on causes, and the results do not change unless the cause is corrected. As shown in Figure 4, different types of databases 430, 440 and different types of access logs 450, 452 can be provided. The deterministic method provided by the audit interface 416 simplifies the processing of access logs 450, 452 and how users track different databases 430, 440. Although the audit interface 416 is described as using a deterministic method, the embodiments described herein can be implemented in other ways, such as by applying artificial intelligence (AI) and machine learning (ML) to generate audit results.
[0062] In the storage system, inputs are received to create a first-level report for analyzing data usage (S550). Referring to Figure 1, the data calculation interface 196 receives inputs to create a first-level report for analysis (197). After data access is granted, the user can generate query inputs to access the data using the query layer or from any user-defined process.
[0063] The process then terminates (S560).
[0064] At least one embodiment of the Method provides central data governance and access control for enterprise data. The Method includes receiving information identifying data of interest in a storage system provided by a data platform; generating an access request for the data of interest in the storage system; automatically forwarding an authorization request to one or more access control entities based on the access request; receiving an authorization response from one or more access control entities; creating a policy for accessing the data of interest in the storage system; and providing access to the data of interest in the storage system based on the policy.
[0065] Figure 6 is a high-level functional block diagram of a processor-based system 600 according to at least one embodiment.
[0066] In at least one embodiment, the processing circuit 600 provides central data governance and access control for enterprise data. The processing circuit 600 implements central data governance and access control for enterprise data using a processor 602. The processing circuit 600 also includes a non-temporary computer-readable storage medium 604 used to implement central data governance and access control for enterprise data. In particular, the non-temporary computer-readable storage medium 604 is encoded, i.e., stored, in instructions 606, i.e., computer program code, which are executed by the processor 602, and the instructions 606 cause the processor 602 to perform operations to provide central data governance and access control for enterprise data. The execution of instructions 606 by the processor 602 represents (at least partially) an application that implements at least a part of the methods described herein (hereinafter, the referred processes and / or methods) in one or more embodiments.
[0067] The processor 602 is electrically coupled to a non-temporary computer-readable storage medium 604 via a bus 608. The processor 602 is electrically coupled to an input / output (I / O) interface 610 via the bus 608. The network interface 612 is also electrically connected to the processor 602 via the bus 608. The network interface 612 is connected to a network 614, thereby connecting the processor 602 and the non-temporary computer-readable storage medium 604 to external elements via the network 614. The processor 602 is configured to execute instructions 606 encoded in the non-temporary computer-readable storage medium 604 in order to make the processing circuit 600 available for performing at least a portion of a process and / or method. In one or more embodiments, the processor 602 is a central processing unit (CPU), a multiprocessor, a distributed processing system, an application-specific integrated circuit (ASIC), and / or a suitable processing unit.
[0068] The processing circuit 600 includes an I / O interface 610. The I / O interface 610 is coupled to an external circuit. In one or more embodiments, the I / O interface 610 includes a keyboard, keypad, mouse, trackball, trackpad, touchscreen, and / or cursor directional keys for communicating information and commands to the processor 602.
[0069] The processing circuit 600 also includes a network interface 612 coupled to the processor 602. The network interface 612 allows the processing circuit 600 to communicate with a network 614 to which one or more other computer systems are connected. The network interface 612 includes wireless network interfaces such as Bluetooth, Wi-Fi, Worldwide Interoperability for Microwave Access (WiMAX), General-Purpose Packet Radio Service (GPRS), or Wideband Code Division Multiple Access (WCDMA®), or wired network interfaces such as Ethernet, Universal Serial Bus (USB), or IEEE 864.
[0070] The processing circuit 600 is configured to receive information via the I / O interface 610. The information received through the I / O interface 610 includes one or more of the following for processing by the processor 602: instructions, data, design rules, cell libraries, and / or other parameters. The information is transferred to the processor 602 via the bus 608. The processing circuit 600 is configured to receive information related to the user interface (UI) via the I / O interface 610. The information is stored as the UI 620 in the non-temporary computer-readable storage medium 604.
[0071] In one or more embodiments, one or more non-temporary computer-readable storage media 604 store (in compressed or uncompressed form) instructions 606 that can be used to program a computer, processor, or other electronic device to perform the processes or methods described herein. One or more non-temporary computer-readable storage media 604 include one or more of the following: electronic storage media, magnetic storage media, optical storage media, quantum storage media, etc.
[0072] For example, the non-temporary computer-readable storage medium 604 may include, but is not limited to, a hard drive, a floppy diskette, an optical disc, read-only memory (ROM), random access memory (RAM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, a magnetic or optical card, a solid-state memory device, or other types of physical media suitable for storing electronic instructions. In one or more embodiments using an optical disc, one or more non-temporary computer-readable storage mediums 604 include compact disc read-only memory (CD-ROM), compact disc read / write (CD-R / W), and / or digital video discs (DVD).
[0073] In one or more embodiments, the non-temporary computer-readable storage medium 604 stores instructions 606 configured to cause a processor 602 to execute at least part of a process and / or method for providing central data governance and access control for enterprise data. In one or more embodiments, the non-temporary computer-readable storage medium 604 also stores information such as algorithms that facilitate the execution of at least part of a process and / or method for providing central data governance and access control for enterprise data.
[0074] Therefore, in at least one embodiment, the processor 602 implements the data discovery and catalog interface 630, the data governance and access control layer 640, and the data storage layer 650 by executing instructions 606 stored in one or more non-temporary computer-readable storage media 604. The processor 602 implements the data access layer 642 and the data governance layer 646 in the data governance and access control layer 640 by executing instructions 606. The processor 602 implements the policy engine 647 and the audit interface 649 in the data governance layer 646 by executing instructions 606. The data storage layer 650 can support various storage technologies 652 such as MinIO, Yugabyte, MySQL, and Kafka. Plugins 654 can be added to support various storage technologies. The processor 602 implements the data query interface 660 and the data calculation interface 670 by executing instructions 606. The processor 602 executes instruction 606 to receive information identifying data of interest in the storage system 690, generates an access request 634 for the data of interest in the storage system 690, automatically forwards an approval request 643 to one or more access control entities 692, such as a data owner 694 and a data governance officer 696, based on the access request 634, receives an approval response 644 from one or more access control entities 692, creates a policy 648 for accessing the data of interest in the storage system 690, and provides access to the data of interest in the storage system 690 based on the policy 648. The processor 602 executes instruction 606 to provide access verification for accessing the data of interest in the storage system 690 based on the policy 648.Processor 602 receives an input at the data query interface 660 for generating a query request 662, forms a search input for metadata and data definitions 632 from the data discovery and catalog interface 630 to identify a dataset based on the input for generating the query request 662, generates a query request 662 to access the dataset in the storage system 690 in response to the metadata and data definitions 632, determines access rights at the policy engine 647 based on policy 648, and executes an instruction 606 to access the dataset in the storage system 690 granted according to the access rights of policy 648 determined based on the query request 662. Processor 602 executes instruction 606 and receives an input at the data calculation interface 670 for creating a first level report 672 for analyzing data usage in the storage system 690. Processor 602 executes instruction 606 to receive an acknowledgment response 644 from one or more access control entities 692 at the data access layer 642, and the policy engine 647 determines one or more of the following: authorized users for access to the data of interest in the storage system 690, the time period set to allow access to the data of interest, the scope of access to the storage system 690, restrictions on access to the storage system 690, or a mask to apply to the data. Processor 602 executes instruction 606 to cause the data storage layer 650 to parse event 656, extract log 658 associated with access to the data of interest in the storage system 690, and send log 658 to the audit interface 649 for auditing the use of the data of interest in the storage system 690 to audit the use of the data of interest in the storage system 690 based on a deterministic process. User interface (UI) 682 is presented on display 680 to allow the user to input and manage data.For example, UI 682 may include a data discovery and catalog interface 683 for receiving user input to generate an access request 634. A data query interface 684 allows the user to input information to generate a query request 662. A data calculation interface 685 allows the user to input data to create a first-level report 672 and make data usage analysis available. A policy interface 686 allows the user to configure policies 648 for accessing the data.
[0075] Embodiments described herein provide methods that offer one or more advantages. For example, a method for providing central data governance and access control for enterprise data provides a one-stop solution for data governance and access-related services. This reduces time and effort for end users because they do not need to create tickets to gain access to specific data and do not need to identify the data owner or data governance office for specific data. Communication with data owners and data governance officers is handled automatically. Trust and control are also provided through centralized and optimized auditing.
[0076] One aspect of this description relates to a method for providing central data governance and access control for enterprise data[1], which includes receiving information identifying data of interest in a storage system provided by a data platform; generating access requests for the data of interest in the storage system; automatically forwarding authorization requests to one or more access control entities based on the access requests; receiving authorization responses from one or more access control entities; creating policies for accessing the data of interest in the storage system; and providing access to the data of interest in the storage system based on the policies.
[0077] The method according to [1], further comprising the storage system performing access validation to access data of interest in the storage system based on a policy.
[0078] Providing access to data of interest in a storage system based on a policy includes providing access to data of interest in one or more storage sources, the one or more storage sources being accessible using one or more plug-ins, according to the method of [1] to [2].
[0079] The method according to [1] to [3] further comprises receiving input for generating a query, forming search inputs for metadata and data definitions from a data discovery catalog to identify a dataset based on the input for generating the query, generating a query request for accessing the dataset in a storage system in response to the metadata and data definitions, determining access rights based on a policy for accessing the dataset in the storage system, and accessing the dataset in the storage system granted according to the determined access rights based on the query request.
[0080] The method according to [1] to [4], further comprising receiving input to create a first-level report for analyzing data usage in a storage system.
[0081] The method according to [1] to [5], wherein receiving an authorization response from one or more access control entities determines one or more of the following: a user authorized to access data of interest in a storage system, a time period set to allow access to the data of interest, the scope of access to the storage system, restrictions on access to the storage system, or a mask to apply to the data.
[0082] The method according to [1] to [6], further comprising: analyzing events and extracting logs associated with access to data of interest in the storage system; sending logs to audit the use of data of interest in the storage system; and auditing the use of data of interest in the storage system based on a deterministic process.
[0083] One aspect of this description relates to a data platform [8] comprising a memory for storing computer-readable instructions and a processor attached to the memory, the processor being configured to execute computer-readable instructions, receive information identifying data of interest in a storage system, generate access requests for data of interest in the storage system, automatically forward authorization requests to one or more access control entities based on the access requests, receive authorization responses from one or more access control entities, create policies for accessing data of interest in the storage system, and provide access to data of interest in the storage system based on the policies.
[0084] The data platform described in [8] is further configured to have a processor that performs access validation to access data of interest in the storage system based on a policy.
[0085] The processor is further configured to provide access to data of interest in a storage system based on a policy by providing access to data of interest in one or more storage sources, and the data platform described in [8] to [9], wherein one or more storage sources are accessible using one or more plug-ins.
[0086] The data platform described in [8] to
[10] is further configured to receive input for generating queries, to form search inputs for metadata and data definitions from a data discovery catalog to identify datasets based on the input for generating queries, to generate query requests for accessing datasets in a storage system in response to the metadata and data definitions, to determine access rights based on a policy for accessing datasets in the storage system, and to access datasets in the storage system granted according to the determined access rights based on the query request.
[0087] The processor is further configured to perform operations to receive input for creating a first-level report for analyzing data usage in the storage system, as described in [8] to
[11] of the data platform.
[0088] The data platform described in [8] to
[12] , further configured to receive authorization responses from one or more access control entities by determining one or more of the following: authorized users for access to data of interest in the storage system, a time period set to allow access to the data of interest, the scope of access to the storage system, restrictions on access to the storage system, or a mask to apply to the data.
[0089] The processor is further configured to analyze events, extract logs associated with access to data of interest in the storage system, send logs to audit the use of data of interest in the storage system, and audit the use of data of interest in the storage system based on a deterministic process, as described in [8] to
[13] .
[0090] One aspect of this specification relates to a non-temporary computer-readable medium
[15] having computer-readable instructions stored thereon, which, when executed by a processor, causes the processor to perform operations including: receiving information identifying data of interest in a storage system provided by a data platform; generating an access request for the data of interest in the storage system; automatically forwarding an authorization request to one or more access control entities based on the access request; receiving an authorization response from one or more access control entities; creating a policy for accessing the data of interest in the storage system; and providing access to the data of interest in the storage system based on the policy.
[0091] The non-temporary computer-readable media described in
[15] further includes, by the storage system, performing access verification to access data of interest in the storage system based on a policy.
[0092] Providing access to data of interest in a storage system based on a policy includes providing access to data of interest in one or more storage sources, the one or more storage sources being non-temporary computer-readable media as described in
[15] to
[16] , which are accessible using one or more plug-ins.
[0093] A non-temporary computer-readable medium as described in
[15] to
[17] , further comprising: receiving input for generating a query; forming search inputs for metadata and data definitions from a data discovery catalog to identify a dataset based on the input for generating a query; generating a query request for accessing the dataset in a storage system in response to the metadata and data definitions; determining access rights based on a policy for accessing the dataset in the storage system; and, based on the query request, accessing the dataset in the storage system granted according to the determined access rights.
[0094] Non-temporary computer-readable media as described in
[15] to
[18] , further comprising receiving input for generating a first-level report for analyzing data usage in a storage system.
[0095] The non-temporary computer-readable media described in
[15] to
[19] further includes analyzing events and extracting logs associated with access to data of interest in the storage system, sending logs to audit the use of data of interest in the storage system, and auditing the use of data of interest in the storage system based on a deterministic process.
[0096] Separate instances of these programs can run or be distributed across any number of separate computer systems. Therefore, while certain steps are described as being performed by a specific device, software program, process, or entity, this is not necessarily required. Various alternative implementations will be understood by those skilled in the art.
[0097] In addition, those skilled in the art will readily recognize that the aforementioned technologies can be used in a variety of devices, environments, and situations. While the embodiments are described in language specific to structural features or methodological actions, the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms for implementing the claims.
Claims
1. A method for providing central data governance and access control for enterprise data, Receiving information that identifies data of interest in the storage system provided by the data platform, Based on the received information, the storage system generates an access request for the data of interest, The process involves automatically forwarding an authorization request to one or more access control entities based on the aforementioned access request, wherein the one or more access control entities include at least one of the data owner or the data governance officer. Receiving an approval response from each of the one or more access control entities, Based on the aforementioned acknowledgment response, a policy is created in the database and the storage system according to the type of data of interest, Applying a policy to access the type of data of interest in the database and the storage system, Based on the aforementioned policy, the storage system provides access to the data of interest. Methods that include...
2. The storage system performs access verification to access the data of interest in the storage system based on the policy. The method according to claim 1, further comprising:
3. The method according to claim 1, wherein providing access to the data of interest in the storage system based on the policy includes providing access to the data of interest in one or more storage sources, the one or more storage sources being accessible using one or more plug-ins.
4. Receiving input to generate queries, To identify a dataset based on the input for generating the query, search inputs for metadata and data definitions are formed from a data discovery catalog, In response to the metadata and data definition, generate a query request in the storage system to access the dataset, In the aforementioned storage system, access rights are determined based on the policy for accessing the aforementioned dataset, Based on the query request, access to the dataset in the storage system granted according to the determined access rights. The method according to claim 1, further comprising:
5. The method according to claim 1, further comprising receiving input for generating a first level report for analyzing data usage in the storage system.
6. The method according to claim 1, wherein receiving the authorization response from one or more access control entities includes determining one or more of the following: a user authorized to access the data of interest in the storage system; a time period set to enable the access to the data of interest; the scope of the access to the storage system; restrictions on the access to the storage system; or a mask to apply to the data.
7. The process involves analyzing the event and extracting logs associated with the data storage layer of the storage system's access to the event and logs, The analyzed events and extracted logs are transmitted to the data storage layer of the storage system. Auditing the analyzed events and extracted logs in the storage system based on a predetermined deterministic process. The method according to claim 1, further comprising:
8. A storage system receives information that identifies data of interest, Based on the received information, the storage system generates an access request for the data of interest. Based on the access request, the authorization request is automatically forwarded to one or more access control entities, and the one or more access control entities include at least one of the data owner or data governance officer. Upon receiving an approval response from each of the one or more access control entities, Based on the aforementioned acknowledgment response, a policy is created in the database and the storage system according to the type of data of interest: Apply a policy to access the type of data of interest in the aforementioned database and storage system. A data platform configured to provide access to the data of interest in the storage system based on the aforementioned policy.
9. The data platform according to claim 8, further configured to perform an operation to perform access verification for accessing the data of interest in the storage system based on the policy.
10. The data platform according to claim 8, further configured to provide access to the data of interest in the storage system based on the policy by providing access to the data of interest in one or more storage sources, wherein the one or more storage sources are accessible using one or more plug-ins.
11. Receiving input for generating a query, To identify a dataset based on the input for generating the query, search inputs for metadata and data definitions are formed from a data discovery catalog. In response to the metadata and data definition, the storage system generates a query request to access the dataset. In the aforementioned storage system, access rights are determined based on the policy for accessing the dataset. The data platform according to claim 8, further configured to perform an operation to access the dataset in the storage system granted in accordance with the determined access rights based on the query request.
12. The data platform according to claim 8, further configured to perform an operation to receive input for generating a first level report for analyzing data usage in the storage system.
13. The data platform according to claim 8, further configured to receive the authorization response from one or more access control entities by determining one or more of the following: a user authorized for the access to the data of interest in the storage system, a time period set to enable the access to the data of interest, the scope of the access to the storage system, restrictions on the access to the storage system, or a mask to apply to the data.
14. Analyze the event and extract logs associated with the access of the event and logs by the data storage layer of the storage system, The analyzed events and extracted logs are transmitted to the data storage layer of the storage system. The data platform according to claim 8, further configured to perform an operation to audit the analyzed events and extracted logs in the storage system based on a predetermined deterministic process.
15. A non-temporary computer-readable medium storing computer-readable instructions, which, when executed, Receiving information that identifies data of interest in the storage system provided by the data platform, Based on the received information, the storage system generates an access request for the data of interest, The process involves automatically forwarding an authorization request to one or more access control entities based on the aforementioned access request, wherein the one or more access control entities include at least one of the data owner or the data governance officer. Receiving an approval response from each of the one or more access control entities, Based on the aforementioned acknowledgment response, a policy is created in the database and the storage system according to the type of data of interest, Applying a policy to access the type of data of interest in the database and the storage system, Based on the aforementioned policy, the storage system provides access to the data of interest. A non-temporary computer-readable medium that enables the operation of including [a specific action].
16. The storage system performs access verification to access the data of interest in the storage system based on the policy. The non-temporary computer-readable medium according to claim 15, further comprising:
17. Providing the access to the data of interest in the storage system based on the policy includes providing the access to the data of interest in one or more storage sources, the one or more storage sources being accessible using one or more plug-ins, the non-temporary computer-readable medium according to claim 15.
18. Receiving input to generate queries, To identify a dataset based on the input for generating the query, search inputs for metadata and data definitions are formed from a data discovery catalog, In response to the metadata and data definition, generate a query request in the storage system to access the dataset, In the aforementioned storage system, access rights are determined based on the policy for accessing the aforementioned dataset, Based on the query request, access to the dataset in the storage system granted according to the determined access rights. The non-temporary computer-readable medium according to claim 15, further comprising:
19. The non-temporary computer-readable medium according to claim 15, further comprising receiving input for generating a first level report for analyzing data usage in the storage system.
20. The process involves analyzing the event and extracting logs associated with the data storage layer of the storage system's access to the event and logs, The analyzed events and extracted logs are transmitted to the data storage layer of the storage system. Auditing the analyzed events and extracted logs in the storage system based on a predetermined deterministic process. The non-temporary computer-readable medium according to claim 15, further comprising:
Citation Information
Patent Citations
Metadata Management System
JP2016520890A
JPP4472775B
Systems and methods for data storage and processing
US20200026710A1
Access control center auto configuration
US8850525B1
Blockchain-implemented system and method
WO2018020372A1