Policy-code for data assets and repairs in cloud environment
By adopting a policy-as-code approach in a cloud environment, and utilizing computer-coded policies to automatically manage and monitor data assets, the challenges of data policy coordination and compliance management in existing technologies are solved. This enables dynamic and real-time governance throughout the data lifecycle, improving the efficiency and accuracy of data management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AMAZON TECH INC
- Filing Date
- 2024-09-19
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies lack the ability to coordinate, manage, and monitor data policies in cloud environments, and cannot effectively support policy compliance and data lifecycle actions for different asset types, making it difficult to link data management tasks to compliance dashboards.
The strategy-as-code approach is adopted to dynamically manage and monitor data assets in the control plane of the cloud environment through computer-coded policies. It automatically associates data assets with policies, responds to violations in real time and executes remedial actions, and uses a computer-coded policy library to provide templates and rules to generate adaptive policies.
It enables dynamic and real-time governance of data policies, simplifies compliance management, ensures privacy compliance throughout the data lifecycle, reduces the need for manual coding, and improves the efficiency and accuracy of data asset management.
Smart Images

Figure CN121986330A_ABST
Abstract
Description
Cross-reference to related applications
[0001] This is U.S. non-provisional patent application No. 18 / 372,390, filed on September 25, 2023, entitled "POLICY-AS-CODE FOR DATA ASSETS AND REMEDIATION IN CLOUD ENVIRONMENTS", the entire contents of which are incorporated herein by reference for all intents and purposes. Background Technology
[0002] Data producers and consumers can benefit from intermediate data governance features comprised of hardware and software within a cloud environment. Data producers can be aggregators of data from diverse sources and accounts, or providers of such data used by data consumers for services such as marketing and analytics for end-users or clients. The data domains of intermediate data governance features allow data producers to catalog data according to business contexts such as sales, marketing, quality, and others. Data domains can support certain policies to provide data governance and can support subdomains for categories or end-users or clients to implement these policies. Policy-constrained data can be referenced through its data assets, which can be representations of the data's metadata, including table names, column names, column types, aliases, business descriptions, categories, etc. The data itself can be presented as data objects such as tables, dashboards, files, or virtual data objects. However, this approach to policy implementation still lacks support for centralized capabilities to coordinate, manage, and monitor data policies and data lifecycle actions. For example, this approach may not fully support policy compliance for different asset types, the invocation of appropriate data lifecycle actions, and reporting to dashboards that link data management tasks to compliance. Attached Figure Description
[0003] The embodiments described herein will be described with reference to the accompanying drawings, wherein: Figure 1 A block diagram of a system for performing corrective actions in response to violations associated with computer coding policies, according to at least one embodiment, is shown.
[0004] Figure 2 A block diagram illustrating interface details in a system for remedial actions in response to violations associated with computer coding policies, according to at least one embodiment, is shown.
[0005] Figure 3 A flowchart or method is shown for a system, according to at least one embodiment, for performing corrective actions in response to violations associated with computer coding policies.
[0006] Figure 4 Another flowchart or method is shown for a system, according to at least one embodiment, for remedial actions in response to violations associated with computer coding policies.
[0007] Figure 5 An example of an inclusive network computing environment in which various implementation schemes can be achieved is shown.
[0008] Figure 6 Example components of a server according to various implementations are shown, which can be used to perform at least a portion of a remedial action in response to a violation associated with a computer coding policy.
[0009] Figure 7 Example components of a computing device are shown that can be used to implement various implementation schemes for input, processing, monitoring, and other aspects. Detailed Implementation
[0010] The systems and methods according to at least one embodiment described herein overcome one or more of the aforementioned deficiencies, as well as other such deficiencies in methods of data governance using code (referred to herein as "policy as code"), to provide the ability for dynamic and real-time data governance of the data plane by the control plane. In at least one embodiment, such a system includes at least one processor to execute instructions from memory to cause the system to receive a computer-coded policy for execution in the control plane associated with a cloud environment. The computer-coded policy may be associated with data governance of one or more data assets in the data plane using the cloud environment. The system automatically associates one or more data assets with the computer-coded policy. For example, the system may use predefined rules associated with the computer-coded policy, as well as annotations associated with one or more data assets. Such a system may support dynamic changes to one or more data assets, in part based on real-time changes to the computer-coded policy. Such a system may also include functionality to monitor one or more data assets according to the computer-coded policy. The system may perform remedial actions associated with one or more data assets in response to violations associated with the computer-coded policy.
[0011] In at least one implementation, this dynamic and real-time data governance approach provides the capability to streamline policy compliance (including compliance with data privacy regulations throughout the data lifecycle) in a cloud environment. The system's computer-coded policy extends the dynamic and real-time data governance features of this paper to allow the writing of computer-coded policies to manage data, monitor policy compliance, and take remedial actions, all of which can include dynamic and real-time data responses throughout the data lifecycle. As the computer-coded policy is defined, the dynamic real-time data governance system can automatically match the policy parameters of the computer-coded policy by executing it in the cloud environment to identify data assets in the data plane and invoke tests (including user-defined tests) and actions to monitor and remediate identified issues. The system benefits from built-in reporting, which, for example, provides a comprehensive view of the compliance status of data assets for data consumers. Furthermore, to support the policy-as-code approach of this paper, a library of computer-coded policies provides templates for one or more of the following: retention rules, deletion rules, data filtering rules, data disclosure rules (including data subject access requests), data sovereignty rules, or territorial rules, enabling the generation of various variations of the computer-coded policy to be executed in the cloud environment.
[0012] In at least one embodiment, a computer-coded policy provides a policy-as-a-code tool that enables policy administrators to define data lifecycle policies that can be used to ensure the compliance of data assets with respect to one or more policies, including data privacy regulations or requirements arising from data sharing or commercial contracts. In at least one embodiment, the systems and methods described herein support and build data discovery, classification, collaboration, and access characteristics using data domains and data assets within those domains through computer-coded policies executed in a control plane associated with a cloud environment, wherein the computer-coded policy is associated with data governance performed using one or more data assets in the data plane. Therefore, computer-coded policies, including retention rules and deletion rules, can be applied as policy parameters (such as including retention periods for retention rules and deletion periods for deletion rules) to metadata associated with data assets. The computer-coded policy is enforced by enforcement and monitoring modules across the control plane and data plane to ensure compliance. Thus, one or more data assets can be automatically associated with a computer-coded policy using predefined rules and annotations associated with the policy.
[0013] In at least one implementation, policy administrators can benefit from testing capabilities that also allow for real-time updates to computer-coded policies and, in part, dynamic changes to one or more data assets based on these real-time updates. Furthermore, readily available computer-coded reporting templates can be populated for use in internal compliance processes. The systems and methods described herein address the challenges faced by policy-based systems that may lack standardized workflows for provisioning and managing data policies for entities throughout the data lifecycle. This can include declarative actions that relate to how data is collected, stored, used, archived, and disposed of. Additionally, the systems and methods described herein can be integrated with underlying components such as compute, storage, and identity resources, as further detailed herein, to provide reporting as part of native analytics services. Methods for performing remedial actions in response to violations associated with computer-coded policies allow data publishers or producers to associate policies with data assets without understanding how the policies apply to the data assets and without having to search thousands of datasets to discover relationships—a potentially cumbersome and error-prone process. Furthermore, policy administrators or data users may wish to apply policies to data assets and can achieve this through the methods described herein. These approaches using computer coding strategies eliminate the need to write code solely for testing and remediation actions that may fail to keep pace with evolving data environments, regulatory changes, and new business contracts requiring modifications to preserve data. Furthermore, the methods described in this paper address the issue of auditors lacking effective methods for reviewing compliance status and potentially relying on manual reporting by providing notifications and reports as part of changes to computer coding strategies or data assets.
[0014] Figure 1A block diagram 100 of an example system for performing remedial actions in response to violations associated with a computer coding policy, according to at least one embodiment, is shown. The system includes at least one processor and memory having instructions that, when executed by the at least one processor, enable one or more of the modules and features 104 to 126 described herein. A policy administrator 126 can provide a computer coding policy using a service provider's administrator or administrator interface 126, which supports specific remedial actions for one or more data assets. The computer coding policy supporting specific remedial methods enables resilient workloads or applications for data consumers. For example, such remediation may include requesting a computer coding policy with an extensible policy language and invoking certain data lifecycle actions for remediation. Furthermore, such remediation may be associated with an undefined relationship between data policies, data storage, and datasets. In at least one embodiment, the systems and methods described herein enable real-time dynamic changes to computer coding policies and annotations, and enable monitoring and testing of the content of data assets based on the automatic association between computer coding policies and data assets, so as to take remedial actions when a violation of the computer coding policy occurs regarding the content of the data assets.
[0015] Cloud environment 102 may include a data plane 104 and a control plane 106. Control plane 106 provides application programming interfaces (APIs) that can be associated with resources in data plane 104. These resources may include computing assets that can be used to process or store data from producer 108 and to provide data to consumer 110. These resources are created, updated, deleted, listed, or used in a manner suitable for the application. In one example, control plane actions may include launching compute instances within cloud environment 102, creating storage services, and describing service queues. Furthermore, after launching compute instances, control plane 106 may first execute tasks associated with determining physical hosts capable of performing tasks, determining network interface allocations, preparing storage volumes within cloud environment 102, and generating secure access credentials. Thus, control plane 106 can implement workflows, business logic, and database systems. Data plane 104 provides tasks or workload aspects or functions of the services associated with control plane 106 and data plane 104. In at least one implementation, the data plane 104 can be used to execute compute instances after the control plane 106 is created, to read and write storage volumes or cloud storage buckets, and to perform network query routing and health checks.
[0016] In at least one embodiment, a data domain may correspond to a virtual portion of a data plane to which a data producer can provide the aforementioned data from producer 108 for cataloging by business context for entities or organizations. Thus, in at least one embodiment, different entities in cloud environment 102 can represent their own organizational hierarchy within a data domain within a data domain. Furthermore, some data domains within a data domain may include, for example, subdomains of subsidiaries or affiliated entities. Different data domains may be associated with different data (such as data from producer 108) that may exist in any account or region. In the example, such data from producer 108 may be associated with multiple sources (such as the aforementioned accounts or regions) of an entity's data publisher. A data publisher or producer can publish data asset 116 to directory 114, which is used for data also stored in cloud environment 102 and belonging to one or more data producers. Directory 114 (described in further detail below) may be a data structure that provides context as part of a data domain to the data. For example, business data from producer 108 may be divided into company data and other data, where company data may be further cataloged as sales, finance, marketing, etc.
[0017] In each of these cataloging divisions, the organizational domain can be defined as the highest level of the catalog, followed by a business terminology table, which can be further associated with metadata as a subcategory, which in turn is associated with the finest granularity, namely data asset 116. Thus, for an entity, data from producer 108 can be cataloged in this manner. For example, an entity may have organizational domains of sales, finance, and marketing (as described above), which can be further subdivided into business terminology tables. Sales-related business terms may involve sales in different geographic regions; while finance-related business terms may involve different financial assets. Then, the sales metadata subcategory as a business terminology table may involve different sales forms from different geographic regions; or, the finance metadata subcategory as a business terminology table may involve specific financial forms of different financial assets. Data assets (which are the finest subcategory under sales metadata) then provide certain sales or product tables for different sales forms; or, for finance metadata, certain financial tables for different financial forms.
[0018] Therefore, in at least one implementation, data asset 116 can be tables, files, dashboards, etc., which can reference aspects of data from producer 108, which can be cataloged into a data domain. These data assets 116 can include content 112, such as table names, table descriptions, and schemas (such as column names, column types, column descriptions, etc.), which will be monitored to determine compliance with computer coding policies. Data asset 116 can be associated with extensible business metadata, such as business descriptions, business aliases, data sensitivity, data classification, etc. A metadata subdirectory (also known as a metadata form) provides entry or input forms that can be used to set recommended or required fields. This allows data producers or publishers from different regions and accounts to publish data to the directory in a consistent format. Data producers or publishers can follow recommended or required fields in the template that can be additionally associated with recommended values to publish their data.
[0019] In at least one implementation, for organizational domains associated with sales, metadata can be set to ensure that different sales regions can provide any data assets to catalog 114. Subsequently, data from producer 108 can include its corresponding attributes, including sales revenue, sales quarter, etc. To ensure consistency of values, metadata, and data assets, values in the metadata form and data assets are divided into different columns, some of which may be sensitively correlated with other columns. In at least one implementation, the system and method described herein support the discovery of data from producer 108, thereby enabling data consumers to search for and easily locate data assets 116 of interest, which can then be provided as data to consumer 110.
[0020] In at least one implementation, one or more data domain items may be provided to aggregate data from producer 108 and consumer 110 using computer-coded policies from policy interface 124 of control plane 106. Each item allows a group of users (such as network service clients 222) to aggregate data from producer 108 and consumer 110. Figure 2In this project, data consumers (who may be data consumers) and data owners (who may be data producers) collaborate for business or other needs, where such collaboration may include publishing, discovering, subscribing to, and using data assets 116 in the catalog 114. In at least one embodiment, each party in the project may be associated with access controls from the data governance module 118 of the control plane 106. Access controls can be applied to enable authorized individuals, groups, and roles to access the underlying project, and consequently, the underlying data assets of that project. Furthermore, access controls can ensure that only those tools defined by the project's permissions can be used with, for example, the underlying data assets. In at least one embodiment, the project can act as an identity subject that receives access authorization for the underlying resources associated with the data assets, and can enable the catalog to operate within computing assets or infrastructure associated with relevant data from the data asset producer 108, without relying on credentials from other individual users.
[0021] In at least one implementation, each component of the project can be used to manage data access for teams and groups. For example, at least one project may be subject to internal access controls that restrict access to the project and its data assets to authorized individuals, groups, and roles only, and that only tools (or capabilities) configured within the project can be used with it. Furthermore, project profiles can be defined as multiple sets of pre-configured resources and functionalities that provide reusable templates for creating projects. In at least one implementation, project profiles define settings such as network service accounts or Virtual Private Clouds (VPCs) in which projects are deployed.
[0022] In at least one implementation, a project has functionality representing a stack of ready-to-use deployment and configuration parameters that can be initiated during project creation. This allows one or more project functionalities to be enabled using a project profile and enables project creation. The project functionalities enabled in the project can be defined or described as tools and services for project members, making them available to process data assets in catalog 114 using computer-coded policies. In at least one implementation, data governance module 118 enables data consumers, producers, or policy administrators to streamline access governance by partitioning data domains (supporting data specialists), projects (supporting data consumers), and subscription approvals (supporting data producers). Data producers can share data assets to satisfy data consumer requests for data access to underlying data from producer 108. In at least one implementation, policy administrator 126 can be an entity different from or the same as a data consumer or data producer to limit the lifecycle of data assets shared between them using computer-coded policies. Policy administrator 126 can provide computer-coded policies to policy interface 124 or associate them via administrator or management interface 126.
[0023] In at least one embodiment, the data plane interface 120 allows for the enforcement and monitoring of at least one computer-coded policy associated with a data asset and a policy interface 124, the data asset being subject to that at least one computer-coded policy. For example, the data plane interface 120 coordinates with the data governance module 118 to provide monitoring of one or more data assets. Monitoring is to ensure that one or more data assets comply with the computer-coded policy. Furthermore, the coordination between the data plane interface 120 and the data governance module 118 is to perform remedial actions associated with one or more data assets in response to violations associated with the computer-coded policy. For example, remedial actions may be metadata or configuration of the data asset's environment. In at least one embodiment, the configuration of the environment includes changes to access controls on the availability of the underlying data and changes to sovereignty associated with the underlying data. Additionally, the administrator or management interface 126 may be a standalone or browser-based web application to support various users, including web service clients 222, data owners 218, policy administrators 126, and compliance reporters 224. For example, users may be able to access catalogs, discover data assets, manage data assets, share data assets and computer-coded policies, and analyze data in a self-service manner. In at least one implementation, the administrator or management interface 126 uses credentials from the identity provider to authenticate the user.
[0024] Figure 2 A block diagram 200 shows interface details in a system for remedial actions in response to violations associated with computer coding policies, according to at least one embodiment. Figure 2 The interface details in the document are at least as follows: Figure 1 The system's data plane interface 120 and policy interface 124 are related. In at least one embodiment, the system in block diagram 200 includes at least one processor and a memory storing instructions that, when executed by the at least one processor, cause the system to execute the interface details described herein. The system can receive computer-coded policies 212 from a policy administrator 126. In at least one embodiment, policy templates may exist in a computer-coded policy library 226 or a dataset association module 210, which can be modified to suit data assets. The policy administrator 126 can write computer-coded policies that can be automatically associated with computing resources and / or data assets.
[0025] In at least one implementation, tests can be performed in the form of executable code to test the computer coding policy in a test environment before deploying a released version of the computer coding policy 212. However, these tests can also be part of enforcement and compliance in a live environment where the computer coding policy is released and deployed. As part of the testing, test parameters can be provided or generated for the computer coding policy 212 for testing purposes. The control, scheduling, and orchestration module 206 is adapted to perform the tests. For example, this module can invoke customer-defined tests and actions applied to at least one computer coding policy and its associated data assets in the test environment. For example, the control, scheduling, and orchestration module 206 can also monitor and remediate identified issues before releasing the computer coding policy 212 to the live environment. The tests can include test parameters that define one or more compliance thresholds for one or more data assets. Multiple remediation actions can then be generated or provided in the control, scheduling, and orchestration module 206 for one or more data assets. For example, at least one of the remediation actions can be executed against one or more data assets and can be based on a violation of at least one of the one or more compliance thresholds.
[0026] In at least one implementation, Cedar® policy coding can be used to write computer coding policies 212. Cedar® provides verified permission codes that are scalable and include fine-grained permission management and authorization for custom applications. Verified permission codes are authorization-related by verifying whether policy administrator 126 is permitted to perform actions on applications and resources in a given context. Therefore, verified permission codes assume that policy administrator 126 has been identified and authenticated by any authentication solution associated with cloud environment 102. Furthermore, verified permission codes ensure that computer coding policies are at least independent of the type of authentication used.
[0027] In at least one implementation, validated permission coding enables developers to build secure applications more quickly through externalized authorization and centralized policy management / administration. The validated permission coding aspects used in Cedar® allow for the definition of fine-grained permissions for policy administrators, application users, or other consumers of data from producer 108. For example, a computer-coded policy 212 is provided for data assets to ensure that only authorized users can access the dataset or computer-coded policy 212. Furthermore, authorization can also ensure that policy administrators are restricted to certain policy functions. Cedar® decouples business logic from policy logic, allowing policies to include a pre-request to the Cedar® authorization engine to first verify that the policy and its associated aspects have been authorized. The computer-coded policy can then be provided as instructions (with a set of predefined rules encoded in these instructions from policy parameters) to the dataset association module 210, which can enforce the policy on the underlying dataset. For example, predefined rules and identifiers of the associated dataset, such as the user subject or resource (as detailed in Table 1), are provided to catalog 114. Catalog 114 may include data assets with annotations 116A from the underlying data of producer 108, which are constrained by one or more policy parameters 128 of computer coding policy 212.
[0028] In at least one embodiment, the computer coding policy 212 can be used via a software development kit (also known as the policy execution module 214). In at least one embodiment, product module 228 provides an abstraction layer that serves as an intuitive user interface within product module 228 to support the writing of the computer coding policy 212. Furthermore, policy execution module 214 allows the execution of the computer coding policy 212 in a control plane associated with a cloud environment. Instructions may be generated by policy execution module 214 in part based on the execution of the computer coding policy in the control plane.
[0029] In the example, computer coding policy 212 can be associated with one or more data assets in a cloud environment and the data governance module 118 in the data plane. Data governance, as used herein, is at least associated with the people, processes, and technologies required to satisfy an entity's data policies, such as those defined in part by data lifecycle management and data policy oversight. While these are subsets of the capabilities an entity can take in data governance, there may be other capabilities as part of a broader data governance implementation to ensure that the computer coding policy addresses at least one or more issues faced by the entity's data policies. As part of data lifecycle management, instructions in the policy enforcement module 214 can be used to execute modifications in annotations 116A of the data assets to enforce remedial actions associated with one or more data assets.
[0030] In at least one implementation, data producers can natively publish structured data assets (such as in XML®, JSON®, CSV, or other formats). These data assets are published to a catalog based on their data source from producer 108. For example, a reservation table can be used to ingest data from the data producer's sources into its appropriate catalog. Data consumers can then access their data assets using, for example, the reservation table. One or more of the data producers, data consumers, network service customers, or policy administrators can manage permissions on the reservation table. When computer-coded policies are defined at the metadata layer, testing and remediation actions can be used to implement underlying policy parameters on the data itself within the data resources associated with the relevant data assets. Therefore, as long as data assets are managed by metadata, policies and guidelines can be defined to handle these data assets, as will be further detailed in later aspects of this document.
[0031] Although described in relation to Cedar®, in at least one implementation, data producers or consumers do not need to write computer-coded policies in Cedar®, but can use other policy languages such as Rego®. In at least one implementation, policy interface 124 also addresses privacy and data challenges by automatically attaching policies to business and other metadata and by continuously testing policies against computer-coded policies. Thus, one or more data assets can be automatically associated with computer-coded policies. For example, predefined rules associated with computer-coded policies and annotations associated with one or more data assets can be used to establish such associations. Due to such associations, in dataset association module 210, one or more dynamic changes can be performed on the annotations of one or more data assets, in part based on real-time changes to the computer-coded policies.
[0032] In at least one implementation, the customer / data owner or policy administrator 222; 218; 126 can use logical expressions based on a business terminology table to set matching rules for predetermined rules associated with a computer coding policy. This can be performed using the computer coding policy module 212 or the computer coding policy library 226. In one example of a computer coding policy provided in Table 1, the matching rule for the computer coding policy may include the statement “Purpose is Provide_Ongoing_Service”, where “Purpose” can be a business terminology table expression, and “Provide_Ongoing_Service” can be a sub-item of the “Purpose” expression indicating that “Personal Data” is associated with a service and should be retained while the service is “ongoing”. Example computer coding policies may be provided for each entry in a retention period table and may be provided from computer coding policy templates in the computer coding policy library. Such computer coding policies may include a business terminology table for defining matching rules and may additionally include the relevant retention period as a policy parameter for each policy.
[0033] Table 1
[0034] Regarding data assets, applications may include strategies with multiple "Purposes," multiple data subjects (such as children and different age groups), and multiple data states (anonymous or non-anonymous). Besides "ongoing" service as a "Purpose," other purposes may include product improvement, personalization, marketing, security, and litigation. One or more of these "Purposes," data subjects, and data states can be associated with a business glossary and thus provided in annotation templates to guide the generation and / or population of data assets.
[0035] The policy enforcement module 214 can manage missing data errors, add fully qualified namespaces, and create computer-coded policies written in Cedar® or other policy languages. When executing computer-coded policies such as those in Table 1, the “Purpose” expression can be used to identify data assets using annotations. For example, a “Purpose” expression can be split into two policy parameters (one for anonymization and one for not using syntax statements). The “Action” argument in the example computer-coded policy refers to a repair action. The example computer-coded policy in Table 1 can extract data assets for which “Purpose” is “product improvement”, as described in its corresponding annotation (which may also include semantic variations). The computer-coded policy’s predefined rules (such as retention rules (coded as policy parameters 2.5 years or 3 months) and deletion rules (in addition to the explicitly included retention years and months)) are then automatically associated with the extracted data assets, ensuring compliance with these policy parameters.
[0036] Furthermore, this paper implements a two-way review of the compliance and enforcement of computer coding policies or one or more data assets. For example, on the one hand, it allows dynamic changes to one or more data assets, where real-time changes to the computer coding policy may alter retention or deletion rules, leading to dynamic changes to one or more data assets to apply new retention, deletion, data filtering, data disclosure rules (including data subject access requests), data sovereignty rules, or territorial rules. On the other hand, it monitors one or more data assets according to the computer coding policy, and when violations associated with the computer coding policy are identified, remedial actions can be performed on one or more data assets in response to the violations. Therefore, it not only allows dynamic changes to data assets to make them compliant with changes to the computer coding policy, but also continuously monitors data assets to ensure compliance regardless of whether the computer coding policy changes in real time.
[0037] In one example, enforcement and monitoring module 202 can span the data plane and control plane, which have sub-modules for control, scheduling, and orchestration module 206, control plan manager module 208, and control manager module 216 to provide such monitoring and dynamic changes to one or more data assets. Changes can then be performed on one or more data assets when it can be determined that a change has occurred in the computer coding policy (such as the addition of new policy parameters or the deletion of existing policy parameters). These changes could include, in part, deleting, retaining, filtering, disclosing (including by providing access) the content of the data asset, subjecting it to data sovereignty, or subjecting it to territorial rules. At a minimum, deletion can serve as a remedial action on information associated with one or more data assets in response to a violation related to the computer coding policy.
[0038] In at least one implementation, the enforcement and monitoring module 202 can monitor the content of one or more data assets according to a computer coding policy and can perform remedial actions associated with one or more data assets in response to violations associated with the computer coding policy. In one example, the policy administrator 126 can formulate a computer coding policy 212 by providing predefined rules and a semantic description of the data type to be applied. Data types may include geographic information, business domains, and the business purpose for which such data is discovered. The semantic description may be associated with semantic variations of the data type, such that literal matching is not a necessary condition for using the predefined rules herein.
[0039] The predefined rules may also include matching rules, partially based on semantic descriptions, to discover data assets 116 to which the computer coding strategy (after being written) should apply. Therefore, the strategy interface 124 may be able to define strategy types and matching rules to make them part of the computer coding strategy. Then, using the semantic subsystem of the dataset association module 210, one or more data assets associated with the strategy type and matching rules can be identified. Data assets can be provided to the dataset association module 210 from the catalog 114. In at least one embodiment, the computer coding strategy includes at least matching rules and compliance thresholds to support different testing and remediation actions on the data assets to measure and meet the compliance thresholds.
[0040] Further predefined rules can describe the maximum timeframes and parameters that data must adhere to, including retention schedules. Executable code can be associated with computer coding policies, but unlike the policy interface, this executable code can measure compliance with these rules. For example, control, scheduling, and orchestration module 206 can monitor violations and changes to computer coding policies, enabling dynamic changes to one or more data assets to be ensured.
[0041] In one example, a policy state change, partially based on a change in computer coding policy 212, can be passed to the control management module 216, which can cause the notification module 220 to notify one or more of the customer / data owners 222, 218, or policy administrator 126. Similarly, changes in the association between computer coding policies and data assets can be similarly passed from the dataset association module 210 to the control manager module 216, allowing the notification module 220 to notify one or more of the customer / data owners 222, 218, or policy administrator 126. Additionally, these updates or changes can be passed to the control, scheduling, and orchestration module 206 to allow the reporting module 204 to report compliance changes and to monitor for compliance threshold violations. The control, scheduling, and orchestration module 206 allows the execution of planned control runs to check compliance with inputs from network service customer 222.
[0042] In at least one implementation, when it is determined that a specific directory 114 is used to obtain data from producer 108 of an entity, users associated with that entity can publish data assets and enrich these data assets using business context. Users can also securely share analytics, data science, or machine learning associated with the data in producer 108. Furthermore, data governance from the provided data governance module 118 is set with specific privacy requirements applicable to the data in producer 108 of the entity. The basis for managing the data in producer 108 throughout its lifecycle, at least in part based on regulations, contracts, and other directives, may include considerations of what data in producer 108 is collected, where and how the data in producer 108 is stored, how it is used, and ensuring that the data in producer 108 is properly retained or deleted.
[0043] In at least one implementation, an entity may have a default retention policy that requires all its subscribers to use only data collected within the last 30 days, unless that data is used solely for product diagnostic purposes. Therefore, a computer-coded policy may be needed to delete data older than 30 days. The computer-coded policy can be written for data processing and defined by data governance or privacy aspects associated with the data. Data producers or publishers do not need to understand how the computer-coded policy can be applied to data assets and do not need to search thousands of datasets to discover relationships. Instead, metadata forms are used to collect data governance or privacy aspects of the entity. Data producers or publishers may include annotations on their data assets using these metadata forms. The computer-coded policy may then include a policy enforcement module 214 that discovers a standard for the relationship between the computer-coded policy and data assets by partially enforcing the computer-coded policy, which allows one or more data assets to be automatically associated with the computer-coded policy using predetermined rules associated with the computer-coded policy and annotations associated with one or more data assets.
[0044] For example, a computer coding policy includes policy parameters such as retention periods and deletion schedules. The policy administrator 126 can also define guidelines within the computer coding policy to instruct how specific policies within the computer coding policy should be implemented. Furthermore, data producers or publishers can use these guidelines to prepare computer codes for compliance monitoring and to execute remedial actions that can be invoked if data assets do not conform to the policy parameters of the computer coding policy. In the example, a data consumer or subscriber processing data in cloud storage and having Structured Query Language (SQL) access can provide the data to be cataloged to directory 114.
[0045] To implement a default retention policy, a computer-coded policy can be written, specifying a retention period of 30 days. The computer-coded policy may include predefined rules that set matching rules for data assets not annotated for product diagnostic purposes, as shown in the example above. Policy enforcement module 214 can execute the matching rules to determine which data assets are within the policy's scope. Data producers or publishers can also write further tests and actions before publishing the computer-coded policy. For example, a data producer or publisher can prepare computer-coded remediation actions to read the storage object of the data asset, check if the collection date of each row has expired, delete the data if it exceeds 30 days, and create new data objects. In at least one implementation, policy enforcement module 214 can also automatically detect conflicts within and across different computer-coded policies using policy parameters therein. In one example, a report can be generated listing more than one retention policy parameter with the same matching rule in the predefined rules. Conflicts must be resolved before writing and activation can proceed.
[0046] Furthermore, data producers or publishers can prepare computer-coded remediation actions in the form of SQL scripts that delete rows older than 30 days from the affected tables forming the data assets. Further, data producers or publishers can prepare computer-coded tests that check for the presence of data older than 30 days in the data objects and data assets. In at least one implementation, data consumers or subscribers can attach the computer-coded tests and computer-coded actions to their working copies. In at least one implementation, notification module 220 sends notifications to the accounts of policy administrator 126, compliance reporter 224, network service customer 222, or data owner 218. The notifications can trigger remediation actions according to a schedule configured by any of the policy administrator 126, compliance reporter 224, network service customer 222, or data owner 218 using control, scheduling, and orchestration module 206.
[0047] Alternatively, data publishers can leverage the sample functional workflow to trigger one or more of the automation processes in the policy interface to generate computer-coded policies, monitor data assets, and perform dynamic changes to data assets, as described throughout the document. The Policy-as-a-Code feature in this document automates policy enforcement and reduces operational overhead for users such as data publishers. When a client adds, deletes, and updates data assets, the computer-coded policy is automatically matched to the data assets using the policy parameters defined therein and annotations within the data assets. The system then begins monitoring policy compliance and publishing reports and underlying data, as new data assets may be included in catalog 114. The system can provide such reports using the reporting module 204, enabling policy administrators to use and perform additional analyses.
[0048] In at least one implementation, once the association of data assets with a computer coding policy is executed, guidelines for implementing the computer coding policy can be described in control plan manager module 208, and these guidelines can be managed by control manager module 216. These guidelines are stored in control plan manager module 208 in the form of a control plan. The control plan can be conceived as a functional contract that defines how data processing requirements (such as data retention) are implemented according to regulations, commercial contracts, or company policies. The control plan can provide coding syntax statements regarding: the intent of what policy is expected to be implemented; the tests and actions regarding how to execute the intent in the policy; whether the control plan applies when the policy is running in an active and / or shadow state; and the frequency for running the tests and actions. The computer coding policy can be associated with matching rules and the control plan before activation. Data producers or publishers can then receive notifications related to the computer coding policy from event and task notifications.
[0049] In the example, events and tasks may be triggered as part of other actions performed by the control, scheduling, and orchestration module 206. These tasks may indicate that a retention governance policy has been activated for a data domain, particularly for data assets controlled by entities in a cloud environment. These tasks may also include indications that controls (such as tests or actions) need to be registered to enforce computer-coded policies, such as retention policies. Policy administrators may be able to review the association between computer-coded policies and data assets, based partly on comments in the catalog and partly on reports from the authors of the computer-coded policies. Policy administrators may manage policy controls and control plans, partly using the control plan manager module 208 and the control manager module 216, and may associate control plans with computer-coded policies.
[0050] In at least one implementation, different policy states can be provided as part of the policy parameters of the computer-coded policy. One or more of the policy administrator, data producer (or owner), or data consumer can be able to set triggers that can be issued to coordinate the execution of the control plan across the control, scheduling, and orchestration module 206, the control plan manager module 208, and the control manager module 216 at a planned time. Furthermore, results can be written back to the reporting module 204. Control blueprints can be provided to simplify the receipt of trigger notifications to run controls, protect control parameters, and write completion status to the reporting module 204.
[0051] Figure 3 A flowchart or method 300 is shown, illustrating a remedial action performed in response to a violation associated with a computer coding policy, according to at least one embodiment. One or more steps in method 300 may be performed on a client device, or using a client device in conjunction with a cloud environment remote from the client device, such as... Figures 5 to 7 As shown and described. Method 300 includes receiving 302 computer-coded policies for data governance using one or more data assets within a cloud environment. For example, the method includes receiving computer-coded policies from a policy administrator using an administrator interface, or using policy templates from a computer-coded policy library or dataset association module, which can be modified to suit the data assets. In one example, policy administrator 126 can use policy templates to write computer-coded policies.
[0052] Method 300 includes providing 304 predetermined rules associated with a computer coding strategy. For example, policy parameters of the computer coding strategy (as shown in Table 1) are used to provide predetermined rules, such as retention rules and deletion rules, by executing the computer coding strategy. A further step in the method includes providing 306 annotations associated with one or more data assets. A further step includes determining or verifying 308 the association between the predetermined rules and the annotations. For example, the application associated with the data assets may include a strategy with multiple “Purposes,” multiple data subjects, and multiple data states, and one or more of such “Purposes,” data subjects, and data states may be associated with a business terminology table and thus can be provided in the annotation template to guide the generation and / or population of the data assets.
[0053] Method 300 includes automatically associating one or more data assets to a computer coding strategy using one or more data assets. For example, the sample computer coding strategy (in Table 1) is capable of extracting a data asset where “Purpose” is “product improvement”, as described in its corresponding notes (which may also include semantic variations). Predefined rules of the computer coding strategy (such as retention rules (provided in the strategy as strategy parameters of 2.5 years or 3 months) and deletion rules (in addition to the explicitly included retention periods and months)) are then automatically associated with the extracted data asset.
[0054] Figure 4 Another flowchart or method 400 is shown, according to at least one embodiment, used by a system for remedial actions in response to violations associated with computer coding policies. One or more steps in method 400 may be performed on a client device, or using a client device in conjunction with a cloud environment remote from the client device, such as... Figures 5 to 7 As shown and described. Method 400 can be used... Figure 3 Method 300 is executed afterward. In at least one embodiment, method 400 includes determining 402 real-time changes to the computer encoding policy. For example, if the computer encoding policy is changed to include different deletion or retention rules, this would cause method 400 to enable 404 dynamic changes associated with one or more data assets, such as performing dynamic changes to annotations in part based on the real-time changes to the computer encoding policy.
[0055] Method 400 includes monitoring 406 the content of one or more data assets according to a computer coding policy as part of a two-way review for compliance and enforcement. Similar to changes in the computer coding policy, data assets may also change, and step 406 monitors such changes. Verification 408 can be performed on the results of monitoring step 406, such as monitoring step 406 indicating that the change to one or more data assets is unrelated to a real-time change in the computer coding policy. Method 400 includes determining 410 that, due to the change to the data assets indicated in step 408, one or more data assets have committed a violation associated with the computer coding policy.
[0056] Method 400 includes performing a remedial action 412 associated with one or more data assets in response to a violation related to a computer coding policy. For example, after monitoring step 406, it can be determined that a change has occurred to the computer coding policy, such as the addition of a new policy parameter or the deletion of an existing policy parameter, which requires changes associated with a data asset or some information therein. Changes can be performed on one or more data assets, such as changes to the metadata or environment configuration of the data assets. In at least one embodiment, the environment configuration includes changes to access control on the availability of the underlying data and changes to sovereignty configuration associated with the underlying data, based in part on the new policy parameters. The content of the data assets can then be further monitored to see the changes regarding the new policy parameters. At least further remedial actions associated with one or more data assets can be performed in response to a violation associated with the new policy parameters of the computer coding policy.
[0057] In at least one embodiment, method 400 is a computer-implemented method that includes, or includes, the further step of: enabling a computer-coded policy to include policy parameters associated with at least retention rules and deletion rules. In at least one embodiment, method 400 includes, or includes, the further step of: providing an interface to enable the definition of policy types and matching rules as part of the computer-coded policy. Furthermore, method 400 includes, or includes, the further step of: using a semantic subsystem to determine one or more data assets associated with the policy types and matching rules.
[0058] In at least one implementation, method 400 is a computer-implemented method that includes further steps or sub-steps: providing test parameters for a computer coding strategy. These test parameters may define one or more compliance thresholds for one or more data assets. Method 400 includes further steps or sub-steps: providing remedial actions for one or more data assets. At least one of the remedial actions may be performed on one or more data assets and may be provided from a predefined remedial action that may be based on a violation of at least one of the one or more compliance thresholds.
[0059] In at least one embodiment, method 400 is a computer-implemented method that includes, or includes, further steps or sub-steps: generating instructions in part based on computer-coded policies executed in the control plane of a cloud environment. Method 400 includes, or includes, further steps or sub-steps: performing deletions or additions in a retention table in part based on the instructions to enforce a remedial action associated with one or more data assets. Method 400 includes, or includes, further steps or sub-steps: making the remedial action a change in access control to the data storage of one or more data assets or performing a soft or hard delete to remove one of the non-compliant data items from one or more data assets.
[0060] In at least one embodiment, method 400 is a computer-implemented method that includes, or includes, the following steps: enabling a preview operation associated with a computer coding policy using a control plane interface. The computer coding policy may be applied to a representation of one or more data assets, such as a table, view, or dashboard. Method 400 includes, or includes, the following steps: providing results associated with remedial actions or violations for the representation of one or more data assets. Method 400 includes, or includes, the following steps: allowing the release of the computer coding policy to take action on one or more data assets.
[0061] In at least one embodiment, method 400 is a computer-implemented method that includes, or includes, the further step of: enabling one or more of the notifications to perform planned or triggered tests on a computer coding policy against a representation of one or more data assets. Here, the representation may include infrastructure built in a testing portion of a cloud environment (also known as a test environment). The representation is used to trigger actions, such as test actions and remediation actions. Furthermore, the representation is used to implement real-time changes to the computer coding policy. In at least one embodiment, method 400 is a computer-implemented method that includes, or includes, the further step of: enabling the triggered tests to be partially based on changes to one or more data assets during workload execution.
[0062] In at least one implementation, using Figure 1 and Figure 2This approach, 300, 400, performed by one or more aspects of the system, allows customers or administrators to create retention and deletion computer-coded policies in a cloud environment. These computer-coded policies specify policy parameters, such as retention rules and deletion rules, including deletion schedules. Such computer-coded policies also define matching rules for policy coverage. One or more of the policy interface 124 or data plane interface 120 then match the policy to the data assets using policy coverage criteria or predefined rules (containing policy parameters) and annotations from the data assets. Customers or administrators can also define guidelines for how to test policy compliance and remedial actions to be taken if necessary. In one example, a data publisher could use these guidelines to implement testing and remedial actions as part of the computer-coded policy.
[0063] Furthermore, once the computer code policy is activated, the enforcement and monitoring module 202 continuously assists customers in monitoring policy compliance within their data environment and enables notifications and changes to user accounts to facilitate the use of data assets. Notifications may include details of the affected data assets, which may be configured or modified to allow for remedial actions. For example, such changes may include deleting data, blocking data, or enforcing least privilege access. One or more of the policy interface 124 or data plane interface 120 allow customers and administrators to set and enforce data processing rules regarding when and how data is retained or deleted. In at least one implementation, while the computer code policy is written using the computer code policy library 226, the policy administrator can add a semantic layer to describe the data assets in the catalog. The policy interface 124 or data plane interface 120 can then use the semantic layer to enable the policy administrator to write computer code policies, discover associations between computer code policies and data assets, regardless of whether the data assets physically reside in a cloud environment.
[0064] In at least one implementation, customers and administrators can use custom tests to monitor and enforce computer-coded policies to identify issues and take remedial actions to correct any problems. Customers and administrators can manage their experience by linking existing computer-coded policies and having one or more of the policy interface 124 or data plane interface 120 automatically invoke remedial actions on their behalf. Furthermore, the automation described herein allows for remediation of policy outcomes based on notifications or signals. For example, a specific row in a retention table can be deleted to enforce retention. Thus, one or more of the policy interface 124 or data plane interface 120 utilize computer-coded policies to span cloud, on-premises, and other environments. Figure 1 and Figure 2 Other products related to the architecture provide a unified approach to data lifecycle management.
[0065] Figure 5 An example inclusive network computing environment 500 is shown, in which various implementations can be achieved. In some implementations, such an environment may be used to provide source servers to one or more customers or administrators of a resource provider as part of a shared or multi-tenant resource environment. Thus, the inclusive network computing environment 500 may be located in different geographic locations, enabling the policy interface 124 or data plane interface 120 to utilize one or more of these different geographic locations and the resources within their respective environments. For example, the provider environment 506 may be a cloud environment that can be used to provide cloud-based network connectivity to users, just as the policy interface 124 or data plane interface 120 can. These resources may also provide networking capabilities for one or more client devices 502 (such as personal computers), which may be able to connect to one or more networks 504 discussed herein.
[0066] In this example, a customer or administrator can use client device 502 to submit requests to service provider environment 506 across at least one network 504. This service provider environment may include one or more of a policy interface or a data plane interface (DP+Policy I / F) 510. The client device may include any suitable electronic device operable to send and receive requests, messages, or other such information over a suitable network and to transmit information back to the user of the device. Examples of such client devices include personal computers, tablets, smartphones, laptops, etc. The at least one network 504 may include any suitable network, including intranets, the Internet, cellular networks, local area networks (LANs), or any other such networks or combinations, and communication on this network may be enabled via wired and / or wireless connections. The service provider environment 506 may include any suitable components for receiving requests and returning information or performing actions in response to those requests. As an example, the service provider environment may include a web server and / or application server for receiving and processing requests and then returning data, web pages, video, audio, or other such content or information in response to those requests. The service provider's environment is protected so that only authorized users can access these resources.
[0067] In various implementations, the service provider environment 506 may include multiple types of resources that can be used by multiple users for various purposes. As used herein, computing and other electronic resources utilized in the network environment may be referred to as “network resources.” For example, these may include servers, databases, load balancers, routers, etc., which can perform tasks such as receiving, transmitting, and / or processing data and / or executing instructions. In at least some implementations, at least for a defined period of time, all or part of a given resource or set of resources may be allocated to a specific user or assigned to a specific task. Sharing these multi-tenant resources from the service provider environment is commonly referred to as resource sharing, web services, or “cloud computing,” among other such terms, and depends on the specific environment and / or implementation. In this example, the service provider environment includes multiple resources 514 of one or more types. For example, these types may include: application servers operable to process instructions provided by users, or database servers operable to process data stored in one or more data storage areas 516 in response to user requests. As known for such purposes, users may also reserve at least a portion of the data storage in a given data storage area. Methods for enabling users to reserve various resources and resource instances are well known in the art, so this document will not discuss the detailed description of the entire process or the explanation of all possible components.
[0068] In at least some implementations, a customer or administrator wishing to utilize a portion of resource 514 can submit a request, which is received by interface layer 508 of service provider environment 506. Interface layer 508 may include application programming interfaces (APIs) or other exposed interfaces 518 that enable users to submit requests to the provider environment. Interface layer 508 in this example may also include other components, such as at least one web server, routing components, load balancers, etc. When a request to configure resources is received by interface layer 508, information for that request can be directed to resource manager 514 or other such systems, services, or components configured to manage customer accounts and information, resource configuration and usage, and other such aspects. Resource manager 514 receiving the request can perform tasks such as authenticating the identity of the customer or administrator submitting the request and determining whether an existing account exists within the resource provider, wherein account data may be stored in at least one account 512 within the provider environment.
[0069] Customers or administrators can provide any of a variety of types of credentials to authenticate their identity to the provider. These credentials may include, for example, username and password pairs, biometric data, digital signatures, or other such information. The provider can verify this information based on the information stored for the customer or administrator. If the customer or administrator has an account with appropriate permissions, status, etc., the resource manager can determine whether sufficient resources are available to satisfy the administrator's request, and if so, can configure resources or otherwise grant access to corresponding portions of those resources for the user's use, for the amount specified in the request. For example, this amount may include the capacity to process a single request or perform a single task, a specified time period, or a cyclical / updatable period, and other such values. If the customer or administrator does not have a valid account with the provider, the account cannot access the type of resource specified in the request, or other such reasons prevent the customer or administrator from obtaining access to such resources, a message may be sent to the customer or administrator to enable the user to create or modify an account, or change the resource specified in the request, and other such options.
[0070] In at least one embodiment, the resources available for use by the client device 502 may include servers and other resources 514, 520, each server and other resource having at least one processor and memory. The memory includes instructions that, when executed by the corresponding processor, cause this document to... Figure 1 and Figure 2 One or more of the modules and storage areas described herein can enable remedial actions in response to violations associated with computer coding policies. The server can be used to issue application programming interface (API) calls to one or more other modules of the service provider environment 506, or to perform one or more functions associated with remedial actions in response to violations associated with computer coding policies. The policy interface or data plane interface (DP+policy I / F) 510 may be unique to each customer or administrator account and can be accessed from client device 502 by the customer or administrator using a command-line interface (CLI) or character-based user interface (GUI). Customers or administrators use... Figure 1 or Figure 2 One or more of the modules or storage areas in the system and one or more of the service provider environment 506 provide services or applications to the user (on the corresponding host) 522.
[0071] Once a client or administrator is authenticated, their account verified, and resources allocated, they can utilize the allocated resources for specified capabilities, data transfer volumes, time periods, or other such values. In at least some implementations, the client or administrator can provide a session token or other such credentials in subsequent requests to enable those requests to be processed within the session. The client or administrator may receive resource identity, a specific address, or other such information that enables client device 502 to communicate with the allocated resources without having to communicate with resource manager 514, at least until changes occur in account-related aspects, the client or administrator is no longer granted access to the resources, or other such changes.
[0072] The policy interface or data plane interface (DP+policy I / F) 510 may include a control manager module 216 for certain control aspects in this example, and may also support the functionality of a virtual layer of hardware and software components that handles control functions in addition to management actions, including configuration, expansion, replication, policy enforcement, compliance, monitoring, remediation, etc. The policy interface or data plane interface (DP+policy I / F) 510 may utilize dedicated APIs in the interface layer 508, where each API can be provided to receive requests for at least one specific action to be performed on the data environment, such as provisioning, expanding, cloning, or hibernating instances, and monitoring, enforcing, and remediating computer coding policies. Upon receiving a request for one of the APIs, the web service portion of the interface layer may parse or otherwise analyze the request to determine the steps or operations required to perform or process the call. For example, a web service call including a request to create a data repository may be received.
[0073] In at least one embodiment, the interface layer 508 includes a set of scalable, user-facing servers that can provide various APIs and return appropriate responses based on API specifications. The interface layer may also include at least one API service layer, which in one embodiment consists of stateless, redundant servers that process external user-facing APIs. The interface layer may be responsible for web service front-end features such as credential-based authentication of users or administrators, authorization of users or administrators, restriction of requests to the API server, input validation, and marshalling or unmarshalling of requests and responses. The API layer may also be responsible for reading database configuration data from and writing database configuration data to the management data store in response to API calls. In many embodiments, the web service layer and / or the API service layer will be the only externally visible components, or the only components visible to and accessible to the administrator or user of the services provided herein. As is known in the art, the servers of the web service layer may be stateless and horizontally scalable. The API servers, and persistent data stores, may be distributed across multiple data centers in a region, for example, enabling these servers to withstand the failure of a single data center.
[0074] Figure 6 An example resource stack 602 of virtual and physical resources 600 is shown, which can be utilized according to various implementation schemes, such as those that can be provided as part of a framework or environment to perform remedial actions in response to violations associated with computer coding policies, such as Figure 1 and Figure 2 As shown. For example, for a directory as application 632, hypervisor 618 can be used to perform tasks such as policy enforcement, monitoring, and compliance tasks, and these tasks can be performed against one or more instances 620, 622 of hypervisor 618. Resource stack 602 includes physical underlying resources such as a central processing unit (CPU) 612 for executing code to perform these tasks, a network interface card (NIC) 606 for transmitting network traffic, and memory for storing instructions and networking data. In some implementations, the entire machine can be allocated for these tasks, or only a portion of the machine can be allocated, such as allocating a portion of the resources as virtual resources in instance 620; 622, which can perform at least some of these tasks.
[0075] This resource stack 602 can be used to provide an allocated environment for an administrator (or a customer of a resource provider) who has configured an operating system on the resources. According to the illustrated embodiment, the resource stack 602 includes multiple hardware resources 604, such as one or more CPUs 612; solid-state drives (SSDs) or other storage devices 610; a NIC 606; one or more peripheral devices (e.g., graphics processing units (GPUs) 608); a BIOS 616 implemented in flash memory; a baseboard management controller (BMC) 614; and so on.
[0076] In at least one embodiment, hardware resource 604 resides on a single computing device (e.g., chassis). In at least one embodiment, hardware resource may reside on multiple devices, racks, chassis, etc. The virtual resource stack running on hardware resource 604 may include a virtualization layer, such as hypervisor 618, a first instance 620, and possibly a second instance 622 capable of executing at least one application 632. If hypervisor 618 is used in a virtualization environment, it can manage the execution of one or more guest operating systems and allow multiple instances of different operating systems to share the underlying hardware resource 604.
[0077] Instances 620 or 622 may include one or more virtualization or paravirtualization drivers 630 and may include one or more backend device drivers 626. When the operating system kernel 628 of the instance wants to invoke an I / O operation, the virtualization or paravirtualization driver 630 can perform the operation by communicating with the backend device driver 626. When the virtualization or paravirtualization driver 630 wants to initiate an I / O operation (e.g., send a network packet), kernel components can identify which physical memory buffer contains the packet (or other data), and the virtualization or paravirtualization driver 630 can either copy the memory buffer to a temporary storage location in the kernel to perform the I / O or obtain a set of pointers to the memory page containing the packet. In at least one embodiment, these locations or pointers are provided to the backend driver 626 of the host kernel 624, which can obtain data access permissions and transmit them directly to a hardware device (such as NIC 606) to send packets over the network.
[0078] It should be noted that Figure 6The resource stack 602 shown is merely one possible example of a set of resources capable of providing a virtualized computing environment, and the various implementations described herein are not necessarily limited to this particular resource stack. In the computing server, the BMC 614 can maintain a list of events occurring in the system, referred to herein as the System Event Log (SEL). In at least one implementation, the BMC 614 can receive the System Event Log from the BIOS 616 on the host processor. The BIOS 616 can provide system event data to the BMC via an appropriate interface (such as an I2C interface) using an appropriate protocol (such as the SMBus System Interface (SSIF) or an LPC-based KCS interface). As previously mentioned, examples of System Event Log events in the BIOS include uncorrectable memory errors indicating RAM module corruption. In at least some implementations, the System Event Log logged by the BMC on various resources can be used for purposes such as monitoring server health, including triggering manual component replacement or instance degradation when a SEL in the BIOS indicates a failure.
[0079] In at least one embodiment, certain portions of physical resource 600 will be inaccessible to the OS. For example, this could include at least a portion of BIOS 616. In at least one embodiment, BIOS 616 is volatile memory, such that any data stored in that memory will be lost if a reboot or power outage occurs. The BIOS may retain at least a portion of unmapped host memory, making it undiscoverable by the host OS. Computing resources such as servers, smartphones, or personal computers will generally include at least a set of standard components configured for general operation, although various proprietary components and configurations may also be used across different embodiments. As previously described, this could include client devices for transmitting and receiving network communications, or servers for performing tasks such as network analysis and rerouting, and other such options.
[0080] Figure 7 Components of an example computing device 700 that can be utilized according to various embodiments are shown. It should be understood that numerous such computing resources and numerous such components can be provided in various arrangements, such as in a local network or across the Internet or the “cloud,” to provide the computing resource capacity discussed elsewhere herein. The computing resource 700 (e.g., a desktop computer or network server) will have one or more processors 702, such as a central processing unit (CPU), a graphics processing unit (GPU), etc., which are electronically and / or communicatively coupled to various components using various buses, traces, and other such mechanisms.
[0081] Processor 702 may include memory register 706 and cache memory 704 for storing instructions, data, etc. In this example, chipset 714 (which in some embodiments may include a northbridge and a southbridge) may work with various system buses to connect processor 702 to components such as system memory 716 in the form of physical RAM or ROM, which may include code for the operating system and various other instructions and data used to operate the computing device. The computing device may also include or communicate with one or more storage devices 720 (such as hard disk drives, flash drives, optical storage devices, etc.) for persistent storage of data and instructions similar to or other than those stored in the processor and memory.
[0082] The processor 702 can also communicate with various other components via a chipset 714 and an interface bus (or graphics bus, etc.), which may include communication devices 724 (such as a cellular modem or network card), media components (graphics or audio cards) 726 (such as a graphics card and audio components), and peripheral interfaces 728 for connecting peripheral devices (such as printers, keyboards, etc.). It may also include at least one cooling fan 732 or other such temperature regulation or reduction components, which may be driven by the processor or triggered by various other sensors or components on or off the device. Various other or alternative components and configurations known in the art for computing devices may also be utilized.
[0083] In some embodiments, at least one processor 702 can obtain data from system memory 716 (such as a dynamic random access memory (DRAM) module) via a coherence configuration. It should be understood that various architectures can be utilized for such a computing device across a range of embodiments, including different choices, numbers, and parameters of buses and bridges. Data in memory can be managed and accessed by a memory controller (such as a DDR controller) via a coherence configuration. In at least some embodiments, data can be temporarily stored in cache memory 704. Computing resource 700 can also support multiple I / O devices using a set of I / O controllers connected via an I / O bus and also supported by a system clock 710. I / O controllers can be present to support corresponding types of I / O devices, such as Universal Serial Bus (USB) devices, data storage devices (e.g., flash or disk storage devices), network cards, high-speed peripheral component interconnect (PCIe) cards or peripheral interfaces 728, communication devices 724, graphics or audio cards 726, and direct memory access (DMA) cards, and other such options. In some implementations, components such as processors, controllers, and caches may be configured on a single card, board, or chip (e.g., a system-on-a-chip implementation), while in other implementations, at least some of the components may be located in different locations, etc.
[0084] The operating system (OS) running on the processor 702 helps manage the various devices that can be used to provide input to be processed. For example, this can include enabling interaction with various I / O devices using associated device drivers, which may be related to data storage devices, device communications, user interfaces, etc. The various I / O devices will typically be connected via various device ports and communicate with the processor and other device components through one or more buses. Specific types of buses may exist that provide communication according to specific protocols, such as Peripheral Component Interconnect (PCI) or Small Computer System Interface (SCSI) communication, and other such options. Communication can occur using registers associated with corresponding ports, including registers such as data input and data output registers. Memory-mapped I / O can also be used for communication, where a portion of the processor's address space is mapped to a specific device, and data is written directly to and from that portion of the address space.
[0085] For example, such a device can be used as a server in a server cluster or data warehouse. Server computers often need to perform tasks outside of the CPU and main memory (e.g., RAM) environment. For example, a server may need to communicate with external entities (e.g., other servers) or use an external processor (e.g., a general-purpose graphics processing unit (GPGPU)) to process data. In such cases, the CPU may interact with one or more I / O devices. In some cases, these I / O devices may be dedicated hardware designed to perform specific roles. For example, an Ethernet network interface controller (NIC) can be implemented as an application-specific integrated circuit (ASIC) that includes digital logic operable to send and receive data packets.
[0086] In illustrative embodiments, the host computing device is associated with various hardware components, software components, and corresponding configurations that facilitate the execution of I / O requests. One such component is an I / O adapter that inputs and outputs data along a communication channel. In one aspect, the I / O adapter device may communicate as a standard bridge component to facilitate access between various physical and emulated components and the communication channel. In another aspect, the I / O adapter device may include an embedded microprocessor to allow the I / O adapter device to execute computer-executable instructions relating to the implementation of management functions or the management of one or more such management functions, or to execute other computer-executable instructions relating to the implementation of the I / O adapter device. In some embodiments, the I / O adapter device may be implemented using multiple discrete hardware elements, such as multiple cards or other devices.
[0087] The management controller can be configured to electrically isolate itself from any other components in the host device besides the I / O adapter device. In some embodiments, the I / O adapter device is externally attached to the host device. In some embodiments, the I / O adapter device is internally integrated into the host device. An external communication port component may also communicate with the I / O adapter device to establish a communication channel between the host device and one or more network-based services or other network-attached or directly attached computing devices. Illustratively, the external communication port component may correspond to a network switch, sometimes referred to as a top-of-rack (TOR) switch. The I / O adapter device can utilize the external communication port component to maintain a communication channel between one or more services (such as health check services, financial services, etc.) and the host device.
[0088] The I / O adapter device can also communicate with a Basic Input / Output System (BIOS) component. The BIOS component may include non-transitory executable code (often referred to as firmware), which can be executed by one or more processors and is used to initialize and identify system devices such as video display cards, keyboards and mice, hard disk drives, optical disk drives, and other hardware on the host device. The BIOS component may also include or locate boot loader software that will be used to boot the host device. For example, in one embodiment, the BIOS component may include executable code that, when executed by the processor, causes the host device to attempt to locate the Preboot Execution Environment (PXE) boot software. Additionally, the BIOS component may include or utilize hardware latches electrically controlled by the I / O adapter device. Hardware latches can restrict access to one or more aspects of the BIOS component, thereby controlling the modification or configuration of the executable code maintained within the BIOS component. The BIOS component may connect to (or communicate with) multiple additional computing device resources such as processors, memory, etc.
[0089] In one embodiment, such computing device resource components can be physical computing device resources that communicate with other components via a communication channel. The communication channel can correspond to one or more communication buses, such as shared buses (e.g., front-side bus, memory bus), point-to-point buses (e.g., PCI or PCI Express buses), etc., on which components of the bare-metal host device communicate. Other types of communication channels, communication media, communication buses, or communication protocols (e.g., Ethernet communication protocols) can also be utilized. Additionally, in other embodiments, one or more of the computing device resource components can be virtualized hardware components emulated by the host device. In such embodiments, an I / O adapter device can implement a management process in which the host device is configured with physical or emulated hardware components based on various standards. The computing device resource components can communicate with the I / O adapter device via the communication channel. Furthermore, the communication channel can connect PCI Express devices to the CPU via the northbridge or host bridge, and other such options.
[0090] One or more controller components may communicate with the I / O adapter device via a communication channel for managing hard disk drives or other forms of memory. An example of a controller component could be a SATA hard disk drive controller. Similar to a BIOS component, a controller component may include hardware latches electrically controlled by the I / O adapter device or utilize their benefits. Hardware latches may restrict access to one or more aspects of the controller component. Illustratively, hardware latches may be controlled together or independently. For example, the I / O adapter device may selectively disable hardware latches for one or more components based on a trust level associated with a particular user. In another example, the I / O adapter device may selectively disable hardware latches for one or more components based on a trust level associated with the author or distributor of executable code to be executed by the I / O adapter device. In a further example, the I / O adapter device may selectively disable hardware latches for one or more components based on a trust level associated with the component itself. The host device may also include additional components that communicate with one or more of the illustrative components associated with the host device. Such components may include devices, such as one or more controllers combined with one or more peripheral devices (such as hard disks or other storage devices). Additionally, other components of the host device may include another set of peripheral devices, such as a graphics processing unit (“GPU”). Peripheral devices may also be associated with hardware latches for restricting access to one or more aspects of the components. As mentioned above, in one embodiment, the hardware latches may be controlled together or independently.
[0091] As discussed, different methods can be implemented in various environments according to the described implementation scheme. It should be understood that although a network-based or web-based environment is used in some examples presented herein for illustrative purposes, various implementation schemes can be implemented in different environments as appropriate. Such a system may include at least one electronic client device, which may include any suitable device operable to send and receive requests, messages, or information on a suitable network and to transmit information back to the user of the device. Examples of such client devices include personal computers, mobile phones, handheld messaging devices, laptop computers, set-top boxes, personal data assistants, e-book readers, etc.
[0092] The network may include any suitable network, including intranets, the Internet, cellular networks, local area networks, or any other such network or combinations thereof. The components used in such systems may depend at least in part on the type of network and / or environment chosen. Protocols and components used for communication via such a network are well known and will not be discussed in detail herein. Communication on the network may be achieved via wired or wireless connections and combinations thereof. In this example, the network includes the Internet because the environment includes a web server for receiving requests and providing content in response to those requests; however, for other networks, alternative devices serving similar purposes may be used, as will be apparent to those skilled in the art.
[0093] The illustrative environment includes at least one application server and a data storage area. It should be understood that there may be several application servers, layers, or other elements, processes, or components, which may be chained or otherwise configured and can interact to perform tasks such as retrieving data from the appropriate data storage area. As used herein, the term "data storage area" refers to any means or combination of means capable of storing, accessing, and retrieving data, which may include any combination and number of data servers, databases, data storage devices, and data storage media in any standard, distributed, or clustered environment. The application server may include any suitable hardware and software integrated with the data storage area and handling most of the application's data access and business logic as needed to execute various aspects of one or more applications on the client device. The application server provides access control services in cooperation with the data storage area and is capable of generating content, such as text, graphics, audio, and / or video, to be delivered to the user; in this example, the content may be provided to the user by a web server in the form of HTML, XML, or another suitable structured language. The handling of all requests and responses, as well as the delivery of content between the client device and the application server, may be handled by the web server. It should be understood that web servers and application servers are not necessary and are merely example components, as the structured code discussed in this article can be executed on any suitable device or host machine as discussed elsewhere in this article.
[0094] A data storage area may include several separate data tables, databases, or other data storage mechanisms and media for storing data related to specific aspects. For example, the data storage area shown includes mechanisms for storing content (e.g., production data) and user information, which can be used to provide content to the producer. The data storage area is also shown to include mechanisms for storing logs or session data. It should be understood that many other aspects may need to be stored in the data repository, such as page image information and access permission information, which may be stored in any of the mechanisms listed above or in additional mechanisms within the data repository, depending on the circumstances. The data repository can be operated through its associated logic to receive instructions from the application server and, in response to said instructions, retrieve, update, or otherwise process data. In one example, a user may submit a search request for a certain type of item. In this case, the data storage area can access user information to verify the user's identity and can access directory details to obtain information about that type of item. The information can then be returned to the user, such as in the form of a list of results on a webpage that the user can view via a browser on the user's device. Information about the specific item of interest can be viewed in a dedicated page or window of the browser.
[0095] Each server will typically include an operating system that provides executable program instructions for the general administration and operation of the server, and will typically include a computer-readable medium storing instructions that, when executed by the server's processor, allow the server to perform its intended functions. Suitable implementations of the server's operating system and general functions are known or commercially available, and these implementations are readily apparent to those skilled in the art, particularly in accordance with this disclosure.
[0096] In one embodiment, the environment is a distributed computing environment utilizing several computer systems and components interconnected via communication links using one or more computer networks or direct connections. However, those skilled in the art will understand that such a system can operate just as well in systems with fewer or more components than those shown. Therefore, the depiction of the system herein should be considered illustrative in nature and is not limited to the scope of this disclosure.
[0097] Various implementation schemes can be further implemented in a wide range of operating environments, in some cases of which may include one or more user computers or computing devices that can be used to operate any of a number of applications. User or client devices may include: any of several general-purpose personal computers, such as desktop or laptop computers running standard operating systems; and cellular, wireless, and handheld devices running mobile software and capable of supporting several networking and messaging protocols. Such systems may also include multiple workstations running various commercially available operating systems and any of other known applications for purposes such as research and development and database management. These devices may also include other electronic devices, such as virtual terminals, thin clients, gaming systems, and other devices capable of communicating via a network.
[0098] Most implementations utilize at least one network familiar to those skilled in the art to support communication using various commercially available protocols, such as TCP / IP, FTP, UPnP, NFS, and CIFS. The network can be, for example, a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), the Internet, an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network, and any combination thereof.
[0099] In implementations utilizing web servers, the web server can run any of a variety of server or middleware applications, including HTTP servers, FTP servers, CGI servers, data servers, Java servers, and business application servers. One or more servers may also be able to execute programs or scripts in response to requests from user devices, such as by executing one or more web applications that can be implemented in any programming language (such as Java®, C, C#, or C++) or any scripting language (such as Perl, Python, or TCL) and combinations thereof. The server also includes database servers, including but not limited to servers available from Oracle®, Microsoft®, Sybase®, and IBM®, as well as open-source servers such as MySQL, Postgres, SQLite, MongoDB, and any other server capable of storing, retrieving, and accessing structured or unstructured data. Database servers may include table-based servers, document-based servers, unstructured servers, relational servers, non-relational servers, or combinations of these and / or other database servers.
[0100] The environment may include a variety of data storage areas as discussed above, as well as other storage and storage media. These may reside in various locations, such as on storage media local to one or more computers (and / or residing in one or more of one or more computers), or remotely to any or all of the computers spread across a network. In a particular set of implementations, information may reside in a storage area network (SAN) familiar to those skilled in the art. Similarly, any necessary files for performing functions belonging to a computer, server, or other network device may be stored locally and / or remotely, as appropriate. Where the system includes computerized devices, each such device may include hardware elements that can be electrically coupled via a bus, including, for example, at least one central processing unit (CPU), at least one input device (e.g., mouse, keyboard, controller, touch-sensitive display element, or keypad), and at least one output device (e.g., display device, printer, or speaker). Such a system may also include one or more storage devices, such as hard disk drives, tape drives, optical storage devices, and solid-state storage devices such as random access memory (RAM) or read-only memory (ROM), as well as removable media devices, memory cards, flash memory cards, etc.
[0101] Such devices may also include computer-readable storage medium readers, communication devices (e.g., modems, (wireless or wired) network interface cards, infrared communication devices), and working memory as described above. A computer-readable storage medium reader may be connected to or configured to receive a computer-readable storage medium, which represents a remote, local, fixed, and / or removable storage device and storage medium for temporarily and / or more permanently accommodating, storing, transmitting, and retrieving computer-readable information. Systems and various devices will also typically include numerous software applications, modules, services, or other elements, including operating systems and applications such as client applications or web browsers, residing within at least one working memory device. It should be understood that alternative embodiments may have numerous variations different from those described above. For example, custom hardware may also be used, and / or specific elements may be implemented in hardware, software (including portable software such as small applications), or both. Furthermore, connectivity with other computing devices, such as network input / output devices, may be employed.
[0102] Storage media and other non-transient computer-readable media used to contain code, or portions thereof, may include any suitable media known or used in the art, such as, but not limited to, volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing information such as computer-readable instructions, data structures, program modules or other data: RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage devices, magnetic cartridges, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by system devices. Based on this disclosure and the teachings provided herein, those skilled in the art will understand other ways and / or methods for implementing various embodiments.
[0103] The description and drawings should therefore be considered illustrative rather than restrictive. However, it will be apparent that various modifications and changes can be made to the invention without departing from the broader spirit and scope of the invention as set forth in the claims.
Claims
1. A system (100; 200; 500 to 700), characterized in that: At least one processor (612; 702); and Memory (610, 612, 712), the memory storing instructions, the instructions causing the system to: when executed by the at least one processor Receive (302) a computer coding policy executed in a control plane (106) associated with a cloud environment (102), the computer coding policy being associated with data governance using one or more data assets (116) of the cloud environment in a data plane (104); The one or more data assets are automatically associated (310) with the computer coding strategy using a set of predetermined rules associated with the computer coding strategy and with annotations associated with the one or more data assets. One or more dynamic changes are made to the annotation, in part based on real-time changes to the computer coding strategy (404); According to the computer coding strategy, monitor (406) the content of the one or more data assets; as well as In response to a violation associated with the computer coding policy, (412) a remedial action associated with the one or more data assets is performed.
2. The system of claim 1, wherein the computer coding policy includes policy parameters (128) associated with one or more of retention rules, deletion rules, data filtering rules, data disclosure rules, data sovereignty rules, or territorial rules.
3. The system of claim 1, wherein the memory includes the instructions, which, when executed by the at least one processor, cause the system to: An interface (124) is provided to enable the definition of a strategy type and matching rules to be part of the computer coding strategy, wherein the matching rules are associated with the directory and the annotation, so that the system can perform automatic association of the one or more data assets to the computer coding strategy association strategy; and The semantic subsystem is used to determine the one or more data assets associated with the policy type and the matching rule.
4. The system of claim 1, wherein the memory includes the instructions, which, when executed by the at least one processor, cause the system to: Provide test parameters for the computer coding strategy, wherein the test parameters define one or more compliance thresholds for the one or more data assets; and Multiple remediation actions are provided for the one or more data assets, wherein the remediation actions performed on the one or more data assets are provided from the multiple remediation actions based on a violation of at least one of the one or more compliance thresholds.
5. The system of claim 1, wherein the memory includes the instructions, which, when executed by the at least one processor, cause the system to: The instructions are generated in part based on the computer coding strategy executed in the control plane; and The deletion or addition is performed in the retention table in part based on the instructions to enforce the remediation action associated with the one or more data assets.
6. The system of claim 1, wherein the remediation action is either changing access control to the data storage of the one or more data assets or performing a soft delete or hard delete to remove non-compliant data from the one or more data assets.
7. A computer-implemented method (300, 400), characterized in that: Receive (302) computer coding policy for data governance using one or more data assets within the cloud environment; The one or more data assets are automatically associated (310) with the computer coding strategy, at least in part based on one or more predefined functions of the one or more data assets; Implement dynamic changes (404) to add or remove the identified data assets from the one or more data assets; as well as In response to a violation associated with the computer coding policy, (412) a remedial action associated with the one or more data assets is performed.
8. The computer-implemented method as described in claim 7, Its further features are: An interface is provided to enable the definition of policy types and matching rules to be part of the computer coding policy, wherein the matching rules are associated with the directory and the annotation, so that the system can perform automatic association of the one or more data assets to the computer coding policy association policy; as well as The semantic subsystem is used to determine the one or more data assets associated with the policy type and the matching rule.
9. The computer-implemented method as described in claim 7, further characterized in that: Provide test parameters for the computer coding strategy, wherein the test parameters define one or more compliance thresholds for the one or more data assets; and Multiple remediation actions are provided for the one or more data assets, wherein the remediation actions performed on the one or more data assets are provided from the multiple remediation actions based on a violation of at least one of the one or more compliance thresholds.
10. The computer-implemented method as described in claim 7, further characterized in that: Instructions are generated in part based on the computer coding strategy executed in the control plane of the cloud environment; as well as The deletion or addition is performed in the retention table in part based on the instructions to enforce the remediation action associated with the one or more data assets.
11. The computer-implemented method as described in claim 7, Its further features are: Enable a preview action associated with the computer coding strategy using the interface of the control plane, wherein the computer coding strategy is applied to the representation of the one or more data assets; Provides results associated with remedial actions or violations for the representation of the one or more data assets; as well as The computer coding policy can be published to take action on the one or more data assets.
12. A non-transitory computer storage medium (610, 612; 712) storing instructions configured to instruct at least one computing device (100; 200; 500 to 700): Receive (302) computer coding policy for data governance using one or more data assets within the cloud environment; At least in part based on the predefined functions of the one or more data assets, the one or more data assets are automatically associated (310) with the computer coding strategy; (404) Enable dynamic changes to add or remove identified data assets from one or more data assets; as well as In response to a violation associated with the computer coding policy, (412) a remedial action associated with the one or more data assets is performed.
13. The non-transitory computer storage medium of claim 12, wherein the instructions are configured to instruct at least one computing device to further: Provides an interface that allows the definition of policy types and matching rules to be part of the computer coding policy, wherein the matching rules are associated with directories and annotations, so that the system can perform automatic association of the one or more data assets to the computer coding policy association policy; and The semantic subsystem is used to determine the one or more data assets associated with the policy type and the matching rule.
14. The non-transitory computer storage medium of claim 12, wherein the instructions are configured to instruct at least one computing device to further: Provide test parameters for the computer coding strategy, wherein the test parameters define one or more compliance thresholds for the one or more data assets; and Multiple remediation actions are provided for the one or more data assets, wherein the remediation actions performed on the one or more data assets are provided from the multiple remediation actions based on a violation of at least one of the one or more compliance thresholds.
15. The non-transitory computer storage medium of claim 12, wherein the instructions are configured to instruct at least one computing device to further: The instructions are generated in part based on the computer coding strategy executed in the control plane; and The deletion or addition is performed in the retention table in part based on the instructions to enforce the remediation action associated with the one or more data assets.