Self-adaptive data management method and system
Through adaptive data governance methods and systems, the governance policy chain is selected and recorded in real time during the data migration process, which solves the problems of inefficient data governance and leakage across cloud platforms and realizes efficient and secure data migration and management.
Patent Information
- Application Number
- CN202510989297.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-10
AI Technical Summary
Existing data governance methods are inefficient in cross-cloud environments, making unified management difficult to achieve. They also pose data leakage and compliance risks, especially during data migration, when improper governance policy configuration can easily lead to data leakage.
By obtaining information about the data to be migrated and the governance requirements, the governance policy chain is adaptively selected, data governance is performed while migrating the data to be migrated, and the governance records are updated on the distributed ledger to ensure the consistency and traceability of the governance policy during the data migration process.
It improves governance efficiency during data migration, reduces the risk of data leakage, and achieves unified management and data compliance traceability across cloud platforms.
Smart Images

Figure CN120763141A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data governance, and in particular, to an adaptive data governance method and system. Background Art
[0002] As enterprises deepen their digital transformation, multi-cloud architectures have become a mainstream choice for supporting business development. To achieve optimal resource allocation and greater business agility, enterprises are generally adopting platforms from multiple cloud service providers. This trend has led to an unprecedented degree of dispersion in enterprise data assets. Data is no longer centralized in a single data center, but is now widely distributed across various cloud platforms. At the same time, data types have become extremely diverse. More importantly, this massive amount of data contains a significant amount of personal privacy information and commercial secrets, with a disproportionately high proportion of sensitive data. To fully tap into the value of data, drive business innovation, and streamline operations, enterprises need to frequently migrate and convert data across clouds. For example, they need to migrate production data to analytical environments for in-depth analysis, or migrate data from one cloud platform to another to optimize costs or improve performance.
[0003] However, most data governance tools and methods currently available are designed for traditional single-cloud environments or on-premises data centers. Their limitations are becoming increasingly prominent when navigating complex and dynamic multi-cloud environments. Traditional data governance solutions face numerous challenges, particularly in data migration and conversion scenarios. First, a lack of unified management capabilities across cloud environments makes it difficult for data governance policies to work together across different cloud platforms, resulting in inefficient data governance. Second, with the explosive growth of data assets and the diversification of data types in multi-cloud environments, traditional manual configuration methods are no longer sufficient. They struggle to respond to data changes and security threats in real time, easily creating data governance bottlenecks. Furthermore, the lack of a unified view of data assets makes it difficult for enterprises to fully understand the distribution and status of data assets across clouds, leading to inefficient data management and utilization. More seriously, during the migration and conversion of sensitive data across clouds, improperly configured or inadequately enforced governance policies can easily lead to data leaks and data compliance violations, posing significant security and compliance risks to enterprises.
[0004] Therefore, in order to solve the technical problems of existing data governance methods such as low data governance efficiency and high risk of data leakage when working across cloud platforms, an adaptive data governance method and system is urgently needed. Summary of the Invention
[0005] The purpose of this application is to provide an adaptive data governance method and system, which selects a governance policy chain based on the data information and governance requirement information of the data to be migrated, performs data governance while migrating the data to be migrated, obtains data governance records, and updates the data governance records to a distributed ledger, thereby solving the problems of low data governance efficiency and high risk of data leakage in existing data governance methods when working across cloud platforms. During data migration, the system adaptively selects governance policies and performs governance simultaneously, and puts governance records on the chain, thereby improving the governance efficiency during data migration.
[0006] In a first aspect, the present application provides an adaptive data governance method, comprising: Responding to user-initiated cross-cloud data migration requests, obtaining data information and governance requirements for the data to be migrated; According to the data information and the governance requirement information, a corresponding data governance policy is selected from a preset database to obtain a governance policy chain; Based on the governance policy chain and the data information, data governance is performed on the data to be migrated while migrating the data to obtain a data governance record; The data governance records are updated on the distributed ledger so that staff can trace the migration path and governance operation records of the data to be migrated through the distributed ledger.
[0007] The adaptive data governance method provided by this application can realize data governance during data migration. By selecting a governance policy chain based on the data information of the data to be migrated and the governance requirement information, data governance is performed while the data to be migrated is being migrated, a data governance record is obtained, and the data governance record is updated on the distributed ledger. This solves the problems of low data governance efficiency and high risk of data leakage in existing data governance methods when working across cloud platforms. During data migration, the governance strategy is adaptively selected and governance is performed synchronously, and the governance record is uploaded to the chain, thereby improving the governance efficiency during data migration.
[0008] Optionally, the data information includes the original location of the data to be migrated, the target migration location, the data type and the data size.
[0009] Optionally, according to the data information and the governance requirement information, a corresponding data governance policy is selected from a preset database to obtain a governance policy chain, including: Extracting a candidate strategy set from the preset database according to the data information and the governance requirement information; Using a rule-based conflict detection method, identifying and removing conflicting policies in the candidate policy set to obtain a streamlined policy set; Based on preset policy arrangement rules, the data governance policies in the streamlined policy set are sorted to obtain a governance policy chain.
[0010] The adaptive data governance method provided in this application can realize data governance during data migration. From preliminary screening, conflict elimination to final sorting, it systematically constructs a governance policy chain suitable for specific data and needs, overcomes the possible defects of simple policy selection, and improves the accuracy, efficiency and reliability of data governance.
[0011] Optionally, according to the data information and the governance requirement information, a candidate policy set is extracted from the preset database, including: Extracting an initial data governance strategy corresponding to the data type from the preset database according to the data type in the data information; Extracting corresponding keyword tags from the governance demand information; The keyword tags are matched with the initial data governance policies to obtain a candidate policy set.
[0012] Optionally, a rule-based conflict detection method is used to identify and remove conflicting policies in the candidate policy set to obtain a streamlined policy set, including: Extracting a conflicting strategy combination from the candidate strategy set based on preset conflicting strategy rules; Determine the low-priority strategy in the conflicting strategy combination according to the preset strategy priority; The low-priority policies are eliminated from the candidate policy set to obtain a streamlined policy set.
[0013] The adaptive data governance method provided in this application can realize data governance during data migration by comparing the priorities of each policy in the identified conflicting policy combination, determining the policy with the lowest priority, and eliminating low-priority policies from the candidate policy set to obtain a streamlined policy set. By removing the policies with lower priority in the conflicting policy combination, conflicts between policies are resolved, ensuring consistency within the streamlined policy set, and improving governance efficiency during data migration.
[0014] Optionally, based on a preset policy arrangement rule, the data governance policies in the simplified policy set are sorted to obtain a governance policy chain, including: Determine corresponding migration scenario information based on the data information; the migration scenario information includes network bandwidth, cloud platform computing resources, and data storage format; According to the migration scenario information and the preset policy arrangement rules, a topological sorting algorithm is used to sort the data governance policies in the streamlined policy set to obtain a governance policy chain.
[0015] Optionally, according to the migration scenario information and the preset policy arrangement rules, a topological sorting algorithm is used to sort the data governance policies in the streamlined policy set to obtain a governance policy chain, including: Based on the migration scenario information, a directed acyclic graph corresponding to the streamlined policy set is constructed, with the data governance goal as the end point, each data governance policy as a node, and the dependency relationship between each data governance policy as an edge; Using a weighted topological sorting algorithm, according to the weight of each edge in the directed acyclic graph, the execution priority of each data governance policy in the streamlined policy set is calculated to obtain a preliminary governance policy chain; According to the preset policy arrangement rules, the preliminary governance policy chain is optimized to obtain a governance policy chain.
[0016] Optionally, the preset policy arrangement rule is to first execute data desensitization policies, then execute data format conversion policies, and finally execute data quality verification policies.
[0017] Optionally, based on the governance policy chain and the data information, data governance is performed on the data to be migrated while migrating the data to obtain a data governance record, including: Extracting the data to be migrated from the original location of the data in the data information; Based on the governance policy chain, data governance is performed on the data to be migrated to obtain the data to be migrated after data governance; After migrating the data to be migrated after data governance to the target migration location in the data information, a data governance record is recorded.
[0018] In a second aspect, the present application provides an adaptive data governance system, comprising: An acquisition module, configured to respond to a cross-cloud data migration request initiated by a user and obtain data information and governance requirement information of the data to be migrated; A selection module is used to select a corresponding data governance policy from a preset database according to the data information and the governance requirement information to obtain a governance policy chain; A governance module, configured to perform data governance while migrating the data to be migrated based on the governance policy chain and the data information, and obtain a data governance record; An update module is used to update the data governance record to the distributed ledger so that staff can trace the migration path and governance operation record of the data to be migrated through the distributed ledger.
[0019] This adaptive data governance system performs data governance while migrating the data by selecting a governance policy chain based on the data information and governance requirement information of the data to be migrated, obtains data governance records, and updates the data governance records to the distributed ledger. It solves the problems of low data governance efficiency and high risk of data leakage in existing data governance methods when working across cloud platforms. It adaptively selects governance strategies during data migration and performs governance simultaneously, and at the same time uploads governance records to the chain, thereby improving the governance efficiency during data migration.
[0020] Beneficial effects: The adaptive data governance method and system provided by this application performs data governance on the data to be migrated while migrating it, obtains data governance records, and updates the data governance records to the distributed ledger through the governance policy chain selected based on the data information and governance demand information of the data to be migrated. This solves the problems of low data governance efficiency and high risk of data leakage in existing data governance methods when working across cloud platforms. It adaptively selects governance strategies and performs governance simultaneously during data migration, and puts governance records on the chain, thereby improving governance efficiency during data migration. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 Flowchart of the adaptive data governance method provided in an embodiment of the present application.
[0022] Figure 2 A schematic diagram of the structure of the adaptive data governance system provided in an embodiment of the present application.
[0023] Explanation of reference numerals: 1. Acquisition module; 2. Selection module; 3. Management module; 4. Update module; 301. Processor; 302. Memory; 303. Communication bus. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.
[0025] It should be noted that like reference numerals and characters refer to like elements throughout the following description and the claims, therefore, once an element is defined in one drawing, it is not necessary to further define and explain it in the subsequent drawings. Also, in the description of the present application, the terms "first", "second", and the like are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.
[0026] Please refer to Figure 1 , Figure 1 The adaptive data governance method is used for data governance during data migration in some embodiments of the present application, comprising the steps of: Step S101, in response to a cross-cloud data migration request initiated by a user, obtaining data information and governance requirement information of data to be migrated; Step S102, according to the data information and the governance requirement information, selecting a corresponding data governance strategy from a preset database to obtain a governance strategy chain; Step S103, based on the governance strategy chain and the data information, performing data governance while performing data migration on the data to be migrated to obtain data governance records; Step S104, updating the data governance records to a distributed ledger, so that the staff can trace the migration path and governance operation records of the data to be migrated through the distributed ledger.
[0027] The adaptive data governance method, by selecting the governance strategy chain based on the data information and the governance requirement information of the data to be migrated, performing data governance while performing data migration on the data to be migrated to obtain data governance records, and updating the data governance records to a distributed ledger, solves the problems of low data governance efficiency and easy data leakage of existing data governance methods when working across cloud platforms. The adaptive selection of governance strategy and synchronous governance during data migration, and the on-chain governance records improve the governance efficiency during data migration.
[0028] Specifically, in step S101, in response to a cross-cloud data migration request initiated by a user, data information and governance requirement information of the data to be migrated are obtained, wherein the data information includes the original location of the data to be migrated, the target migration location, the data type, and the data size. The original location of the data indicates the specific location where the data to be migrated is currently stored, such as a specific cloud storage bucket, database instance, or file path. The target migration location indicates the specific location where the data to be migrated will be stored after the migration is completed, such as another cloud storage bucket, database instance, or file path. The data type indicates the structure and content attributes of the data to be migrated, such as structured data, unstructured data, semi-structured data, or a data type containing sensitive information. The data size indicates the scale of the data to be migrated, such as the capacity in bytes, gigabytes, or terabytes. These data information are collected by the acquisition module as input for the subsequent selection of data governance strategies and the execution of data migration governance operations.
[0029] Governance requirements are the specific requirements for the governance of migrated data that users submit when initiating a data migration request. This information reflects the type of governance operations users wish to perform on their data and the governance objectives they aim to achieve. Governance requirements can include various types of information, such as data masking, which users may request to mask sensitive information (such as names, ID numbers, and contact information) in their data to protect personal privacy or meet regulatory compliance requirements.
[0030] Specifically, in step S102, according to the data information and the governance requirement information, a corresponding data governance policy is selected from a preset database to obtain a governance policy chain, including: Based on data information and governance requirements, a candidate strategy set is extracted from the preset database; A rule-based conflict detection method is used to identify and remove conflicting policies in the candidate policy set to obtain a streamlined policy set. Based on the preset policy orchestration rules, the data governance policies in the streamlined policy set are sorted to obtain a governance policy chain.
[0031] Specifically, in step S102, based on the data information and the governance requirement information, a candidate policy set is extracted from a preset database, including: According to the data type in the data information, the initial data governance strategy of the corresponding data type is extracted from the preset database; Extract corresponding keyword tags from governance demand information; Match the keyword tags with the initial data governance policy to obtain a set of candidate policies.
[0032] In step S102, the data type of the data to be migrated is used as the first-level filtering criterion. Based on the data type (e.g., relational database tables, unstructured documents, streaming data, etc.), an initial set of governance policies applicable to that type of data is quickly located and extracted from a pre-set policy database. This step significantly narrows the search scope and improves efficiency. Next, the governance requirements information entered by the user is analyzed, and key business or technical terms are extracted from it, which are used as keyword tags. These tags reflect the specific governance operations (e.g., data cleansing, desensitization, encryption, format conversion) or goals (e.g., meeting compliance requirements, improving data quality) that the user desires to perform on the data. Finally, the extracted keyword tags are matched against the initial set of data governance policies obtained in the first step. This matching process can be based on the policy metadata, descriptive text, or a pre-set tagging system. Only initial policies that successfully match the keyword tags (i.e., their descriptions, functions, or tags match or have similar meanings to the keyword tags) are selected for inclusion in the final set of candidate policies. By combining data type and governance requirement keywords for joint screening, we can more accurately identify the governance policies most relevant to the current data migration task, avoiding the extraction of a large number of irrelevant or inapplicable policies, thereby providing more streamlined and effective input for subsequent policy conflict detection and sorting, and improving the accuracy and efficiency of the entire policy selection process.
[0033] For example, the data type submitted by a user for migration is "relational database table," and the governance requirement information submitted includes "customer names and addresses need to be desensitized and converted to Parquet format for big data analysis." First, based on the data type "relational database table," all initial governance policies applicable to relational databases are extracted from the pre-set policy database, such as "SQL data cleansing policy," "relational data desensitization policy," "relational data encryption policy," and "relational data format conversion policy." Keyword tags are also extracted from the governance requirement information, such as "desensitization," "name," "address," "format conversion," "Parquet," and "big data analysis." The system then matches these keyword tags with the extracted initial policies. For example, the "relational data desensitization policy" might contain metadata or descriptions related to "desensitization," "name," and "address," and is therefore matched and selected. The "relational data format conversion policy" might contain metadata related to "format conversion" and "Parquet," and is therefore also matched and selected. The "SQL data cleansing policy" and "relational data encryption policy" might not match the extracted keyword tags and are therefore not selected. Ultimately, the candidate strategy set may include "relational data desensitization strategy" and "relational data format conversion strategy".
[0034] Specifically, in step S102, a rule-based conflict detection method is used to identify and remove conflicting policies in the candidate policy set, thereby obtaining a streamlined policy set including: Based on the preset conflict strategy rules, a conflict strategy combination is extracted from the candidate strategy set; According to the preset policy priority, determine the low priority policy in the conflicting policy combination; Low-priority policies are eliminated from the candidate policy set to obtain a streamlined policy set.
[0035] In step S102, to address the issue of a candidate policy set potentially containing conflicting policies, a conflict detection and resolution process is performed after the candidate policy set is extracted. First, the policies in the candidate policy set are scanned and compared according to preset conflict policy rules to identify all conflicting policy combinations. For example, if the candidate set contains a policy requiring the complete deletion of a sensitive field and a policy requiring format conversion of the field, and the preset rules define these two operations as conflicting on the same field, then these two policies are identified as a conflicting policy combination. Next, each policy in the conflicting policy combination is evaluated according to the preset policy priority. For example, if the deletion policy has a higher priority than the format conversion policy, then the format conversion policy is determined to be a low-priority policy. Finally, the determined low-priority policies are removed from the candidate policy set, resulting in a streamlined policy set that contains no internal conflicts. This streamlined policy set provides a reliable foundation for subsequent policy orchestration, avoiding data governance errors or inefficiencies caused by policy conflicts. The preset conflict policy rule is to only execute one policy of the same type for the same field or for sentences or paragraphs corresponding to the same field.
[0036] For example, assume that the candidate policy set extracted from the preset database includes Policy A (data encryption), Policy B (data desensitization), and Policy C (data compression). The preset conflicting policy rules define Policy A and Policy B as a conflicting policy combination. The preset policy priorities set Policy A to a priority of 10, Policy B to a priority of 8, and Policy C to a priority of 5. Based on the preset conflicting policy rules, it is identified that Policy A and Policy B constitute a conflicting policy combination. Based on the preset policy priorities, Policy A (priority 10) is compared with Policy B (priority 8), and Policy B is determined to have a lower priority. Policy B is removed from the candidate policy set {Policy A, Policy B, Policy C}, resulting in the streamlined policy set {Policy A, Policy C}.
[0037] Specifically, in step S102, based on the preset policy arrangement rules, the data governance policies in the streamlined policy set are sorted to obtain a governance policy chain, including: Determine the corresponding migration scenario information based on the data information; migration scenario information includes network bandwidth, cloud platform computing resources, and data storage format; According to the migration scenario information and the preset policy orchestration rules, a topological sorting algorithm is used to sort the data governance policies in the streamlined policy set to obtain a governance policy chain.
[0038] In step S102, the corresponding migration scenario information is determined based on the data information of the data to be migrated, such as the data size and data type, in combination with the current data migration environment. This migration scenario information may include factors that affect the efficiency and feasibility of policy execution, such as the current network bandwidth conditions, the computing resource availability of the target cloud platform, and the storage format of the data to be migrated. This information reflects the actual environmental constraints during policy execution.
[0039] Specifically, in step S102, based on the migration scenario information and the preset policy arrangement rules, a topological sorting algorithm is used to sort the data governance policies in the streamlined policy set to obtain a governance policy chain, including: Based on the migration scenario information, a directed acyclic graph corresponding to the streamlined policy set is constructed, with the data governance goal as the end point, each data governance policy as the node, and the dependencies between data governance policies as the edges. A weighted topological sorting algorithm is used to calculate the execution priority of each data governance policy in the streamlined policy set based on the weight of each edge in the directed acyclic graph, thus obtaining a preliminary governance policy chain. According to preset policy arrangement rules, the preliminary governance policy chain is optimized to obtain a governance policy chain.
[0040] In step S102, a directed acyclic graph is constructed according to current data information (e.g., data type, data volume) and migration scenario information (e.g., source / target cloud platform type, network bandwidth, available computing resources). The nodes in the graph correspond to each strategy in the set of refinement strategies, and the terminal node represents the data governance completion state. The dependency relationship between strategies (e.g., data encryption must be performed before data compression) is represented as a directed edge. Further, according to the migration scenario information, a weight is assigned to each edge in the graph. For example, in a network bandwidth limited scenario, the edge weight corresponding to the subsequent strategy (such as data transmission) of the data compression strategy may be adjusted to reflect the importance or execution efficiency of the compression strategy. A weighted topological sorting algorithm is used to process the directed acyclic graph to calculate the execution priority of each strategy and generate a preliminary strategy execution order chain. The preliminary strategy chain satisfies the dependency relationship between strategies and takes into account the relative execution cost or efficiency of the strategy under the current scenario. Finally, according to the preset strategy arrangement rule (e.g., the data desensitization strategy is always executed before the data quality verification strategy), the preliminary strategy chain is checked and adjusted to obtain the final strategy chain for data migration and governance. The final strategy chain takes into account the strategy dependency, scenario characteristics and business rules, improving the efficiency and effectiveness of data governance. The preset strategy arrangement rule is a strategy arrangement rule that prioritizes the execution of data desensitization strategies, followed by data format conversion strategies, and finally data quality verification strategies.
[0041] For example, suppose the data to be migrated requires the application of four governance policies: data desensitization (S1), data format conversion (S2), data quality verification (S3), and data compression (S4). The known dependencies are that S1 must be executed before S2, S2 must be executed before S4, and S3 has no direct preceding dependencies. The migration scenario involves low network bandwidth and sufficient computing resources on the target cloud platform. First, construct a directed acyclic graph: nodes S1, S2, S3, S4, and endpoint E. The edges are S1→S2, S2→S4, S3→E, and S4→E. Next, assign weights based on the migration scenario. Because low network bandwidth significantly impacts subsequent transmission efficiency, S4 (data compression) can be given a higher weight, indicating that the steps before S4 are prioritized. The weights of other edges are determined based on factors such as policy execution time. Using a weighted topological sorting algorithm, a preliminary policy chain may be obtained: S1→S2→S4→S3. This order satisfies the dependencies and takes into account the importance of the compression policy. Finally, optimization is performed based on the pre-set policy orchestration rules. For example, the pre-set rules dictate that data desensitization policies (S1) must be executed first, and data quality verification policies (S3) must be executed last. The initial policy chain, S1 → S2 → S4 → S3, complies with the S1 priority and S3 last rule. Therefore, the final governance policy chain is determined to be S1 → S2 → S4 → S3. This policy chain not only satisfies dependencies and rules, but also takes into account network bandwidth limitations through weighted sorting, optimizing overall execution efficiency.
[0042] Specifically, in step S103, based on the governance policy chain and data information, data governance is performed on the data to be migrated while migrating the data, and a data governance record is obtained, including: Extract the data to be migrated from the original location of the data in the data information; Based on the governance policy chain, data governance is performed on the data to be migrated to obtain the data to be migrated after data governance; After migrating the data to be migrated after data governance to the target migration location in the data information, a data governance record is recorded.
[0043] In step S102, the original location of the data to be migrated is determined based on the data information, and the data to be migrated is obtained from this location. Using a predetermined governance policy chain, a series of governance operations are performed on the data to be migrated, such as desensitizing sensitive information, standardizing data formats, cleaning data values, etc., to ensure that the data is effectively processed before or during migration, and to generate data that meets governance requirements. The governed data is transferred to the target migration location specified by the data information. After the data is successfully migrated, the system records the detailed information of this operation, including the source and destination of the data, the applied governance policy, the execution results, etc., to form a data governance record. By embedding the data governance steps into the data migration process, it is ensured that only governed data will be migrated to the target location, thereby reducing the security and compliance risks faced by the data during the migration process and improving the reliability of data migration.
[0044] Specifically, in step S104, the generated data governance record is written to a distributed ledger. The tamper-proof nature of the distributed ledger ensures the authenticity and integrity of the data governance record. By querying the distributed ledger, staff or auditors can easily trace the complete migration path of the migrated data and all governance operations performed on it during the migration process. This provides end-to-end data processing transparency, greatly enhancing the credibility and compliance of data processing, and effectively reducing the risk of data leakage and illegal operations.
[0045] As can be seen from the above, the adaptive data governance method obtains data information and governance requirement information of the data to be migrated by responding to the cross-cloud data migration request initiated by the user, selects the corresponding data governance policy from the preset database according to the data information and governance requirement information, obtains the governance policy chain, and based on the governance policy chain and data information, performs data governance on the data to be migrated while migrating the data to obtain data governance records, and updates the data governance records to the distributed ledger so that the staff can trace the migration path and governance operation records of the data to be migrated through the distributed ledger; thus, by selecting the governance policy chain based on the data information and governance requirement information of the data to be migrated, performs data governance on the data to be migrated while migrating the data to obtain data governance records, and updates the data governance records to the distributed ledger, solving the problems of low data governance efficiency and easy data leakage in the existing data governance methods when working across cloud platforms, adaptively selects governance strategies and performs governance simultaneously during data migration, and uploads the governance records to the chain, thereby improving the governance efficiency during data migration.
[0046] refer to Figure 2 , this application provides an adaptive data governance system for governing data during data migration, including: Acquisition module 1, used to respond to a cross-cloud data migration request initiated by a user and obtain data information and governance requirement information of the data to be migrated; Selection module 2 is used to select the corresponding data governance policy from the preset database based on the data information and governance requirement information to obtain the governance policy chain; Governance module 3 is used to perform data governance while migrating the data to be migrated based on the governance policy chain and data information, and obtain data governance records; Update module 4 is used to update the data governance records to the distributed ledger so that staff can trace the migration path and governance operation records of the data to be migrated through the distributed ledger.
[0047] This adaptive data governance system performs data governance while migrating the data by selecting a governance policy chain based on the data information and governance requirement information of the data to be migrated, obtains data governance records, and updates the data governance records to the distributed ledger. It solves the problems of low data governance efficiency and high risk of data leakage in existing data governance methods when working across cloud platforms. It adaptively selects governance strategies during data migration and performs governance simultaneously, and at the same time uploads governance records to the chain, thereby improving the governance efficiency during data migration.
[0048] Specifically, when the acquisition module 1 is executed, it responds to the cross-cloud data migration request initiated by the user and obtains the data information and governance requirement information of the data to be migrated, wherein the data information includes the original location of the data to be migrated, the target migration location, the data type and the data size. The original location of the data indicates the specific location where the data to be migrated is currently stored, such as a specific cloud storage bucket, database instance or file path. The target migration location indicates the specific location where the data to be migrated will be stored after the migration is completed, such as another cloud storage bucket, database instance or file path. The data type indicates the structure and content attributes of the data to be migrated, such as structured data, unstructured data, semi-structured data, or a data type containing sensitive information. The data size indicates the scale of the data to be migrated, such as the capacity in bytes, gigabytes or terabytes. These data information are collected by the acquisition module as input for the subsequent selection of data governance strategies and the execution of data migration governance operations.
[0049] Governance requirements are the specific requirements for the governance of migrated data that users submit when initiating a data migration request. This information reflects the type of governance operations users wish to perform on their data and the governance objectives they aim to achieve. Governance requirements can include various types of information, such as data masking, which users may request to mask sensitive information (such as names, ID numbers, and contact information) in their data to protect personal privacy or meet regulatory compliance requirements.
[0050] Specifically, when the selection module 2 selects the corresponding data governance policy from the preset database based on the data information and governance requirement information and obtains the governance policy chain, it executes: Based on data information and governance requirements, a candidate strategy set is extracted from the preset database; A rule-based conflict detection method is used to identify and remove conflicting policies in the candidate policy set to obtain a streamlined policy set. Based on the preset policy orchestration rules, the data governance policies in the streamlined policy set are sorted to obtain a governance policy chain.
[0051] Specifically, when the selection module 2 extracts a candidate strategy set from a preset database based on data information and governance requirement information, it executes: According to the data type in the data information, the initial data governance strategy of the corresponding data type is extracted from the preset database; Extract corresponding keyword tags from governance demand information; Match the keyword tags with the initial data governance policy to obtain a set of candidate policies.
[0052] During execution, Selection Module 2 uses the data type of the data to be migrated as the first-level filtering criteria. Based on the data type (e.g., relational database tables, unstructured documents, streaming data, etc.), it quickly locates and extracts an initial set of governance policies applicable to that data type from the pre-set policy database. This step significantly narrows the search scope and improves efficiency. Next, it analyzes the governance requirements information entered by the user, extracting key business or technical terms from it and using them as keyword tags. These tags reflect the specific governance operations (e.g., data cleansing, desensitization, encryption, format conversion) or goals (e.g., meeting compliance requirements, improving data quality) that the user wishes to perform on the data. Finally, the extracted keyword tags are matched against the initial set of data governance policies obtained in the first step. This matching process can be based on the policy metadata, descriptive text, or a pre-set tagging system. Only initial policies that successfully match the keyword tags (i.e., their descriptions, functions, or tags match or have similar meanings to the keyword tags) are selected for inclusion in the final set of candidate policies. By combining data type and governance requirement keywords for joint screening, we can more accurately identify the governance policies most relevant to the current data migration task, avoiding the extraction of a large number of irrelevant or inapplicable policies, thereby providing more streamlined and effective input for subsequent policy conflict detection and sorting, and improving the accuracy and efficiency of the entire policy selection process.
[0053] For example, the data type submitted by a user for migration is "relational database table," and the governance requirement information submitted includes "customer names and addresses need to be desensitized and converted to Parquet format for big data analysis." First, based on the data type "relational database table," all initial governance policies applicable to relational databases are extracted from the pre-set policy database, such as "SQL data cleansing policy," "relational data desensitization policy," "relational data encryption policy," and "relational data format conversion policy." Keyword tags are also extracted from the governance requirement information, such as "desensitization," "name," "address," "format conversion," "Parquet," and "big data analysis." The system then matches these keyword tags with the extracted initial policies. For example, the "relational data desensitization policy" might contain metadata or descriptions related to "desensitization," "name," and "address," and is therefore matched and selected. The "relational data format conversion policy" might contain metadata related to "format conversion" and "Parquet," and is therefore also matched and selected. The "SQL data cleansing policy" and "relational data encryption policy" might not match the extracted keyword tags and are therefore not selected. Ultimately, the candidate strategy set may include "relational data desensitization strategy" and "relational data format conversion strategy".
[0054] Specifically, when the selection module 2 adopts a rule-based conflict detection method to identify and remove conflicting policies in the candidate policy set to obtain a streamlined policy set, it executes: Based on the preset conflict strategy rules, a conflict strategy combination is extracted from the candidate strategy set; According to the preset policy priority, determine the low priority policy in the conflicting policy combination; Low-priority policies are eliminated from the candidate policy set to obtain a streamlined policy set.
[0055] When selection module 2 is executed, to address the issue of conflicting policies within the candidate policy set, a conflict detection and resolution process is performed after extracting the candidate policy set. First, the policies in the candidate policy set are scanned and compared according to the preset conflict policy rules to identify all conflicting policy combinations. For example, if the candidate set includes a policy requiring the complete deletion of a sensitive field and a policy requiring format conversion of the field, and the preset rules define these two operations as conflicting on the same field, the two policies are identified as a conflicting policy combination. Next, each policy in the conflicting policy combination is evaluated according to the preset policy priority. For example, if the deletion policy has a higher priority than the format conversion policy, the format conversion policy is determined to be a low-priority policy. Finally, the identified low-priority policies are removed from the candidate policy set, resulting in a streamlined policy set that contains no internal conflicts. This streamlined policy set provides a reliable foundation for subsequent policy orchestration, avoiding data governance errors or inefficiencies caused by policy conflicts. The preset conflict policy rule is to execute only one policy of the same type for the same field or for sentences or paragraphs corresponding to the same field.
[0056] For example, assume that the candidate policy set extracted from the preset database includes Policy A (data encryption), Policy B (data desensitization), and Policy C (data compression). The preset conflicting policy rules define Policy A and Policy B as a conflicting policy combination. The preset policy priorities set Policy A to a priority of 10, Policy B to a priority of 8, and Policy C to a priority of 5. Based on the preset conflicting policy rules, it is identified that Policy A and Policy B constitute a conflicting policy combination. Based on the preset policy priorities, Policy A (priority 10) is compared with Policy B (priority 8), and Policy B is determined to have a lower priority. Policy B is removed from the candidate policy set {Policy A, Policy B, Policy C}, resulting in the streamlined policy set {Policy A, Policy C}.
[0057] Specifically, when selecting module 2 sorts the data governance policies in the streamlined policy set based on the preset policy arrangement rules and obtains the governance policy chain, it executes: Determine the corresponding migration scenario information based on the data information; migration scenario information includes network bandwidth, cloud platform computing resources, and data storage format; According to the migration scenario information and the preset policy orchestration rules, a topological sorting algorithm is used to sort the data governance policies in the streamlined policy set to obtain a governance policy chain.
[0058] The selection module 2, when executed, determines corresponding migration scenario information according to data information of the data to be migrated, such as data size, data type, and the like, in combination with the current data migration environment. The migration scenario information can include current network bandwidth conditions, availability of computing resources of the target cloud platform, and storage formats of the data to be migrated, and the like, which affect the efficiency and feasibility of policy execution. These information reflect the actual environmental constraints during policy execution.
[0059] Specifically, when the selection module 2 sorts the data governance strategies in the simplified strategy set according to the migration scenario information and the preset strategy arrangement rule, and adopts a topological sorting algorithm to obtain a governance strategy chain, the following is performed: Based on the migration scenario information, a directed acyclic graph corresponding to the simplified strategy set is constructed, with data governance targets as the terminal points, each data governance strategy as a node, and the dependency relationship between each data governance strategy as an edge. A topological sorting algorithm considering weights is adopted to calculate the execution priority of each data governance strategy in the simplified strategy set according to the weights of each edge in the directed acyclic graph, to obtain a preliminary governance strategy chain. The preliminary governance strategy chain is optimized according to the preset strategy arrangement rule to obtain a governance strategy chain.
[0060] The selection module 2, when executed, constructs a directed acyclic graph according to the current data information (such as data type, data size) and the migration scenario information (such as source / target cloud platform type, network bandwidth, available computing resources). The nodes in the graph correspond to each strategy in the simplified strategy set, and the terminal node represents the data governance completion state. The dependency relationship between strategies (for example, data encryption must be performed before data compression) is represented as a directed edge. Further, according to the migration scenario information, a weight is assigned to each edge in the graph. For example, in a network bandwidth limited scenario, the edge weight corresponding to the subsequent strategy (such as data transmission) of the data compression strategy may be adjusted to reflect the importance or execution efficiency of the compression strategy. A weighted topological sorting algorithm is used to process the directed acyclic graph to calculate the execution priority of each strategy and generate a preliminary strategy execution order chain. The preliminary strategy chain satisfies the dependency relationship between strategies and takes into account the relative execution cost or efficiency of the strategy in the current scenario. Finally, the preliminary strategy chain is checked and adjusted according to the preset strategy arrangement rule (for example, the data desensitization strategy is always executed before the data quality verification strategy), to obtain a final strategy chain for data migration and governance. The final strategy chain comprehensively considers strategy dependency, scenario characteristics, and business rules, improving the efficiency and effectiveness of data governance. Among them, the preset strategy arrangement rule is a strategy arrangement rule that prioritizes the execution of data desensitization strategies, then executes data format conversion strategies, and finally executes data quality verification strategies.
[0061] For example, suppose the data to be migrated requires the application of four governance policies: data desensitization (S1), data format conversion (S2), data quality verification (S3), and data compression (S4). The known dependencies are that S1 must be executed before S2, S2 must be executed before S4, and S3 has no direct preceding dependencies. The migration scenario involves low network bandwidth and sufficient computing resources on the target cloud platform. First, construct a directed acyclic graph: nodes S1, S2, S3, S4, and endpoint E. The edges are S1→S2, S2→S4, S3→E, and S4→E. Next, assign weights based on the migration scenario. Because low network bandwidth significantly impacts subsequent transmission efficiency, S4 (data compression) can be given a higher weight, indicating that the steps before S4 are prioritized. The weights of other edges are determined based on factors such as policy execution time. Using a weighted topological sorting algorithm, a preliminary policy chain may be obtained: S1→S2→S4→S3. This order satisfies the dependencies and takes into account the importance of the compression policy. Finally, optimization is performed based on the pre-set policy orchestration rules. For example, the pre-set rules dictate that data desensitization policies (S1) must be executed first, and data quality verification policies (S3) must be executed last. The initial policy chain, S1 → S2 → S4 → S3, complies with the S1 priority and S3 last rule. Therefore, the final governance policy chain is determined to be S1 → S2 → S4 → S3. This policy chain not only satisfies dependencies and rules, but also takes into account network bandwidth limitations through weighted sorting, optimizing overall execution efficiency.
[0062] Specifically, the governance module 3 performs data governance while migrating the data to be migrated based on the governance policy chain and data information. When the data governance record is obtained, it executes: Extract the data to be migrated from the original location of the data in the data information; Based on the governance policy chain, data governance is performed on the data to be migrated to obtain the data to be migrated after data governance; After migrating the data to be migrated after data governance to the target migration location in the data information, a data governance record is recorded.
[0063] When the governance module 3 is executed, it determines the original location of the data to be migrated based on the data information and obtains the data to be migrated from this location. Using the predetermined governance policy chain, a series of governance operations are performed on the data to be migrated, such as desensitizing sensitive information, standardizing data formats, cleaning data values, etc., to ensure that the data is effectively processed before or during migration and to generate data that meets governance requirements. The governed data is transferred to the target migration location specified by the data information. After the data is successfully migrated, the system records the detailed information of this operation, including the source and destination of the data, the applied governance policy, the execution results, etc., to form a data governance record. By embedding the data governance steps into the data migration process, it is ensured that only governed data will be migrated to the target location, thereby reducing the security and compliance risks faced by the data during the migration process and improving the reliability of data migration.
[0064] Specifically, when executed, Update Module 4 writes the generated data governance records to a distributed ledger. The tamper-proof nature of the distributed ledger ensures the authenticity and integrity of the data governance records. By querying the distributed ledger, staff or auditors can easily trace the complete migration path of the migrated data and all governance operations performed on it during the migration process. This provides end-to-end data processing transparency, greatly enhancing the credibility and compliance of data processing, and effectively reducing the risk of data leakage and illegal operations.
[0065] As can be seen from the above, the adaptive data governance system obtains data information and governance requirement information of the data to be migrated by responding to the cross-cloud data migration request initiated by the user, selects the corresponding data governance policy from the preset database according to the data information and governance requirement information, obtains the governance policy chain, and performs data governance on the data to be migrated while migrating the data based on the governance policy chain and data information, obtains data governance records, and updates the data governance records to the distributed ledger so that the staff can trace the migration path and governance operation records of the data to be migrated through the distributed ledger; thus, by selecting the governance policy chain based on the data information and governance requirement information of the data to be migrated, data governance is performed on the data to be migrated while migrating the data, obtains data governance records, and updates the data governance records to the distributed ledger, solving the problems of low data governance efficiency and easy data leakage in existing data governance methods when working across cloud platforms, adaptively selecting governance strategies and performing governance simultaneously during data migration, and at the same time uploading governance records to the chain, thereby improving the governance efficiency during data migration.
[0066] In the embodiments of the present application, it should be understood that the disclosed system and method can be implemented in other manners. The embodiments described above are merely exemplary, for example, the division of the units is only a logical function division, and there can be another division manner in actual implementation; for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, or the among different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0067] In addition, the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, and can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0068] In addition, the functional modules in the various embodiments of the present application can be integrated together to form a separate part, or each module can exist independently, or two or more modules can be integrated to form a separate part.
[0069] In this document, the terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations.
[0070] The above description is merely exemplary of the application, and is not intended to limit the protection scope of the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall be included in the protection scope of the application.
Claims
1. An adaptive data governance method for managing data during data migration, characterized in that: Including steps: Responding to user-initiated cross-cloud data migration requests, obtaining data information and governance requirements for the data to be migrated; According to the data information and the governance requirement information, a corresponding data governance policy is selected from a preset database to obtain a governance policy chain; Based on the governance policy chain and the data information, data governance is performed on the data to be migrated while migrating the data to obtain a data governance record; The data governance records are updated on the distributed ledger so that staff can trace the migration path and governance operation records of the data to be migrated through the distributed ledger.
2. The adaptive data governance method according to claim 1, characterized in that: The data information includes the original location of the data to be migrated, the target migration location, the data type and the data size.
3. The adaptive data governance method according to claim 2, characterized in that: According to the data information and the governance requirement information, a corresponding data governance policy is selected from a preset database to obtain a governance policy chain, including: Extracting a candidate strategy set from the preset database according to the data information and the governance requirement information; Using a rule-based conflict detection method, identifying and removing conflicting policies in the candidate policy set to obtain a streamlined policy set; Based on preset policy arrangement rules, the data governance policies in the streamlined policy set are sorted to obtain a governance policy chain.
4. The adaptive data governance method according to claim 3, characterized in that: According to the data information and the governance requirement information, a candidate strategy set is extracted from the preset database, including: Extracting an initial data governance strategy corresponding to the data type from the preset database according to the data type in the data information; Extracting corresponding keyword tags from the governance demand information; The keyword tags are matched with the initial data governance policies to obtain a candidate policy set.
5. The adaptive data governance method according to claim 3, characterized in that: A rule-based conflict detection method is used to identify and remove conflicting policies in the candidate policy set to obtain a streamlined policy set, including: Extracting a conflicting strategy combination from the candidate strategy set based on preset conflicting strategy rules; Determine the low-priority strategy in the conflicting strategy combination according to the preset strategy priority; The low-priority policies are eliminated from the candidate policy set to obtain a streamlined policy set.
6. The adaptive data governance method according to claim 3, characterized in that: Based on the preset policy arrangement rules, the data governance policies in the simplified policy set are sorted to obtain a governance policy chain, including: Determine corresponding migration scenario information based on the data information; the migration scenario information includes network bandwidth, cloud platform computing resources, and data storage format; According to the migration scenario information and the preset policy arrangement rules, a topological sorting algorithm is used to sort the data governance policies in the streamlined policy set to obtain a governance policy chain.
7. The adaptive data governance method according to claim 6, characterized in that: According to the migration scenario information and the preset policy arrangement rules, a topological sorting algorithm is used to sort the data governance policies in the streamlined policy set to obtain a governance policy chain, including: Based on the migration scenario information, a directed acyclic graph corresponding to the streamlined policy set is constructed, with the data governance goal as the end point, each data governance policy as a node, and the dependency relationship between each data governance policy as an edge; Using a weighted topological sorting algorithm, according to the weight of each edge in the directed acyclic graph, the execution priority of each data governance policy in the streamlined policy set is calculated to obtain a preliminary governance policy chain; According to the preset policy arrangement rules, the preliminary governance policy chain is optimized to obtain a governance policy chain.
8. The adaptive data governance method according to claim 6 or 7, characterized in that: The preset policy arrangement rule is to first execute data desensitization policies, then execute data format conversion policies, and finally execute data quality verification policies.
9. The adaptive data governance method according to claim 1, characterized in that: Based on the governance policy chain and the data information, data governance is performed on the data to be migrated while migrating the data to obtain a data governance record, including: Extracting the data to be migrated from the original location of the data in the data information; Based on the governance policy chain, data governance is performed on the data to be migrated to obtain the data to be migrated after data governance; After migrating the data to be migrated after data governance to the target migration location in the data information, a data governance record is recorded.
10. An adaptive data management system for managing data during data migration, characterized in that: include: An acquisition module, configured to respond to a cross-cloud data migration request initiated by a user and obtain data information and governance requirement information of the data to be migrated; A selection module, configured to select a corresponding data governance policy from a preset database according to the data information and the governance requirement information, and obtain a governance policy chain; A governance module, configured to perform data governance while migrating the data to be migrated based on the governance policy chain and the data information, and obtain a data governance record; An update module is used to update the data governance record to the distributed ledger so that staff can trace the migration path and governance operation record of the data to be migrated through the distributed ledger.
Citation Information
Patent Citations
Method for creating data governance task and electronic equipment
CN114817227A
Data governance method and system capable of realizing different governance degrees
CN116226108A
Data migration method and device, equipment, storage medium and product
CN118550898A
Data management method and system based on big data
CN118626800A
Data management method and device, equipment and storage medium
CN118761618A
Cited By
Data management method and device, computer equipment and storage medium
CN121743312A