Data lineage-based data governance method and apparatus, and electronic device
Patent Information
- Application Number
- PCT/CN2026/070442
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-01-05
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026070442_01102026_PF_FP_ABST
Abstract
Description
A data governance method, device, and electronic device based on data lineage Technical Field
[0001] This application relates to the technical field of data governance, specifically to a data governance method, apparatus, and electronic device based on data lineage. Background Technology
[0002] In modern data governance, assessing the importance of data is a key step in developing governance strategies, allocating resources, and optimizing data management processes.
[0003] Currently, the importance of data is determined through manual rules of thumb or simple qualitative analysis. For example, the governance priority of data is directly assessed manually based on departmental affiliation, frequency of use, or business needs. However, this qualitative approach lacks unified quantitative standards, leading to significant subjectivity and inconsistency in the assessment results. This method, which relies solely on manual assessment of data governance priorities, is no longer sufficient to meet the needs of precise governance.
[0004] Therefore, there is an urgent need for a data governance method, device, and electronic equipment based on data lineage. Summary of the Invention
[0005] This application provides a data governance method, apparatus, and electronic device based on data lineage, which facilitates precise data governance.
[0006] This application provides a data governance method based on data lineage, the method comprising: acquiring metadata of target data corresponding to target governance requirements; constructing a data lineage graph of the target data based on the metadata; calculating, through the data lineage graph, a data quality propagation score, a data lineage path complexity, a data global impact score, a data leakage risk score, and a data usage efficiency score of the target data; generating a comprehensive data governance score based on the data quality propagation score, the data lineage path complexity, the data global impact score, the data leakage risk score, and the data usage efficiency score; and performing data governance on the target data according to the comprehensive data governance score.
[0007] By adopting the above technical solutions, multiple key factors such as data governance quality, complexity, security, and efficiency are comprehensively considered through data quality propagation scores, path complexity, global impact scores, data leakage risk scores, and usage efficiency scores, avoiding the one-sidedness of single-indicator evaluation. This approach satisfies both quality optimization needs and efficiency improvement and security risk management, adapting to governance requirements in different scenarios. Quantifying various scores provides an objective basis for data priority assessment, avoiding the subjective bias of traditional qualitative methods. Constructing a data lineage graph provides visualization and precise analysis of data dependencies, ensuring a complete depiction of data flow paths and impact scope. Generating a comprehensive score using multiple indicators provides a unified governance priority ranking, helping to rationally allocate governance resources and improve the scientific nature of decision-making. Weights can be dynamically adjusted according to changes in needs or indicators, and the comprehensive score is updated in real time to adapt to governance goals at different stages. The automatic construction of the data lineage graph and indicator calculation achieves efficient data evaluation, reduces manual intervention, and increases analysis speed. Comprehensive score-driven priority governance ensures that governance resources are concentrated on the most important data, improving the overall efficiency and effectiveness of governance. Data leakage risk scoring helps identify potential security risks in advance, reducing data leakage and compliance risks. It identifies critical data through global impact scoring, preventing cascading business risks from erroneous handling of important data. Therefore, it facilitates precise data governance.
[0008] Optionally, constructing the data lineage graph of the target data based on the metadata includes: parsing the metadata to obtain data entities, the data entities including database tables, fields, and files; extracting edges based on data entities with dependencies among multiple data entities; obtaining additional data between the data entities and the edges; determining attributes based on the additional data; and constructing the data lineage graph of the target data based on the data entities, the edges, and the attributes.
[0009] By adopting the above technical solutions, data entities such as database tables, fields, and files are uniformly parsed, ensuring that the data lineage graph construction covers the entire data flow path. By extracting dependencies between entities, the upstream and downstream flow and correlation of data are revealed, supporting comprehensive tracking of data sources and impacts. Additional data is used to determine attributes, further supplementing the detailed information of the data lineage graph and making its semantics clearer. Parsing dependencies between entities ensures that edge generation is based on actual data flow logic, avoiding redundant or misleading information. The introduction of attributes allows the data lineage graph to not only display data entities and dependencies but also provide contextual information, supporting more precise analysis. Metadata parsing and automatic dependency extraction significantly reduce manual intervention and improve construction efficiency. The constructed data lineage graph can directly serve multi-dimensional calculations such as data quality propagation and global impact analysis, significantly improving the efficiency of data governance decisions. The combination of entities, edges, and attributes presents complex metadata relationships in a structured graph form, facilitating understanding and operation. The constructed lineage graph can be visualized using graphical tools, enabling data governance teams to quickly locate key data and dependency paths. Additional attributes in the lineage graph can be dynamically updated as business needs change, ensuring its timeliness and usability. This method supports various data entities (such as tables, fields, and files), adapting to the complex governance requirements of different data sources.
[0010] Optionally, the data quality propagation score is calculated using the following formula:
[0011] Where Q is the data quality propagation score, i is the i-th metadata, and q i Let be the quality score of the i-th metadata, n be the total number of metadata, e be an edge, e∈E represent the set of edges in path e, and d(e) be the quality decay rate of edge e in path e.
[0012] By adopting the above technical solution and using formula calculations, the previously ambiguous issue of data quality propagation is transformed into concrete numerical values, providing a basis for objective assessment. Combining the quality score of each metadata element with the quality decay rate of edges along the path, the changes in data quality during the flow process are comprehensively considered, reflecting the actual situation. The formula incorporates the quality decay rate of path edges, accurately capturing the quality loss of data during complex flow processes, solving the problem of insufficient estimation of the impact of data transmission in traditional methods. This method can adapt to complex data lineage graph structures with multiple paths and nodes, and can be flexibly applied to various complex data environments. Based on the quality propagation score, key nodes or paths with low quality or high decay can be quickly located, helping to prioritize the governance of problematic data. Through quantitative analysis results, governance resources are concentrated and allocated to key paths and nodes, improving the overall efficiency of governance. This formula is based on dynamic data input and can be updated in real time according to changes in data quality and flow, adapting to governance needs at different points in time. By providing quantitative diagnostics of data quality, it assists decision-makers in determining which paths or nodes have the greatest impact on overall quality. The data quality propagation score can serve as a priority reference indicator for data governance, providing a strong basis for governance planning.
[0013] Optionally, the complexity of the data lineage path is calculated using the following formula:
[0014] Where L is the data lineage path complexity, t is time, e is an edge, e∈E represents the set of edges in path e, w(e,t) is the weight of edge e at time t, and f(v) is the complexity function associated with node v.
[0015] By adopting the above technical solutions, the formula, by introducing the weights of path edges at different times, can reflect the dynamic changes of complex paths over time and capture the temporal characteristics in real-world scenarios. By quantifying the complexity associated with nodes through functions, it ensures that path complexity assessment not only depends on the path itself but also comprehensively considers the complex attributes of nodes. The formula can handle complex data lineage networks with multiple paths and nodes, and is particularly suitable for large-scale, deep-layered data governance scenarios. By combining edge and node attributes on the path, it avoids one-sided results caused by single-dimensional complexity assessment, and more comprehensively describes the complex characteristics of the data lineage graph. By calculating the complexity of data lineage paths, key high-complexity paths or nodes can be quickly located, providing a clear direction for optimizing data flow. High-complexity paths may negatively impact data processing performance or governance efficiency; the calculation results of this formula can serve as a basis for resource allocation and process optimization. The formula considers the characteristics of path edge weight changes over time, enabling real-time complexity analysis and adapting to dynamic changes in business needs and data structures. Complexity assessment can be combined with real-time monitoring to provide data support for the dynamic adjustment of governance strategies. Quantitative complexity assessment can help governance teams optimize data flow paths, reduce redundant operations, and improve data processing efficiency. Identifying complex paths helps reduce data quality degradation or security risks caused by overly complex paths.
[0016] Optionally, the global impact score of the data is calculated using the following formula:
[0017] Where I represents the global impact score of the data, and V d For the set of nodes directly affected, α v Let σ(v) be the weight coefficient of node v, σ(v) be the sensitivity score of node v, K be the maximum propagation level, k be the k-th level, and β be the weight coefficient of node v. k I represents the weight of level k. k The score represents the influence of level k.
[0018] By employing the aforementioned technical solutions, the formula accurately assesses the scope and depth of data's impact throughout the entire lineage graph by combining the set of directly affected nodes and the propagation hierarchy. Node weight coefficients and hierarchy weights assign differentiated weights to different nodes and levels, reflecting the priority and importance in actual business needs. Node sensitivity scores consider the importance of data nodes in business processes, ensuring that impact assessments are more aligned with real-world scenarios. The high weighting of sensitivity scores helps to quickly identify and prioritize the protection of data at critical nodes, reducing the risk of data leakage or misuse. The formula introduces the maximum hierarchy and impact scores for each hierarchy, enabling a deep characterization of the global diffusion characteristics of data in multi-level networks. For complex multi-path lineage graphs, this formula can comprehensively assess the impact of different paths and levels, without overlooking any potential critical nodes. Weight coefficients support real-time adjustment to adapt to dynamic changes in different business scenarios or data environments. When data or business changes, this method can quickly recalculate the global impact score, ensuring the timeliness of governance strategies. The global data impact score provides a quantitative basis for data governance priorities, helping to quickly identify the critical data nodes and paths with the greatest impact on business. Based on global impact scores, governance resources can be concentrated on high-impact nodes or paths, improving governance efficiency and effectiveness. By identifying high-impact paths or nodes, global risks that may arise from chain failures or misoperations can be detected and mitigated in a timely manner. Global data impact analysis can help protect critical business data flow links and ensure the stable operation of core businesses.
[0019] Optionally, the data leakage risk score is calculated using the following formula:
[0020] Where R is the data leakage risk score, v is the node, and V is the data leakage risk score. s Let be the set of nodes containing sensitive data, π(v) be the probability that node v is accessed, and r(v) be the leakage risk score of node v.
[0021] By adopting the above technical solution, the formula focuses on the set of sensitive data nodes, clarifying that sensitive data is the core of leakage risk assessment and ensuring that resources are prioritized for protecting high-risk data. The probability of a node being accessed reflects the dynamic nature of leakage risk, helping to accurately predict the potential risks of frequently accessed data. Node access probability and leakage risk score can be dynamically updated based on real-time data, helping to address constantly changing business environments and security threats. By combining access probability and leakage risk score, the formula comprehensively considers the frequency of data usage and the risk characteristics of the node itself. The calculation results can quickly locate the node with the highest leakage risk, helping to formulate targeted security strategies and avoid resource dispersion and waste. The formula provides a quantitative basis for the classification, access control, and encryption protection of sensitive data, improving the scientific nature of security governance. By calculating the data leakage risk score in the data lineage graph, a global visual map of data leakage risk can be generated, clearly showing high-risk nodes and their impact range. As node access behavior changes, the data leakage risk score can be updated in real time, helping the team continuously monitor and prevent potential security risks. Based on the data lineage graph, the formula extends the risk assessment of a single node to the entire data flow path, comprehensively considering the leakage risks that may be caused by chain propagation. Focusing on sensitive data and access behavior, we help companies better meet legal and compliance requirements for data privacy protection.
[0022] Optionally, the data utilization efficiency score is calculated using the following formula:
[0023] Where M is the data usage efficiency score, T is the maximum time, t is the time, U(t) is the number of data accesses at time t, and C(t) is the data maintenance cost at time t.
[0024] By adopting the above technical solutions, the formula combines data access frequency and data maintenance cost, providing a comprehensive quantitative assessment of data usage efficiency and avoiding the one-sidedness of considering only one aspect. By comparing access frequency and maintenance cost, it identifies efficient data usage while helping to discover resource waste or overconsumption. The formula introduces a time variable, capturing the dynamic changes in data usage efficiency at different points in time and reflecting time-series characteristics. As time progresses, by dynamically calculating efficiency scores, data access and maintenance strategies can be identified and optimized in real time. The formula helps quickly locate inefficient data with high maintenance costs and low access frequency, providing the governance team with a clear direction for optimization. Based on the efficiency score, data storage, computing, and operation and maintenance resources can be rationally allocated, improving overall data governance efficiency. The data usage efficiency score provides a scientific basis for data governance prioritization, supporting tiered governance of efficient or inefficient data. By analyzing the time series changes in efficiency scores, long-term trends can be identified, leading to more scientific resource investment and usage plans. By analyzing data maintenance costs, costly operations can be identified and optimized in a targeted manner, thereby reducing overall operating costs. Resources are concentrated on frequently used, higher-value data, avoiding excessive resource consumption by inefficient data.
[0025] Optionally, generating a comprehensive data governance score based on the data quality propagation score, the data lineage path complexity, the data global impact score, the data leakage risk score, and the data usage efficiency score includes: determining the ratio between the data quality propagation score and the data lineage path complexity; calculating a first product between the ratio and a first weight; calculating a second product between the data global impact score and a second weight; calculating a third product between the data leakage risk score and a third weight; calculating a fourth product between the data usage efficiency score and a fourth weight; summing the first, second, and fourth product results and subtracting the third product result to obtain the comprehensive data governance score.
[0026] By adopting the above technical solutions, the method comprehensively considers multiple key dimensions such as data quality propagation score, path complexity, global impact, leakage risk, and usage efficiency to fully reflect the overall status of data governance. By combining different dimensions with weights, the method achieves a balance among multiple indicators, avoiding the limitations of a single indicator deviating from reality. Different weight settings can reflect business objectives, such as prioritizing data quality or security, helping to make more accurate decisions. The ratio of data quality propagation score to path complexity serves as input, reflecting the propagation effect of data quality under the influence of path complexity, enhancing the logical rigor of the scoring. Through weighted summation and difference operations on different indicators, the scoring results are more scientific and can accurately reflect the quality of data governance. The comprehensive score provides an intuitive quantitative result for different data objects, which can be used for rapid sorting and determining the priority of data governance. Comparative analysis of high-scoring and low-scoring data helps the governance team quickly identify governance areas that require key attention. The comprehensive score can be updated in real time as the scores of each indicator change, reflecting the dynamic changes in data governance effectiveness and providing a basis for continuous optimization. The comprehensive score helps enterprises regularly evaluate the effectiveness of governance strategies to ensure the gradual achievement of governance goals. The comprehensive scoring provides a unified standard for data governance, facilitating communication and collaboration among different departments and teams. Based on the scoring results, governance resources can be prioritized for high-priority data, significantly improving governance efficiency.
[0027] A second aspect of this application provides a data governance apparatus based on data lineage. The data governance apparatus includes an acquisition module and a processing module. The acquisition module is used to acquire metadata of target data corresponding to a target governance requirement. The processing module is used to construct a data lineage graph of the target data based on the metadata. The processing module is further used to calculate, through the data lineage graph, a data quality propagation score, a data lineage path complexity, a data global impact score, a data leakage risk score, and a data usage efficiency score for the target data. The processing module is also used to generate a comprehensive data governance score based on the data quality propagation score, the data lineage path complexity, the data global impact score, the data leakage risk score, and the data usage efficiency score. The processing module is further used to perform data governance on the target data according to the comprehensive data governance score.
[0028] A third aspect of this application provides an electronic device, which includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, and both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method described above.
[0029] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described above. Attached Figure Description
[0030] Figure 1 is a flowchart illustrating a data governance method based on data lineage provided in an embodiment of this application;
[0031] Figure 2 is another flowchart illustrating a data governance method based on data lineage provided in an embodiment of this application;
[0032] Figure 3 is a schematic diagram of a data governance device based on data lineage provided in an embodiment of this application;
[0033] Figure 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0034] Explanation of reference numerals in the attached figures: 31. Acquisition module; 32. Processing module; 41. Processor; 42. Communication bus; 43. User interface; 44. Network interface; 45. Memory. Detailed Implementation
[0035] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0036] In the description of the embodiments in this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0037] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0038] This application provides a data governance method based on data lineage. Figure 1 is a flowchart illustrating such a method. This data governance method is applied to a server and includes:
[0039] S110. Obtain the metadata of the target data corresponding to the target governance requirements.
[0040] Specifically, a server refers to the computing device or system that hosts data governance work; it is the main execution unit for data analysis and processing. Target governance requirements are specific requirements or problems related to data governance, such as improving data quality, optimizing query efficiency, and reducing security risks. Target data refers to a specific dataset or data resource directly related to the governance requirements, such as a database table or a specific data file. Metadata describes the target data, including data structure (table name, field name, field type); data source (data generation time, collection channel); data relationships (primary key, foreign key, dependency relationships); and data operation logs (modification time, access count, etc.). The server obtains this metadata through interfaces or tools (such as a metadata management system) to support subsequent data governance.
[0041] For example, suppose a company wants to optimize the efficiency of report generation in its financial system. The target governance requirement is to improve the efficiency of financial report generation and reduce generation delays. The target data is the report-related data stored in the financial database, such as "Revenue Report" and "Expense Report." The server extracts the metadata of this data through a metadata management system or database management tools. Metadata includes the structure of the data tables, such as field names and field types, as well as data storage relationships, such as primary and foreign key relationships between tables. In addition, it includes data volume information, such as the number of rows in the data table and field sizes. It also includes data access logs, such as which queries frequently use this data and the access paths when generating reports. Metadata is used to analyze the reasons for slow report generation, such as whether it is due to missing indexes on certain fields or overly complex join queries. At the same time, it optimizes query paths, reduces access to useless fields, and improves data retrieval speed.
[0042] S120. Construct a data lineage diagram of the target data based on the metadata.
[0043] Specifically, the server extracts data lineage information from metadata, including: entity information, such as specific data entities like tables, fields, and files; relationship information, such as dependencies between data entities, for example, a field in one table originating from another; and additional attributes, such as data generation time, operation records, and sensitive information markers. Each data entity is treated as a node; for example, the `Sales` table is one node, and the `Revenue` field is another. If one data entity depends on another, an edge is added between them. For example, the `Revenue` field might be calculated from the `Orders` table via ETL. A graph structure is built based on these nodes and edges and displayed visually. Examples include: intra-table lineage (fields depend on other fields); cross-table lineage (tables are connected via primary / foreign keys or ETL processes).
[0044] In one possible implementation, a data lineage graph of the target data is constructed based on metadata, specifically including: parsing the metadata to obtain data entities, which include database tables, fields, and files; extracting edges from data entities with dependencies among multiple data entities; obtaining additional data between data entities and edges; determining attributes based on the additional data; and constructing a data lineage graph of the target data based on data entities, edges, and attributes.
[0045] Specifically, the server parses metadata to extract core information, including data entities and their relationships. Metadata may come from database management systems (DBMS), data warehouse tools, or ETL process logs. Data entities are the specific building blocks of data, including: database tables: collections storing structured data, such as the Orders table; fields: detailed units within a data table, such as OrderID, Price, etc.; files: unstructured or semi-structured data storage units, such as CSV and JSON files. Edges represent dependencies between data entities, reflecting the flow of data from source to destination. Edges may be defined by foreign key relationships, data computation logic (such as ETL), application operations, etc. Additional data describes the characteristics of data entities and edges, such as: edge weights: indicating the importance of dependencies; timestamps: marking the time of data updates or processing; operation types: insert, delete, modify, etc. Attributes are used to further characterize the features of nodes and edges, such as node sensitivity and edge transmission efficiency.
[0046] Further, firstly, the server parses the metadata, extracts data entities, and extracts entity information such as tables, fields, and files from the metadata, along with basic descriptions of the entities (table names, field names, data types, etc.). Secondly, the server extracts the dependencies between data entities, determining these dependencies based on information in the metadata, such as foreign key relationships, ETL logs, or data processing scripts. Each dependency is represented by an edge, indicating the direction from the source entity to the target entity. Then, the server acquires additional data, collecting additional information about the entities and edges, such as data transfer frequency, access permissions or sensitivity levels, and the calculation logic of the dependencies. The server uses this additional data to determine the attributes of nodes (data entities) and edges (dependencies), for example: nodes: table update time, field sensitivity level; edges: dependency strength (e.g., high-frequency dependencies have higher weights). Finally, the server uses data entities as nodes, dependencies as edges, and attributes as graph features to generate the final data lineage graph. The data lineage graph is a directed graph representing the entire path of data from the source to the target.
[0047] S130. Calculate the data quality propagation score, data lineage path complexity, global data impact score, data leakage risk score, and data usage efficiency score of the target data through the data lineage graph.
[0048] Specifically, the server calculates various scores for the target data using a data lineage graph, describing the process of multi-dimensional analysis based on the data lineage graph. The purpose of these scores is to comprehensively evaluate the target data's performance in terms of quality, complexity, impact, risk, and efficiency of use in a quantitative manner, thereby providing a reliable basis for data governance. The data quality propagation score is used to assess the transmission and changes in data quality along the data lineage path, reflecting the overall quality performance of data from source to target.
[0049] Data lineage graphs trace the flow of data from upstream to downstream, analyzing the quality of each data entity and considering the impact of each path on the quality of the target data. Assume SalesReport.NetProfit depends on Orders.Revenue and Products.Cost. If the data quality of Orders.Revenue is low (e.g., many missing values), this defect will propagate along the lineage path to NetProfit, causing its quality score to decrease. Data lineage path complexity measures the complexity of the target data's lineage path structure, including the number of nodes and edges, path depth, and the complexity of dependencies. The server analyzes the topology of the data lineage graph to identify the number of nodes and edges, the depth of levels, and the degree of nesting of relationships in the path. If the generation path of SalesReport.NetProfit is complex, requiring multiple intermediate nodes (such as Orders.Revenue and Products.Profit) and multiple edges, the complexity score will be high. This means that path optimization is needed in governance to reduce maintenance costs. The global data impact score assesses the scope and extent of the impact of changes or problems in the target data on the entire system. The server quantifies the impact of target data on downstream data entities and the system as a whole by analyzing the downstream nodes and paths of target data in the data lineage graph. For example, suppose Products.Cost is a key data source for SalesReport.NetProfit, and reports from multiple other departments also rely on Cost. If an error occurs in Cost, it will affect a large number of downstream reports, resulting in a high global impact score.
[0050] Furthermore, the data leakage risk score is used to assess the potential leakage risk of target data along the lineage path, especially the exposure probability of sensitive data. The server identifies sensitive data nodes involved in the lineage path, analyzing the frequency of data access, the security level of storage locations, and encryption and access control during transmission. For example, if Products.Cost is a highly sensitive field and it is transmitted to SalesReport.NetProfit via an unencrypted path, there is a high risk of data leakage. The data usage efficiency score measures the relationship between the value of target data in the system and its maintenance cost, reflecting the efficiency of data resource utilization. The server analyzes data utilization efficiency by statistically analyzing the access frequency and call volume of target data, combined with its storage and processing costs. For example, if Orders.Revenue is accessed multiple times daily for different analytical tasks, but its storage and computation costs are low, its usage efficiency score is high. Conversely, some data, although computationally complex, has a low usage frequency, resulting in a lower efficiency score. By analyzing the data lineage graph, the server can evaluate target data from multiple dimensions, including quality, structural complexity, global impact, leakage risk, and usage efficiency. These scores are quantitative analyses based on data lineage diagrams, providing comprehensive control over the entire data lifecycle, helping to identify key and weak points in data governance, and providing data-driven quantitative evidence for data governance decisions.
[0051] In one possible implementation, the data quality propagation score is specifically calculated using the following formula:
[0052] Where Q is the data quality propagation score, i is the i-th metadata, and q i Let be the quality score of the i-th metadata, n be the total number of metadata, e be an edge, e∈E represent the set of edges in path e, and d(e) be the quality decay rate of edge e in path e.
[0053] Specifically, the data quality propagation score is a quantitative metric used to assess the degradation and accumulation of data quality as it is transmitted from upstream source data to target data. Its purpose is to measure the contribution of the quality of each data entity along the data lineage path to the quality of the target data, considering the quality degradation of data along each path. Here, the data quality propagation score represents the overall quality of the target data. The i-th metadata refers to each data entity or field in the lineage path. The quality score of the i-th metadata is derived based on assessments of field completeness, accuracy, etc. The total number of metadata is the number of nodes in the path. Edges in the path represent dependencies between data entities. The quality degradation rate of path edge e represents the degree of quality reduction during the transmission of data from upstream to downstream.
[0054] The server first extracts the paths related to the target data from the data lineage graph, including upstream nodes (metadata) and edges (dependencies), and calculates a quality score for each metadata node. The quality score can be evaluated across the following dimensions: data integrity (whether there are missing values), data accuracy (whether it conforms to business logic), and data consistency (whether it matches across systems). Next, the server analyzes the decay rate of each edge in the path. The decay rate may be related to information loss caused by data transformations (such as aggregation and filtering) and errors or biases present during data processing. Finally, the server calculates a comprehensive quality propagation score for the target data based on the quality scores of the nodes in the path and the decay rates of the edges.
[0055] In one possible implementation, the data lineage path complexity is specifically calculated using the following formula:
[0056] Where L is the data lineage path complexity, t is time, e is an edge, e∈E represents the set of edges in path e, w(e,t) is the weight of edge e at time t, and f(v) is the complexity function associated with node v.
[0057] Specifically, data lineage path complexity is a quantitative metric used to assess the complexity of dependency paths of target data in a data lineage graph. It reflects the difficulty of data flow and dependencies by analyzing the edges, nodes, their weights, and complexity functions within the path. This complexity helps identify and optimize complex paths in data governance, improving the efficiency and accuracy of data processing. Data lineage path complexity measures the overall complexity of dependency paths of target data. Time is used to handle dynamic scenarios, reflecting how complexity changes over time. The weight of edge e in the path at time t represents the importance of a dependency relationship. For example, some edges may have higher weights due to higher frequency of use or greater importance. f(v) is the complexity function associated with node v, representing the complexity of the node itself. For example, the number of fields contained in the data node, the difficulty of data operations (such as aggregation and cleaning) involved in the node, and the computational resource consumption of the node.
[0058] The server first extracts the paths that the target data depends on from the data lineage graph, including edges e and nodes v. The server calculates the weight w(e,t) for each edge e in the path. The weight calculation may be based on factors such as the edge's access frequency (more frequent edges have higher weights) and the data volume (edges involving large data flows have higher weights). Then, the server evaluates the complexity f(v) of each node, possibly considering the number of fields in the node; more fields indicate higher complexity. This may also consider whether the node contains complex aggregation operations. Finally, the server combines the weights on the path and the node complexity to calculate the path complexity L according to the formula.
[0059] In one possible implementation, the global impact score of the data is calculated using the following formula:
[0060] Where I represents the global impact score of the data, and V d For the set of nodes directly affected, α v Let σ(v) be the weight coefficient of node v, σ(v) be the sensitivity score of node v, K be the maximum propagation level, k be the k-th level, and β be the weight coefficient of node v. k I represents the weight of level k. k The score represents the influence of level k.
[0061] Specifically, the Global Data Impact Score is a quantitative indicator used to measure the degree of influence of target data on other data entities throughout the data lineage graph. This indicator comprehensively evaluates the impact from both direct influence and multi-level propagation perspectives, focusing on analyzing the set of nodes directly associated with the target data (directly affected nodes) and the combined impact of each level as the influence propagates along the path. The Global Data Impact Score measures the global influence of the target data throughout the entire data lineage graph. The set of nodes directly affected by the target data includes entities directly associated with it. The weight coefficient of node v indicates the importance of that node; for example, core business nodes may have higher weights, while non-critical field nodes may have lower weights. The sensitivity score of node v assesses the node's sensitivity to changes in data quality; for example, nodes with a greater impact on business due to data errors are more sensitive. K is the maximum propagation level, defining the depth of influence propagation, which depends on the complexity of the business scenario. The weight of level k reflects the importance of propagation at that level. For example, levels farther from the target data may have lower weights. The influence score of level k represents the combined impact of all nodes at level k.
[0062] The server first identifies all directly related nodes to the target data. These nodes are directly affected by changes in the target data. Then, the server calculates the influence of each directly affected node; weight and sensitivity jointly determine the node's importance. Next, the server calculates the influence score for each level of propagation, with the score for each level being the sum of the influence of all nodes at that level. Hierarchical weights are used to adjust the contribution of each level of propagation. Finally, the server performs a weighted sum of the direct influence and the propagated influence to obtain the global influence score I.
[0063] In one possible implementation, the data leakage risk score is calculated using the following formula:
[0064] Where R is the data leakage risk score, v is the node, and V is the data leakage risk score. s Let be the set of nodes containing sensitive data, π(v) be the probability that node v is accessed, and r(v) be the leakage risk score of node v.
[0065] Specifically, the data leakage risk score is a metric used to assess the leakage risk of target data or its related nodes in a data lineage graph. This score comprehensively considers factors such as the set of sensitive data nodes in the data lineage graph, the probability of each node being accessed reflecting its exposure level, and the inherent leakage risk score of each node representing the impact of data characteristics on the likelihood of leakage. By weighted summing of these factors, the overall risk of sensitive data leakage in the entire data lineage graph is derived. The data leakage risk score measures the overall likelihood of leakage of the target data. Nodes in the data lineage graph represent data entities (such as tables, fields, or files), containing a set of nodes with sensitive data, such as fields involving personal privacy like user_id and email_address; tables involving trade secrets like financial_records; and the probability of node v being accessed. The access probability can be estimated based on logs, permission assignments, or historical records. The leakage risk score of node v depends on the data type, storage security, and its potential impact in a leakage scenario. Essentially, the formula is the sum of the products of the access probability and leakage risk of all nodes in the sensitive node set.
[0066] The server first identifies a set of sensitive data nodes by extracting sensitive information from metadata, such as marking nodes containing fields containing identity information or financial information. The server then categorizes the data according to business rules or sensitivity levels. Based on access logs, the server calculates the probability of nodes being exposed to users by analyzing historical access frequencies and considering user permission distribution. Next, the server comprehensively considers factors such as data storage methods, encryption strategies, and security vulnerabilities. For example, unencrypted fields have higher leakage risk scores, while encrypted fields or fields with lower sensitivity levels have lower risk scores. Finally, the server sums the products of the access probability and leakage risk score for each node in the sensitive node set to obtain the overall data leakage risk score R.
[0067] In one possible implementation, the data usage efficiency score is calculated using the following formula:
[0068] Where M is the data usage efficiency score, T is the maximum time, t is the time, U(t) is the number of data accesses at time t, and C(t) is the data maintenance cost at time t.
[0069] Specifically, the data utilization efficiency score is an indicator used to measure the actual utilization efficiency of target data at different points in time. It assesses the value of data use and the rationality of resource investment by calculating the ratio of data access frequency to maintenance cost and combining it with the time dimension. The data utilization efficiency score represents the overall utilization efficiency of the target data. The maximum time point refers to the last point in time within the data lifecycle or monitoring period. Time t is used to break down data usage. The number of data accesses at time t reflects the frequency of data use. The data maintenance cost at time t represents the resource consumption of the data at that point in time, such as storage, backup, or computing resources. This formula measures whether the data has been utilized efficiently throughout its lifecycle by dividing the number of accesses by the maintenance cost over time and then summing the results.
[0070] The process begins with the server collecting data access logs, extracting access records from the database, file system, or data lake, and calculating the access frequency by time. For example, if a table is accessed 10 times per day, then U(t) = 10. Next, the server analyzes the consumption of data storage and operational resources, such as hardware, network bandwidth, and computing power. For example, if the data storage cost is 5 yuan per GB per day, the cost can be calculated based on the storage volume: C(t) = 5 × storage volume. Then, at each time point t, the server divides the number of accesses by the maintenance cost to determine the data usage efficiency at that moment. Finally, the server sums up the efficiency at all individual moments to obtain the overall data usage efficiency score M.
[0071] S140. Based on the data quality propagation score, data lineage path complexity, data global impact score, data leakage risk score, and data usage efficiency score, a comprehensive data governance score is generated.
[0072] Specifically, the comprehensive data governance score is the result of a multi-dimensional quantitative evaluation of the target data, taking into account five dimensions: data quality, data complexity, global data impact, leakage risk, and usage efficiency. The server calculates a weighted sum and difference of these scores to generate a unified indicator, which serves as the basis for evaluation and decision-making.
[0073] In one possible implementation, refer to another flowchart of a data governance method based on data lineage in Figure 2. A comprehensive data governance score is generated based on data quality propagation score, data lineage path complexity, global data impact score, data leakage risk score, and data usage efficiency score. Specifically, this includes: S210, determining the ratio between the data quality propagation score and the data lineage path complexity; S220, calculating the first product between the ratio and a first weight; S230, calculating the second product between the global data impact score and a second weight; S240, calculating the third product between the data leakage risk score and a third weight; S250, calculating the fourth product between the data usage efficiency score and a fourth weight; and S260, summing the first, second, and fourth product results and subtracting the third product result to obtain the comprehensive data governance score.
[0074] Specifically, the server calculates the ratio of data quality propagation score to data lineage path complexity, reflecting the level of data quality relative to data complexity. The aim is to identify whether data maintains high quality even within a complex lineage structure. The first weight is the weight of the data quality-to-complexity ratio, indicating the importance of this ratio. The second weight is the weight of the global impact score, reflecting the degree of importance placed on the data's global impact on other parts of the system. The third weight is the weight of the data leakage risk score, used to measure the security risk of sensitive data. The fourth weight is the weight of the data usage efficiency score, indicating the importance placed on the actual value of the data. These weights can be customized according to the target data governance needs. The server sums the weighted products of the above calculations, aggregating the contribution of each dimension to the overall score. Simultaneously, the product of the leakage risk weight is subtracted, as data leakage risk is a negative impact and should be deducted. Finally, a quantitative overall score is obtained; the higher the value, the higher the priority of the data in governance.
[0075] For example, assuming a comprehensive score of 39.6, which is higher than the preset scoring threshold, indicates that this customer data table has a high priority in data governance. Reasons include: high data quality propagation and relatively controllable path complexity; significant global impact and strong correlation with other parts of the system; high usage efficiency, high access frequency, and controllable maintenance costs. While the leakage risk score presents some potential risks, it still falls under the category of high-priority data overall. Therefore, by comprehensively considering multi-dimensional data performance, the one-sidedness of a single indicator is avoided. Flexible weight configuration can balance the governance goals of different enterprises, such as emphasizing quality, efficiency, or security. It supports complex data environments, such as multi-table dependencies, high access frequency, or high-risk data. It eliminates subjective interference, making the data governance decision-making process more scientific and efficient. Specifically, it prioritizes the governance of high-impact data, such as core business-related data. It quickly identifies high-risk data, such as by planning resources in advance for enhanced security. It improves data governance efficiency, such as directly locating high-value data through scoring, thereby enhancing governance effectiveness.
[0076] S150. Perform data governance on the target data according to the comprehensive data governance score.
[0077] Specifically, based on the previously calculated comprehensive data governance score, the server allocates governance resources, determines governance strategies, or adjusts governance priorities for target data according to the score. The goals of data governance include: improving data quality, such as fixing errors, removing duplicate data, and completing missing data; optimizing data structure, such as optimizing data lineage paths and reducing complexity; enhancing data security, such as strengthening the protection of sensitive data and reducing the risk of data leakage; and improving efficiency, such as reducing unnecessary data storage or redundancy and improving access efficiency.
[0078] The server prioritizes high-priority data by ranking all data based on a comprehensive score, prioritizing data with higher scores for governance. High-priority data may be allocated more manpower, computing resources, and time. Low-priority data may be managed using automated batch governance tools to reduce costs. The server implements targeted governance measures based on data characteristics and score composition. For example, if a piece of data has a low score primarily due to a low quality propagation score, the focus is on fixing errors within the data. Similarly, if a low score is due to a low data leakage risk score, access control for that data is strengthened. The score can be dynamically adjusted during the governance process. If the governance measures are effective, the score will improve, and the server can update its strategies accordingly.
[0079] This application also provides a data governance device based on data lineage, as shown in Figure 3, which is a schematic diagram of the modules of the data governance device based on data lineage provided in this application. The data governance device is a server, which includes an acquisition module 31 and a processing module 32. The acquisition module 31 acquires the metadata of the target data corresponding to the target governance requirements; the processing module 32 constructs a data lineage graph of the target data based on the metadata; the processing module 32 calculates the data quality propagation score, data lineage path complexity, data global impact score, data leakage risk score, and data usage efficiency score of the target data through the data lineage graph; the processing module 32 generates a comprehensive data governance score based on the data quality propagation score, data lineage path complexity, data global impact score, data leakage risk score, and data usage efficiency score; and the processing module 32 performs data governance on the target data according to the comprehensive data governance score.
[0080] In one possible implementation, the processing module 32 constructs a data lineage graph of the target data based on the metadata, specifically including: the processing module 32 parses the metadata to obtain data entities, which include database tables, fields, and files; the processing module 32 extracts edges from data entities with dependencies among multiple data entities; the acquisition module 31 acquires additional data between data entities and edges; the processing module 32 determines attributes based on the additional data; and the processing module 32 constructs a data lineage graph of the target data based on the data entities, edges, and attributes.
[0081] In one possible implementation, the processing module 32 calculates the data quality propagation score, data lineage path complexity, global data impact score, data leakage risk score, and data usage efficiency score of the target data through the data lineage graph. The data quality propagation score is specifically calculated using the following formula:
[0082] Where Q is the data quality propagation score, i is the i-th metadata, and q i Let be the quality score of the i-th metadata, n be the total number of metadata, e be an edge, e∈E represent the set of edges in path e, and d(e) be the quality decay rate of edge e in path e.
[0083] In one possible implementation, the processing module 32 calculates the data quality propagation score, data lineage path complexity, global data impact score, data leakage risk score, and data usage efficiency score of the target data through the data lineage graph. The data lineage path complexity is specifically calculated using the following formula:
[0084] Where L is the data lineage path complexity, t is time, e is an edge, e∈E represents the set of edges in path e, w(e,t) is the weight of edge e at time t, and f(v) is the complexity function associated with node v.
[0085] In one possible implementation, the processing module 32 calculates the data quality propagation score, data lineage path complexity, global data impact score, data leakage risk score, and data usage efficiency score of the target data through the data lineage graph. The global data impact score is specifically calculated using the following formula:
[0086] Where I represents the global impact score of the data, and V d For the set of nodes directly affected, α v Let σ(v) be the weight coefficient of node v, σ(v) be the sensitivity score of node v, K be the maximum propagation level, k be the k-th level, and β be the weight coefficient of node v. k I represents the weight of level k. k The score represents the influence of level k.
[0087] In one possible implementation, the processing module 32 calculates the data quality propagation score, data lineage path complexity, global data impact score, data leakage risk score, and data usage efficiency score of the target data through the data lineage graph. The data leakage risk score is specifically calculated using the following formula:
[0088] Where R is the data leakage risk score, v is the node, and V is the data leakage risk score. s Let be the set of nodes containing sensitive data, π(v) be the probability that node v is accessed, and r(v) be the leakage risk score of node v.
[0089] In one possible implementation, the processing module 32 calculates the data quality propagation score, data lineage path complexity, global data impact score, data leakage risk score, and data utilization efficiency score of the target data through the data lineage graph. The data utilization efficiency score is specifically calculated using the following formula:
[0090] Where M is the data usage efficiency score, T is the maximum time, t is the time, U(t) is the number of data accesses at time t, and C(t) is the data maintenance cost at time t.
[0091] In one possible implementation, the processing module 32 generates a comprehensive data governance score based on the data quality propagation score, data lineage path complexity, data global impact score, data leakage risk score, and data usage efficiency score. Specifically, the processing module 32 determines the ratio between the data quality propagation score and the data lineage path complexity; the processing module 32 calculates the first product result between the ratio and a first weight; the processing module 32 calculates the second product result between the data global impact score and a second weight; the processing module 32 calculates the third product result between the data leakage risk score and a third weight; the processing module 32 calculates the fourth product result between the data usage efficiency score and a fourth weight; the processing module 32 sums the first, second, and fourth product results and subtracts the third product result to obtain the comprehensive data governance score.
[0092] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0093] This application also provides an electronic device, as shown in Figure 4, which is a structural schematic diagram of an electronic device according to an embodiment of this application. The electronic device may include: at least one processor 41, at least one network interface 44, a user interface 43, a memory 45, and at least one communication bus 42.
[0094] The communication bus 42 is used to enable communication between these components.
[0095] User interface 43 may include a display screen and a camera. Optionally, user interface 43 may also include a standard wired interface and a wireless interface.
[0096] Network interface 44 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0097] Processor 41 may include one or more processing cores. Processor 41 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 45, and by calling data stored in memory 45. Optionally, processor 41 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 41 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 41 and may be implemented as a separate chip.
[0098] The memory 45 may include random access memory (RAM) or read-only memory. Optionally, the memory 45 may include a non-transitory computer-readable storage medium. The memory 45 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 45 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 45 may also be at least one storage device located remotely from the aforementioned processor 41. As shown in FIG4, the memory 45, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program of a data governance method based on data lineage.
[0099] In the electronic device shown in Figure 4, the user interface 43 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 41 can be used to call the application program based on data lineage data governance method stored in the memory 45. When executed by one or more processors, the electronic device executes one or more methods as described in the above embodiments.
[0100] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0101] This application also provides a computer-readable storage medium storing instructions. When executed by one or more processors, these instructions cause an electronic device to perform one or more of the methods described in the above embodiments.
[0102] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0103] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0104] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0105] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0106] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0107] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A data governance method based on data lineage, characterized in that, The method includes: Obtain metadata of the target data corresponding to the target governance requirements; Based on the metadata, construct a data lineage diagram of the target data; The data lineage graph is used to calculate the data quality propagation score, data lineage path complexity, global data impact score, data leakage risk score, and data usage efficiency score of the target data. A comprehensive data governance score is generated based on the data quality propagation score, the data lineage path complexity, the data global impact score, the data leakage risk score, and the data usage efficiency score. Data governance is performed on the target data based on the comprehensive data governance score.
2. The data governance method based on data lineage according to claim 1, characterized in that, The step of constructing the data lineage diagram of the target data based on the metadata includes: The metadata is parsed to obtain data entities, which include database tables, fields, and files; Edges are extracted from data entities that have dependencies among the multiple data entities; Obtain additional data between the data entity and the edge; Based on the additional data, determine the attributes; Based on the data entities, the edges, and the attributes, a data lineage graph of the target data is constructed.
3. The data governance method based on data lineage according to claim 2, characterized in that, The dependency is defined by at least one of foreign key relationships, ETL process logs, or application operations; the additional data includes at least one of edge weights, timestamps, operation types, data transmission frequency, access permissions, and sensitivity levels; the attributes are used to characterize the features of nodes and edges, including node sensitivity and edge transmission efficiency.
4. The data governance method based on data lineage according to claim 1, characterized in that, The data lineage diagram for constructing the target data also includes: The data entity is used as a node; Treat the dependencies as edges and indicate the direction from the source entity to the target entity; Using the aforementioned attributes as features of a graph, a directed graph-like data lineage graph is generated, which represents the entire path of data from source to destination.
5. The data governance method based on data lineage according to claim 1, characterized in that, The data quality propagation score is calculated using the following formula: Where Q is the data quality propagation score, i is the i-th metadata, and q i Let be the quality score of the i-th metadata, n be the total number of metadata, e be an edge, e∈E represent the set of edges in path e, and d(e) be the quality decay rate of path edge e.
6. The data governance method based on data lineage according to claim 5, characterized in that, In the calculation of the data quality propagation score, the quality score of the i-th metadata is evaluated based on at least one of field integrity, data accuracy, and data consistency. The field integrity evaluates whether the data has missing values, the data accuracy evaluates whether the data conforms to business logic, and the data consistency evaluates whether the data matches across systems.
7. The data governance method based on data lineage according to claim 1, characterized in that, The complexity of the data lineage path is calculated using the following formula: Where L is the data lineage path complexity, t is time, e is an edge, e∈E represents the set of edges in path e, w(e,t) is the weight of edge e at time t, and f(v) is the complexity function associated with node v.
8. The data governance method based on data lineage according to claim 7, characterized in that, In the calculation of the data lineage path complexity, the weight w(e,t) is dynamically adjusted based on the edge access frequency and data size, and the complexity function f(v) is determined based on the number of node fields and the difficulty of operation.
9. The data governance method based on data lineage according to claim 1, characterized in that, The global impact score of the data is calculated using the following formula: Where I represents the global impact score of the data, and V d For the set of nodes directly affected, α v Let σ(v) be the weight coefficient of node v, σ(v) be the sensitivity score of node v, K be the maximum propagation level, k be the k-th level, and β be the weight coefficient of node v. k I represents the weight of level k. k The score represents the impact of level k.
10. The data governance method based on data lineage according to claim 1, characterized in that, In the calculation of the global impact score of the data, the weight coefficient of node v represents the importance of the node, with the weight of core business nodes being higher than that of non-critical field nodes; the sensitivity score of node v is evaluated based on the degree of impact of data errors on the business, with nodes whose data errors have a greater impact on the business having a higher sensitivity score; the weight of the level k reflects the propagation importance of the level, with levels farther away from the target data having a lower weight.
11. The data governance method based on data lineage according to claim 1, characterized in that, The data leakage risk score is calculated using the following formula: Where R is the data leakage risk score, v is the node, and V is the data leakage risk score. s Let be the set of nodes containing sensitive data, π(v) be the probability that node v is accessed, and r(v) be the leakage risk score of node v.
12. The data governance method based on data lineage according to claim 1, characterized in that, The data usage efficiency score is calculated using the following formula: Where M is the data usage efficiency score, T is the maximum time, t is the time, U(t) is the number of data accesses at time t, and C(t) is the data maintenance cost at time t.
13. The data governance method based on data lineage according to claim 1, characterized in that, The comprehensive data governance score is generated based on the data quality propagation score, the data lineage path complexity, the data global impact score, the data leakage risk score, and the data usage efficiency score, including: Determine the ratio between the data quality propagation score and the data lineage path complexity; Calculate the first product between the ratio and the first weight; Calculate the second product between the global impact score of the data and the second weight; Calculate the third product between the data leakage risk score and the third weight; Calculate the fourth product between the data usage efficiency score and the fourth weight; The data governance comprehensive score is obtained by summing the first product result, the second product result, and the fourth product result, and then subtracting the third product result.
14. The data governance method based on data lineage according to claim 13, characterized in that, The method further includes: The first weight, the second weight, the third weight, and the fourth weight are adjusted according to changes in data governance needs or dynamic adjustments to indicators. The overall data governance score is recalculated based on the adjusted weights to adapt to governance objectives at different stages.
15. A data governance device based on data lineage, characterized in that, The data governance device includes an acquisition module (31) and a processing module (32), wherein, The acquisition module (31) is used to acquire metadata of the target data corresponding to the target governance requirements; The processing module (32) is used to construct a data lineage diagram of the target data based on the metadata; The processing module (32) is also used to calculate the data quality propagation score, data lineage path complexity, data global impact score, data leakage risk score and data usage efficiency score of the target data through the data lineage graph. The processing module (32) is also used to generate a comprehensive data governance score based on the data quality propagation score, the data lineage path complexity, the data global impact score, the data leakage risk score, and the data usage efficiency score. The processing module (32) is also used to perform data governance on the target data according to the data governance comprehensive score.
16. The data governance device based on data lineage according to claim 15, characterized in that, The processing module (32) uses the following formula to calculate the data quality propagation score: Where Q is the data quality propagation score, i is the i-th metadata, and q i Let be the quality score of the i-th metadata, n be the total number of metadata, e be an edge, e∈E represent the set of edges in path e, and d(e) be the quality decay rate of path edge e.
17. The data governance device based on data lineage according to claim 15, characterized in that, The processing module (32) uses the following formula to calculate the complexity of the data lineage path: Where L is the data lineage path complexity, t is time, e is an edge, e∈E represents the set of edges in path e, w(e,t) is the weight of edge e at time t, and f(v) is the complexity function associated with node v.
18. The data governance device based on data lineage according to claim 15, characterized in that, The processing module (32) uses the following formula when calculating the global impact score of the data: Where I represents the global impact score of the data, and V d For the set of nodes directly affected, α v Let σ(v) be the weight coefficient of node v, σ(v) be the sensitivity score of node v, K be the maximum propagation level, k be the k-th level, and β be the weight coefficient of node v. k I represents the weight of level k. k The score represents the impact of level k.
19. The data governance device based on data lineage according to claim 15, characterized in that, The processing module (32) uses the following formula to calculate the data usage efficiency score: Where M is the data usage efficiency score, T is the maximum time, t is the time, U(t) is the number of data accesses at time t, and C(t) is the data maintenance cost at time t.
20. An electronic device, characterized in that, The electronic device includes a processor (41), a memory (45), a user interface (43), and a network interface (44). The memory (45) is used to store instructions. The user interface (43) and the network interface (44) are both used to communicate with other devices. The processor (41) is used to execute the instructions stored in the memory (45) to cause the electronic device to perform the method as described in any one of claims 1 to 14.