Data blood relationship tracking method and device, computer equipment and storage medium
By building blood chains of data sources, ETL tasks and indicator nodes in the graph database, the maintenance problem of traditional data blood tracing technology when the data table structure changes is solved, visualization and full-link tracking of data blood tracing are realized, and the transparency and collaboration efficiency of data management are improved.
Patent Information
- Application Number
- CN202510401857.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional data blood tracing technology is difficult to effectively maintain blood tracing charts when facing changes in the data table structure, resulting in blind spots in data management and difficulty in collaboration.
The graph database is used to build data source nodes, ETL task nodes and indicator nodes. By obtaining data source, ETL task and indicator system metadata, a data blood chain is built, and the blood chain is displayed in the graph database to realize full-link blood tracking and task monitoring.
It realizes the visual display of data blood ties, reduces the workload of manual sorting, improves the transparency and collaboration efficiency of data management, promotes data sharing and collaboration, and improves the level of enterprise data management.
Smart Images

Figure CN120256506A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of database technologies, and in particular, to a data lineage tracing method, apparatus, computer device, and storage medium. Background Art
[0002] In the field of data processing, data lineage tracing technology, as a core tool for ensuring data quality and compliance, is mainly used to trace the complete flow path of data from the original source to business metrics. Traditional methods usually rely on static metadata collection, and build table-level lineage relationships by parsing database table structures or ETL task logs. However, they have significant limitations in practical applications. Traditional technologies often have difficulty maintaining the lineage graph well when faced with the situation where the database table structures often change. Summary of the Invention
[0003] The purpose of this application aims to solve at least one of the above technical defects, especially the problem that the lineage graph cannot be well maintained in the prior art.
[0004] In a first aspect, this application provides a data lineage tracing method, including:
[0005] Obtain data source metadata, ETL task metadata, and metric system metadata respectively;
[0006] In a graph database, construct a data source node according to the data source metadata, an ETL task node according to the ETL task metadata, and a metric node according to the metric system metadata, and determine the node association relationship according to the data source metadata, the ETL task metadata, and the metric system metadata;
[0007] Construct a data lineage chain with the data source node as the starting point and the metric node as the ending point according to the node association relationship;
[0008] Display the data lineage chain through the graph database.
[0009] In one embodiment, the obtaining of the data source metadata includes:
[0010] Send a metadata acquisition request to the target data source to obtain the data source metadata.
[0011] In one embodiment, the obtaining of the data source metadata further includes:
[0012] Receive a data definition language change stream through a streaming processing task;
[0013] Determine the change target according to the received data definition language change statement;
[0014] Obtain the changed data source metadata according to the change target;
[0015] Update the data source metadata according to the changed data source metadata.
[0016] In one embodiment, the acquisition of ETL task metadata includes:
[0017] Extract the ETL task configuration table from the scheduling system;
[0018] Obtain the ETL task metadata according to the ETL task configuration table.
[0019] In one embodiment, the acquisition of the metric system metadata includes:
[0020] Obtain the metric system of the target data warehouse, and obtain the metric system metadata according to the definitions of the metrics in the metric system.
[0021] In one embodiment, the data lineage tracing method further includes:
[0022] Monitor the execution status of the ETL task nodes;
[0023] In the case where the execution status is failed, restart the ETL task nodes in response to the user's rerun instruction.
[0024] In one embodiment, the data lineage tracing method further includes:
[0025] For any data lineage chain, determine whether there is any other path between the data source node and the metric node of the data lineage chain;
[0026] If so, count the average calculation time of the other path and the current path;
[0027] Select the one with the minimum average calculation time as the target path;
[0028] Send a path optimization prompt to the user according to the target path.
[0029] In a second aspect, the present application provides a data lineage tracing device, including:
[0030] A data acquisition module, configured to acquire data source metadata, ETL task metadata, and metric system metadata respectively;
[0031] A node construction module, configured to construct a data source node according to the data source metadata, an ETL task node according to the ETL task metadata, and a metric node according to the metric system metadata in the graph database respectively, and determine the node association relationship according to the data source metadata, the ETL task metadata, and the metric system metadata;
[0032] A lineage chain construction module, configured to construct a data lineage chain starting from the data source node and ending at the metric node according to the node association relationship;
[0033] A display module for displaying the data lineage chain through a graph database.
[0034] In a third aspect, the present application provides a computer device, including one or more processors and a memory. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the one or more processors, the steps of the data lineage tracing method in any of the above embodiments are executed.
[0035] In a fourth aspect, the present application provides a storage medium in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the data lineage tracing method in any of the above embodiments.
[0036] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0037] The data lineage tracing system based on a graph database proposed in the present application integrates data warehouse tables, ETL scheduling tasks, and metric systems to achieve full-link lineage tracing and task monitoring and operation and maintenance. By globally graphically displaying the entire data processing lineage, it intuitively presents the whole process of data from the data source through ETL tasks to the generation of metrics, greatly reducing the workload of manually sorting out data relationships. It also covers data warehouse table metadata, ETL task metadata, and data metric layer metadata, and the data lineage can change with the change of metadata. This enables users to examine the data processing and analysis process from an all-round perspective, clearly understand the source, processing logic, and destination of the data, providing a complete information basis for data governance and avoiding blind spots in data management. The integration of multiple types of metadata is another prominent advantage of this method. It integrates the metadata of data warehouse tables, ETL scheduling tasks, and metric systems to form a one-stop data management and monitoring platform. Personnel in different departments, such as data analysts, developers, and business decision-makers, can obtain the data information they need on this platform, promoting data sharing and collaboration and improving the overall data management level of the enterprise. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0039] Figure 1 It is a schematic flowchart of the data lineage tracing method provided for an embodiment of the present application;
[0040] Figure 2 Schematic diagram of the process for updating data source metadata in an embodiment of the present application;
[0041] Figure 3 Schematic diagram of the process for optimizing the data processing path in an embodiment of the present application;
[0042] Figure 4 Internal structure diagram of a computer device provided in an embodiment of the present application. Detailed implementation manners
[0043] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. The embodiments described in the specification are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0044] The embodiment of the present application provides a data lineage tracking method, including steps S102 to S108.
[0045] S102, respectively obtain data source metadata, ETL task metadata, and metric system metadata.
[0046] It can be understood that data source metadata is data that describes the basic characteristics and attributes of a data source, including metadata of data tables and field metadata within a data source (i.e., a certain database), such as field names, field types, index information, etc., and can also include information such as the type of the database and the update frequency. When extracting data source metadata, it can also be normalized to ensure accuracy and consistency. ETL task metadata is a description of the attributes related to an ETL (Extract, Transform, Load, i.e., data extraction, transformation, and loading) task, covering task names, task dependencies (i.e., other tasks that need to be completed before this task is executed), scheduling frequencies (such as executing at 2 am every day, 10 am every Monday, etc.), execution status (success, failure, running, waiting to execute, etc.). The metadata of the indicator system is a collection of information about the definition, calculation logic, source, etc. of indicators, including indicator names, source tables, calculation formulas, update frequencies, business restrictions (such as indicator calculation rules under specific conditions), etc. Obtaining data source metadata provides basic information for subsequent construction of data source nodes, enabling the system to identify and connect to data sources and understand their data structures, preparing for data processing and lineage tracking. Obtaining ETL task metadata helps to master the execution plans and status of each task in the data processing process, as well as the dependencies between tasks, thus realizing effective scheduling and monitoring of tasks. Obtaining the metadata of the indicator system can clarify the definition and calculation method of indicators, facilitating the display of the source and calculation basis of indicators in data lineage tracking and meeting the business's analysis and monitoring requirements for data indicators. The data lineage relationship in this embodiment is jointly constructed from three perspectives: the data source, ETL tasks, and the final indicator system of the data warehouse, and all are obtained based on the metadata of these three aspects. When they change, they will all be reflected in the metadata, so that the data lineage can be updated synchronously.
[0047] S104. In the graph database, construct data source nodes according to data source metadata, construct ETL task nodes according to ETL task metadata, construct indicator nodes according to the metadata of the indicator system, and determine the node association relationship according to the data source metadata, ETL task metadata, and the metadata of the indicator system.
[0048] It can be understood that a graph database is a database that stores and queries data in a graph structure, with nodes and edges being its core components. Data source nodes are nodes created in the graph database based on data source metadata, representing data sources, and node attributes contain information related to data source metadata. ETL task nodes are nodes constructed based on ETL task metadata, used to represent ETL tasks, and node attributes record ETL task metadata. Metric nodes are nodes constructed based on metric system metadata, reflecting metric information, and node attributes include metric definitions, calculation logics, etc. Node association relationships are determined based on the internal connections between data sources, ETL tasks, and metric systems. For example, a data source is the data input source of an ETL task, and the output data of an ETL task can be used to calculate metrics or be the input of other ETL tasks. These relationships are represented by edges in the graph database. Constructing nodes and determining association relationships are key steps in implementing data lineage tracing. Data source nodes, as the starting points of data, provide the data foundation for the entire data lineage graph. ETL task nodes record the data processing process and are connected to data source nodes through association relationships, showing the data flow and processing logic. Metric nodes, as the results of data processing, are associated with ETL task nodes and data source nodes, clarifying the source and calculation path of metric data. In this way, the graph structure constructed in the graph database can intuitively display the entire flow process of data from the data source to the metric, facilitating data lineage analysis and monitoring.
[0049] S106. Construct a data lineage chain with the data source node as the starting point and the metric node as the ending point according to the node association relationship.
[0050] It can be understood that a data lineage chain is a path in the graph database that starts from a data source node, passes through a series of ETL task nodes and / or other metric nodes, and finally reaches a metric node. It intuitively displays the processing flow and evolution process of data from the original data source to the final metric data. Metric nodes can be specifically divided into atomic metrics, derived metrics, and application metrics. Derived metrics are calculated using atomic metrics, and application metrics are calculated using derived metrics. In the graph database, node association relationships determine the connection methods and data flows between nodes. By traversing these association relationships, starting from the data source node, following the logical order of data processing, and gradually advancing along the associated edges, the association path with the metric node is finally found, thus constructing a complete data lineage chain.
[0051] Building a data lineage chain is the core function for realizing data lineage tracking. It organically connects data sources, ETL tasks, and the metric system, enabling users to clearly understand the context of data. Through the data lineage chain, the source of data problems can be quickly located. For example, when abnormal metric data occurs, it is possible to trace back along the lineage chain to the data source or related ETL tasks to find the problem, which helps with data quality management and troubleshooting. At the same time, in data compliance audits, the data lineage chain can clearly show whether the data processing process complies with regulations and meets compliance requirements.
[0052] S108, display the data lineage chain through a graph database.
[0053] It can be understood that the graph database provides rich visualization tools and interfaces to display the constructed data lineage chain in an intuitive graphical way. In the visual display, data source nodes, ETL task nodes, and metric nodes are usually represented by different graphical elements (such as circles, squares, diamonds, etc.), and the association relationships between nodes are represented by lines, and the direction of the lines indicates the data flow direction. Through this visual way, users can clearly see the whole process of data from the data source through ETL processing and finally generating metric data at a glance.
[0054] Displaying the data lineage chain through the graph database greatly improves the visualization and comprehensibility of data lineage tracking. For data analysts and managers, the intuitive graphical display makes data lineage analysis more convenient, and they can quickly understand the data source and processing process without complex queries and analysis. In the data governance process, the visual data lineage chain helps to discover data quality problems and potential data risks, and make timely optimizations and adjustments.
[0055] The data lineage tracking system based on the graph database proposed in this application integrates data warehouse tables, ETL scheduling tasks, and the metric system to achieve full-link lineage tracking and task monitoring and operation and maintenance. By globally graphically displaying the entire data processing lineage, it intuitively presents the whole process of data from the data source through ETL tasks to the generation of metrics, greatly reducing the workload of manually sorting out data relationships. It also covers the metadata of data warehouse tables, ETL task metadata, and data metric layer metadata, and the data lineage can change with the change of metadata. This enables users to examine the data processing and analysis process from an all-round perspective, clearly understand the source, processing logic, and destination of data, providing a complete information basis for data governance and avoiding blind spots in data management. The integration of multiple types of metadata is another prominent advantage of this method. It integrates the metadata of data warehouse tables, ETL scheduling tasks, and the metric system to form a one-stop data management and monitoring platform. Personnel from different departments, such as data analysts, developers, and business decision-makers, can obtain the data information they need on this platform, promoting data sharing and collaboration and improving the overall data management level of the enterprise.
[0056] In one embodiment, the acquisition of data source metadata includes: sending a metadata acquisition request to a target data source to obtain data source metadata. Sending a metadata acquisition request is a process of interacting with a data source based on a network communication protocol. In a network environment, a client (i.e., a data lineage tracking system) constructs a request message based on a specific communication protocol (such as HTTP, JDBC, etc. in the TCP / IP protocol family), and the message contains key information such as the request content and the target data source. When the request is sent to the target data source, the data source receives the request and parses the instructions therein, retrieves relevant information from the system directory, data dictionary, and other areas storing metadata according to its own data storage structure and management mechanism, and then encapsulates these data source metadata in a response message in a prescribed format, and then returns it to the client through the network. For example, if the target data source is a relational database MySQL, the system may send an SQL query statement through the JDBC protocol to obtain metadata information such as table structure and field type; if the target data source is a file system, it may send a corresponding request instruction based on the interface specification of the file system to obtain file metadata, such as file size, creation time, file type, etc. The target data source may include dimension tables and fact tables. Dimensions are perspectives used to observe and analyze business data. They support data aggregation, drilling, and slicing analysis, and can be applied to the GROUP BY condition in SQL. Most dimensions have a hierarchical structure, such as: geographic dimensions (including content at the country, region, province, and city levels), time dimensions (including content at the year, quarter, and month levels). The metadata associated with the dimension table includes the fields, field types, and lengths contained in the table. Fact tables store quantitative information about specific topics and are used to analyze numerical measurements of business processes. The metadata associated with fact tables includes the fields, field types, and lengths contained in the table. Fact tables can be associated with required dimension tables.
[0057] Before obtaining the metadata of the data source, you can configure the connection method of the data source in the system. The system will try to establish a connection with the data source. When the user issues a connection test instruction, the system will try to establish a connection with the target data source based on the data source information entered by the user, such as the database connection URL, user name, password, etc. This process uses network communication protocols (such as JDBC, ODBC, etc.) to send the connection request to the target data source. If the data source receives the request and verifies that the user information is correct, it will return a successful response, indicating that the connection is successful; otherwise, if the user information is incorrect, the data source is unreachable, or there is a network problem, the data source will return an error message and the connection fails.
[0058] After the data source is successfully connected, the system will send specific query requests to the data source. These requests vary according to the type of data source. For example, for a relational database, SQL query statements (such as "SHOW TABLES", "DESCRIBE TABLE", etc.) will be sent to obtain metadata such as table structure and field information; for a file system, file metadata (such as file size, creation time, etc.) will be obtained according to the file system's API. After receiving the query request, the data source will retrieve relevant information from its own metadata storage area and return it to the system, and then the system will store this metadata on the platform.
[0059] In one embodiment, please refer to Figure 2 , the acquisition of data source metadata further includes the step of updating the data source metadata, specifically including steps S202 to S208.
[0060] S202, receive the data definition language change stream through a streaming processing task.
[0061] It can be understood that the data definition language (DDL) is used to define the structure of a database, such as creating, modifying, and deleting database objects (tables, views, indexes, etc.). The change stream refers to the real-time data stream of these DDL statements, which reflects the dynamic changes in the database structure. The streaming processing task can continuously receive and process these real-time data streams to ensure that the system can timely perceive the structural changes of the data source. In principle, the streaming processing task is generally implemented based on a message queue (such as Kafka) or a real-time data platform (such as Flink). When the database executes DDL statements, it will send these change messages to the message queue, and the streaming processing task will consume this information from the message queue. For example, when a database administrator executes an ALTER TABLE statement in a MySQL database to modify the table structure, the database will send the change message to the Kafka topic, and the streaming processing task will subscribe to this topic to receive the DDL change stream.
[0062] S204, determine the change target according to the received data definition language change statement.
[0063] It can be understood that after receiving the DDL change statement, it needs to be parsed to determine the target object of the change. For example, which table, which field, or which index in which data source of which database has changed. The changes here can be adding a new table or adding, deleting, or modifying fields in an existing table. This step mainly performs syntax analysis on the DDL statement. Different database systems have different DDL syntax rules, but usually include keywords (such as CREATE, ALTER, DROP) and the target object name (such as the namespace area). By parsing these keywords and names, the change target can be determined. For example, for the DDL statement ALTER TABLE users ADD COLUMN age INT, it can be determined through analysis that the change target is the users table and the change operation is to add an integer type field named age.
[0064] S206, obtain the metadata of the changed data source according to the change target.
[0065] It can be understood that after determining the change target, the latest metadata related to this target needs to be obtained from the data source. For example, if the change target is a certain table, the latest structure information (field names, field types, indexes, etc.) of this table needs to be obtained. By interacting with the data source, corresponding query operations are executed to obtain the metadata. Different data sources have different metadata query methods. For example, relational databases can query metadata through system views or specific SQL statements, and file systems can obtain file metadata through the file system API.
[0066] S208, update the metadata of the data source according to the metadata of the changed data source.
[0067] It can be understood that after obtaining the metadata of the changed data source, it needs to be updated to the metadata of the data source maintained in the system to ensure the accuracy and consistency of the metadata. The newly obtained metadata is merged and replaced with the original metadata. For the newly added metadata information, it is directly added to the original metadata; for the modified metadata information, the original information is replaced with the new information; for the deleted metadata information, it is removed from the original metadata. In addition, the update of the data source metadata can also be achieved by setting a scheduled task to synchronize the metadata information regularly; it can also be actively triggered to synchronize the metadata through a page button.
[0068] In one of the embodiments, the acquisition of ETL task metadata includes: extracting the ETL task configuration table from the scheduling system. Obtaining the ETL task metadata according to the ETL task configuration table.
[0069] It is understandable that the ETL task configuration table is a structured data set storing ETL task-related configuration information, including key information such as task name, task dependencies, scheduling frequency, execution time, data input and output paths, transformation rules, etc. It is an important basis for the operation and management of ETL tasks. The scheduling system is responsible for managing and coordinating the execution of ETL tasks. By maintaining the ETL task configuration table, it realizes the unified scheduling and monitoring of tasks.
[0070] As the management center of ETL tasks, the scheduling system stores the ETL task configuration table in a specific data storage method (such as a relational database table, an XML configuration file, or a JSON format file). The extraction operation is based on the interfaces or data access mechanisms provided by the scheduling system to obtain the configuration table data. This process is like looking up the borrowing information of a specific book in the library management system. The scheduling system is the library management system, the ETL task configuration table is the database table recording the book borrowing information, and the extraction operation is to obtain the borrowing information through the query function of the management system. The storage methods and interfaces of different scheduling systems vary, but the purpose is to facilitate the management and scheduling of tasks. If the scheduling system uses a relational database to store the ETL task configuration table, SQL query statements can be used to extract data. For example, "SELECT * FROM etl_task_config_table" can obtain all configuration information. If an XML configuration file is used, XML parsing tools (such as DOM and SAX parsers in Java) can be used to read the file content and extract the required configuration. If a JSON format file is used, a JSON parsing library (such as the json library in Python) can be used to parse the file content into a data structure for extraction.
[0071] The principle of obtaining ETL task metadata from the ETL task configuration table is to parse, transform, and supplement the data in the configuration table. The data in the configuration table is stored in a specific format and needs to be parsed according to predefined rules. For example, the scheduling frequency in the configuration table may be stored in a specific time expression, such as "0 0 2 * * *" indicating execution at 2 am every day. Through a parsing program, this expression is converted into time interval and execution time information that is easy to understand and process, and filled into the scheduling frequency field of the ETL task metadata. For task dependencies, the configuration table may be stored in a form of a dependency graph (such as an adjacency list, an edge list). Through a parsing algorithm, it is converted into a dependency relationship structure that can be directly used in the metadata, such as a directed acyclic graph (DAG) structure, for dependency checking and task sorting during task scheduling and execution.
[0072] In addition, ETL tasks can be carried out separately for different data layers. The data layers can include a common dimension layer, a detailed data layer, a summary data layer, and an application data layer. The common dimension layer is a level in the data warehouse that contains general and shared dimension information, such as time, location, customers, etc. These dimension information can be reused by multiple business topics and analysis scenarios, providing a unified dimension standard for data analysis. The detailed data layer is usually obtained based on the common dimension layer, without much summarization and processing, retaining the integrity and accuracy of the data. The summary data layer is a data layer obtained by performing summarization and aggregation operations on the data in the detailed data layer. It groups and statistically analyzes the detailed data according to certain dimensions and metrics to generate higher-level summary data for meeting the data analysis needs at the macro level. The application data layer is a data layer constructed for specific business applications and analysis requirements. It extracts the required data from the common dimension layer, the detailed data layer, and the summary data layer according to different business scenarios and analysis purposes, and performs further processing and handling to meet the requirements of specific business applications.
[0073] At the common dimension layer, the main goal of the ETL task is to extract, clean, and organize general dimension information and load it into the data warehouse. These dimension information are the foundation of the entire data warehouse, providing a unified dimension standard for subsequent data queries and analysis. By standardizing and ensuring the consistency of the dimension data, data redundancy and inconsistency can be avoided, improving the quality and usability of the data. The ETL task at the detailed data layer focuses on extracting raw business data from the source system and performing necessary cleaning and transformation operations. Since detailed data usually contains a large amount of detailed information, the ETL task needs to ensure the integrity and accuracy of the data, while dealing with issues such as data format conversion, data verification, and data cleaning. The ETL task at the summary data layer is to perform summarization and aggregation operations on the data in the detailed data layer. According to different business requirements and analysis metrics, the detailed data is grouped and statistically analyzed according to certain dimensions to generate higher-level summary data. These summary data can greatly reduce the data storage volume and improve the data query efficiency, being suitable for data analysis and decision-making support at the macro level. The ETL task at the application data layer is to extract the required data from other data layers according to specific business applications and analysis requirements and perform further processing and handling. These data usually require complex calculations, associations, and filtering operations to meet the requirements of specific business scenarios. By integrating and processing the data from different data layers, more personalized and accurate data analysis services can be provided for business users.
[0074] When obtaining the metadata of ETL tasks, it is necessary to extract and organize the corresponding metadata information according to the characteristics of ETL tasks in different data layers. For ETL tasks in the common dimension layer, the metadata should include information such as the definition of dimension tables, the sources of dimension data, and the update frequency of dimension data; for ETL tasks in the detail data layer, the metadata should include information such as the structure of detail data tables, the rules for data extraction, and the logic for data cleaning and transformation; for ETL tasks in the aggregated data layer, the metadata should include information such as the definition of aggregated metrics, the dimensions and grouping methods for aggregation, and the calculation methods for aggregated data; for ETL tasks in the application data layer, the metadata should include information such as the definition of application data tables, the conditions for data extraction, and the logic for data processing and handling. By managing and maintaining these metadata, the execution process of ETL tasks can be better understood and controlled, and the development and maintenance efficiency of data warehouses can be improved.
[0075] In one of the embodiments, the acquisition of metadata for the indicator system includes: obtaining the indicator system of the target data warehouse and obtaining the metadata for the indicator system according to the definitions of the indicators in the indicator system.
[0076] It can be understood that the target data warehouse is the data warehouse for which data lineage tracking needs to be performed in this method. The indicator system is a set of a series of interrelated indicators, which are used to measure and evaluate aspects such as the business status, operational performance, and market performance of an enterprise. The indicator system has a clear hierarchical structure and logical relationship, and can reflect the key characteristics and development trends of the business from different dimensions and levels. It can be divided into atomic indicators, derivative indicators, and application indicators. Among them, atomic indicators are the most basic and indivisible indicators, generally calculated directly from the data in the detail data tables. They directly correspond to a specific field or a single calculation rule, such as "order quantity" and "commodity unit price", etc. Derivative indicators are indicators obtained by performing certain calculations or combinations based on atomic indicators. For example, "sales amount = order quantity × commodity unit price". Derivative indicators are often associated with aggregated data tables. Application indicators are indicators obtained by further processing and applying derivative indicators in combination with specific business scenarios and requirements, and are used to solve specific business problems, such as "customer loyalty indicators" and "market share indicators", etc. Application indicators are generally associated with application data tables.
[0077] The core principle of obtaining the indicator system of the target data warehouse and obtaining the metadata for the indicator system lies in deeply mining and analyzing the data and business logic in the data warehouse. The target data warehouse, as the centralized storage place for data, contains various data related to the business, and the indicator system is constructed based on this data and is used to quantify and evaluate the business status. By analyzing the data structure, table relationships, field meanings, etc. in the data warehouse, the data elements related to the indicator system can be identified.
[0078] After obtaining the index system, further extract and organize the metadata of the index system according to the definitions of each index. The definition of an index usually includes information such as the index name, index calculation formula, data source table and fields, and business interpretation. For any kind of index, it is necessary to clarify its calculation logic, the indexes or table fields it depends on, and the fields stored after output, etc. These are all declared in the definition of the index, and can be extracted by means of keyword matching.
[0079] In one embodiment, the data lineage tracing method further includes: monitoring the execution status of the ETL task node. In the case where the execution status is failed, in response to the user's rerun instruction, restart the ETL task node.
[0080] It can be understood that the execution status refers to the stage and situation of the ETL task node during execution. Common execution statuses include success, failure, running, waiting to execute, etc. The execution status reflects the execution result and current progress of the task node, and is an important basis for monitoring and managing the ETL task. Rerun instruction: When the ETL task node fails to execute, the user can issue an instruction to re-execute the task node according to the specific situation, and this instruction is the rerun instruction. The user can trigger the rerun instruction manually or automatically through rules preset by the system. Monitoring the execution status of the ETL task node is a key link to ensure the smooth completion of the ETL task. The ETL task may be affected by various factors during execution, such as data source failure, network problems, data format errors, etc., and these factors may all cause the task node to fail to execute. By monitoring the execution status of the task node in real time, the system can timely discover problems that occur during task execution.
[0081] When the system detects that the execution status of a certain ETL task node is failed, record the relevant information of the node, such as the task name, execution time, failure reason, etc. At this time, if the user issues a rerun instruction, the system will re-initialize the environment and parameters required for the task node according to the previously recorded task node information, and try to execute the task node again. The purpose of restarting the task node is to try to solve the temporary problems that caused the task to fail before, such as a brief network interruption, the data source being temporarily unavailable, etc. Through the rerun mechanism, the success rate of the ETL task can be improved, the cost of manual intervention can be reduced, and the continuity and accuracy of data processing can be guaranteed.
[0082] When a task node fails to execute, the system can not only record the failure information but also automatically perform fault diagnosis. By using natural language processing technology to analyze the error information in the log files and combining with a preset fault knowledge base, the cause of the fault can be quickly located. For some common faults, the system can automatically attempt to repair them, such as reconnecting to the data source, adjusting the data format, etc. If the automatic repair is successful, the task node is directly rerun; if the automatic repair fails, the user is notified in a timely manner and a detailed fault analysis report is provided. After each rerun of the task node, the effect of the rerun can also be evaluated. Information such as the execution time, resource consumption, and data processing results of the rerun is recorded and compared with the previous execution situation. By analyzing the rerun effect, it is judged whether the rerun has solved the previous problem and whether new problems have been introduced. According to the evaluation results, the rerun strategy and the fault handling process are further optimized to improve the stability and reliability of the entire ETL system.
[0083] In one of the embodiments, please refer to Figure 3 , the data lineage tracing method further includes steps S302 to S308.
[0084] S302. For any data lineage chain, determine whether there is any other path between the data source node and the metric node of the data lineage chain.
[0085] It can be understood that other paths refer to other possible data flow methods from the data source node to the metric node, except for the path represented by the data lineage chain currently being analyzed. Different paths may involve different data processing steps, algorithms, or data storage locations. During the data processing process, there may be multiple different data processing methods and processes from the data source node to the metric node, thus forming different data flow paths. The purpose of determining whether there are other paths is to provide a basis for subsequent path optimization. If there is only one path, then there is no room for path selection and optimization; while if there are multiple paths, it is possible to improve the data processing efficiency, reduce costs, or enhance data quality by selecting a better path. Through a comprehensive analysis of the data lineage chain, all possible paths are found for further evaluation and comparison. Specifically, a breadth-first search or a depth-first search can be started from the data source node to find all paths that can reach the metric node. If the search result shows that the number of paths is greater than 1, it means that there are other paths.
[0086] S304. If so, calculate the average calculation time consumption of the other path and the current path.
[0087] It can be understood that the average calculation time refers to the average of the time spent on each calculation during multiple executions of a certain path. It reflects the overall performance of the path over a period of time and is an important indicator for evaluating the path efficiency. Different data processing paths may have different calculation times due to the algorithms used, the computing resources employed, and the data volume. Statistically calculating the average calculation time can more objectively evaluate the efficiency of each path. By executing the same path multiple times, recording the calculation time for each execution, and then taking the average, the influence of accidental factors that may occur in a single execution can be reduced, and a more stable and reliable performance indicator can be obtained. In this way, when selecting a path subsequently, a decision can be made based on more accurate data.
[0088] S306, select the one with the minimum average calculation time as the target path.
[0089] It can be understood that the target path is the path selected as the optimal one after evaluation among all paths from the data source node to the metric node. This path performs optimally in terms of the average calculation time and can improve efficiency and reduce time costs during the data processing. Selecting the path with the minimum average calculation time as the target path is based on the consideration of improving data processing efficiency. In data processing tasks, especially in scenarios involving large-scale data processing or high real-time requirements, the calculation time is a key factor. By selecting the path with the minimum time consumption, the waiting time for data processing can be reduced, the response speed of the system can be increased, and thus the efficiency of the entire business process can be enhanced.
[0090] S308, send a path optimization prompt to the user according to the target path.
[0091] It can be understood that the path optimization prompt is to organize the relevant information of the target path and the comparison with the current path into prompt information and provide it to the user to help the user understand the directions and methods for optimization, so as to adjust the existing data processing process.
[0092] In one embodiment, a data governance report can also be generated regularly, including the metadata update situation, the ETL task execution situation, the metric system update situation, etc.
[0093] The present application provides a data lineage tracing device, which includes a data acquisition module, a node construction module, a lineage chain construction module, and a display module. The data acquisition module is used to acquire data source metadata, ETL task metadata, and metric system metadata respectively. The node construction module is used to construct data source nodes in the graph database according to the data source metadata, construct ETL task nodes according to the ETL task metadata, construct metric nodes according to the metric system metadata, and determine the node association relationships according to the data source metadata, ETL task metadata, and metric system metadata. The lineage chain construction module is used to construct a data lineage chain with the data source node as the starting point and the metric node as the ending point according to the node association relationships. The display module is used to display the data lineage chain through the graph database.
[0094] For the specific limitations of the data lineage tracing device, reference can be made to the limitations of the data lineage tracing method in the foregoing text, which will not be elaborated here. Each module in the above data lineage tracing device can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, merely a logical function division, and there can be other division methods in actual implementation.
[0095] The present application provides a computer device, which includes one or more processors, and a memory. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by one or more processors, the steps of the data lineage tracing method in any of the foregoing embodiments are executed.
[0096] Schematically, as Figure 4 shown, Figure 4 is an internal structure schematic diagram of a computer device provided by an embodiment of the present application. Referring to Figure 4 , the computer device 400 includes a processing component 402, which further includes one or more processors, and memory resources represented by a memory 401 for storing instructions executable by the processing component 402, such as application programs. The application programs stored in the memory 401 can include one or more than one, each corresponding to a set of instruction modules. In addition, the processing component 402 is configured to execute instructions to execute the steps of the data lineage tracing method in any of the above embodiments.
[0097] The present application provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the data lineage tracing method in any of the foregoing embodiments.
[0098] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0099] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.
[0100] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data lineage tracing method, characterized in that, Including: Respectively obtain data source metadata, ETL task metadata, and metric system metadata; In the graph database, respectively construct a data source node according to the data source metadata, construct an ETL task node according to the ETL task metadata, construct a metric node according to the metric system metadata, and determine the node association relationship according to the data source metadata, ETL task metadata, and metric system metadata; Construct a data lineage chain with the data source node as the starting point and the metric node as the ending point according to the node association relationship; Display the data lineage chain through the graph database.
2. The data lineage tracing method according to claim 1, wherein The obtaining of the data source metadata includes: Send a metadata acquisition request to the target data source to obtain the data source metadata.
3. The data lineage tracing method according to claim 2, wherein The obtaining of the data source metadata further includes: Receive a data definition language change stream through a streaming processing task; Determine the change target according to the received data definition language change statement; Obtain the changed data source metadata according to the change target; Update the data source metadata according to the changed data source metadata.
4. The data lineage tracing method according to claim 1, wherein The obtaining of the ETL task metadata includes: Extract the ETL task configuration table from the scheduling system; Obtain the ETL task metadata according to the ETL task configuration table.
5. The data lineage tracing method according to claim 1, wherein The obtaining of the metric system metadata includes: Obtain the metric system of the target data warehouse, and obtain the metric system metadata according to the definitions of the metrics in the metric system.
6. The data lineage tracing method according to claim 1, wherein Also including: Monitor the execution status of the ETL task node; In the case where the execution status is failed, restart the ETL task node in response to the user's rerun instruction.
7. The data lineage tracing method according to claim 1, wherein Also including: For any one of the data lineage chains, determine whether there is another path between the data source node and the metric node of the data lineage chain; If so, calculate the average calculation time of the other path and the current path; Select the one with the minimum average calculation time as the target path; Send a path optimization prompt to the user according to the target path.
8. A data lineage tracking device, characterized in that, Including: A data acquisition module for respectively obtaining data source metadata, ETL task metadata, and metric system metadata; A node construction module for respectively constructing a data source node according to the data source metadata, constructing an ETL task node according to the ETL task metadata, constructing a metric node according to the metric system metadata in the graph database, and determining the node association relationship according to the data source metadata, ETL task metadata, and metric system metadata; A lineage chain construction module for constructing a data lineage chain with the data source node as the starting point and the metric node as the ending point according to the node association relationship; A display module for displaying the data lineage chain through the graph database.
9. A computer device, characterized in that, Including one or more processors and a memory, where computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the one or more processors, the steps of the data lineage tracing method according to any one of claims 1-7 are executed.
10. A storage medium, characterized in that, The storage medium stores computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the data lineage tracing method according to any one of claims 1-7.
Citation Information
Patent Citations
Data blood relationship generation method and device, storage medium and computer equipment
CN113204594A
Field retrieval and path display method and system for data consanguinity
CN113220945A
Data blood relationship determination method and device, storage medium and electronic device
CN114691786A
Data blood relationship full-link monitoring method and system, terminal and storage medium
CN118939839A
Data-based blood relationship analysis method, apparatus, and device and computer-readable storage medium
WO2021218021A1
Cited By
Multi-dimensional index calculation method based on graph database
CN120744193A
Multi-modal data lineage modeling method and device, electronic equipment and storage medium
CN121117265A