Data lineage analysis method, medium, apparatus, and computing device

By calculating the lineage coverage of the data governance platform, the problem of the inability to assess data lineage quality in existing technologies is solved, thus enabling the assessment of the lineage quality and improvement of the coverage of the data governance platform.

CN117149872BActive Publication Date: 2026-02-24HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311103588.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-29
Publication Date
2026-02-24
Estimated Expiration
2043-08-29

AI Technical Summary

Technical Problem

Existing technologies cannot effectively assess the quality of the bloodline relationships revealed by the parsing.

Method used

By obtaining the number of tasks executed at the platform layer and the number of tasks stored at the lineage storage layer, the lineage coverage of the data governance platform is calculated to measure lineage quality.

Benefits of technology

It enables the assessment of the lineage quality of the data governance platform, and can identify and improve tasks that have not generated data lineage, thereby increasing lineage coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117149872B_ABST
    Figure CN117149872B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a data lineage analysis method, medium, device and computing device for determining the lineage quality of a data governance platform, the data governance platform comprising a platform layer for executing tasks within the data governance platform, a lineage storage layer for extracting and storing lineage relationships based on the analysis plan corresponding to the tasks, by obtaining a first number of tasks executed by the platform layer, obtaining a second number of tasks corresponding to the lineage relationships stored in the lineage storage layer, determining the lineage coverage of the data governance platform based on the first number and the second number, since the task is running, the data flow is accompanied by the data flow, and the data flow should generate the lineage, so that the number of tasks executed by the platform layer and the number of tasks stored in the lineage storage layer can measure the lineage coverage of the data governance platform, and the lineage quality of the analyzed lineage relationship is evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure relate to the field of data governance technology, and more specifically, the embodiments of this disclosure relate to a data lineage analysis method, medium, apparatus, and computing device. Background Technology

[0002] This section is intended to provide background or context for the embodiments of this disclosure as set forth in the claims. The description herein is not intended to be a prior art simply because it is included in this section.

[0003] In the era of big data, the widespread adoption of big data computing engines such as Hive and Spark has greatly improved data processing efficiency. At the same time, data flow has become increasingly complex and diverse, with intricate relationships between data tables. Clarifying the dependencies between data tables is crucial for troubleshooting, tracing the source of problems, and building organizational user relationships. Therefore, data governance has gained increasing attention in the industry, and establishing data lineage is at the core of data governance.

[0004] Currently, the construction of data lineages is relatively mature. For example, data lineage parsing tools can be used to perform targeted pattern matching and parse lineages for specific SQL scenarios. However, existing technologies do not provide a method for determining the quality of the parsed lineage relationships. Summary of the Invention

[0005] This disclosure provides a data lineage analysis method, medium, apparatus, and computing device to evaluate the lineage quality of the resolved lineage relationships.

[0006] In a first aspect of this disclosure, a data lineage analysis method is provided, applied to a data governance platform. The data governance platform includes a platform layer and a lineage storage layer. The platform layer executes tasks within the data governance platform, and the lineage storage layer extracts and stores lineage relationships based on the parsing plan corresponding to the task. The lineage relationship is the relationship between an input table and an output table marked by the task. The method includes:

[0007] Obtain the first number of tasks executed by the platform layer; obtain the second number of tasks corresponding to blood relations stored in the blood relation storage layer;

[0008] The lineage coverage rate of the data governance platform is determined based on the first quantity and the second quantity.

[0009] In a second aspect of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the method provided in the first aspect.

[0010] In a third aspect of this disclosure, a data lineage analysis apparatus is provided, applied to a data governance platform. The data governance platform includes a platform layer and a lineage storage layer. The platform layer executes tasks within the data governance platform, and the lineage storage layer extracts and stores lineage relationships based on a parsing plan corresponding to the task. The lineage relationship is the relationship between an input table and an output table marked by the task. The apparatus includes:

[0011] The acquisition module is used to acquire a first number of tasks executed by the platform layer; and to acquire a second number of tasks corresponding to blood relations stored in the blood relation storage layer.

[0012] The first determining module is used to determine the lineage coverage rate of the data governance platform based on the first quantity and the second quantity.

[0013] In a fourth aspect of this disclosure, a computing device is provided, comprising: at least one processor and a memory; the memory storing computer execution instructions; and at least one processor executing the computer execution instructions stored in the memory, such that at least one processor performs the method provided in the first aspect.

[0014] In this embodiment of the disclosure, the lineage quality of a data governance platform is determined. The data governance platform includes a platform layer for executing tasks within the data governance platform and a lineage storage layer for extracting and storing lineage relationships based on the parsing plan corresponding to the task. By obtaining a first number of tasks executed by the platform layer and a second number of tasks corresponding to lineage relationships stored in the lineage storage layer, the lineage coverage of the data governance platform is determined based on the first and second numbers. Since data flow accompanies task execution, lineage should be generated. Therefore, the lineage coverage of the data governance platform can be measured by using the number of tasks executed by the platform layer and the number of tasks stored in the lineage storage layer, thereby enabling the evaluation of the lineage quality of the parsed lineage relationships. Attached Figure Description

[0015] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which:

[0016] Figure 1 A schematic diagram illustrating an application scenario provided according to an embodiment of this disclosure is shown.

[0017] Figure 2 A schematic flowchart of a data lineage analysis method according to an embodiment of the present disclosure is shown.

[0018] Figure 3 A schematic flowchart of another data lineage analysis method provided according to an embodiment of the present disclosure is shown.

[0019] Figure 4 The illustration schematically shows a diagram of data lineage analysis based on a basic lineage coverage set according to an embodiment of the present disclosure;

[0020] Figure 5 A schematic flowchart of another data lineage analysis method provided according to an embodiment of the present disclosure is shown.

[0021] Figure 6 This illustration schematically shows a method for verifying a bloodline analysis tool using a third-party bloodline analysis tool, according to an embodiment of the present disclosure.

[0022] Figure 7 A schematic diagram illustrating continuous output of bloodline coverage according to an embodiment of the present disclosure is shown.

[0023] Figure 8 A schematic diagram illustrating a continuous output of lineage accuracy according to an embodiment of the present disclosure is shown.

[0024] Figure 9 A schematic diagram of the structure of a computer-readable storage medium provided according to an embodiment of the present disclosure is shown.

[0025] Figure 10 A schematic diagram of the structure of a data lineage analysis device provided according to an embodiment of the present disclosure is shown.

[0026] Figure 11 A schematic diagram of the structure of a computing device provided according to an embodiment of the present disclosure is shown.

[0027] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0028] The principles and spirit of this disclosure will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0029] Those skilled in the art will recognize that embodiments of this disclosure can be implemented as a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0030] According to embodiments of this disclosure, a data lineage analysis method, medium, apparatus, and computing device are proposed.

[0031] In this document, it should be understood that the terminology used is for convenience of understanding only and does not imply any limitation on its meaning. Furthermore, any number of elements in the accompanying drawings is for illustrative purposes only and not for limitation, and any naming is for distinction only and has no limiting meaning.

[0032] In addition, the data involved in this disclosure may be data authorized by the user or fully authorized by all parties. The collection, dissemination and use of the data shall comply with the requirements of relevant national laws and regulations. The implementation methods / executives of this disclosure may be combined with each other.

[0033] The following is a description of the terminology used in this disclosure:

[0034] Hive: An open-source big data warehouse processing engine that enables the extraction, transformation, and loading of big data by writing SQL (Structured Query Language) statements;

[0035] Spark: A data processing engine similar to Hive, it is a distributed batch processing engine that performs computations in memory. Spark has a greater performance advantage than Hive.

[0036] ETL (Extract-Transform-Load): The process of extracting, transforming, and loading data from the source to the destination; a synonym for data processing.

[0037] Data lineage: A representation of the data dependencies generated during the ETL process, visualizing the input and output data as parent-child relationships.

[0038] Parsing plan: A tree structure obtained by parsing the schema information (database, table, column, etc.) in the abstract syntax tree (derived from SQL statements). It can be used to generate the actual execution plan. The parsing plan contains information about the input table, output table, and the task to which the SQL statement belongs.

[0039] Platform layer: An integrated platform system that includes services such as user authentication, task submission, task management, monitoring and alerting, such as big data platforms and data middleware.

[0040] Engine layer: Located below the platform layer, it is the layer to which the components that actually execute computing tasks (such as computing engines like Hive and Spark) belong.

[0041] Job: Tasks on big data platforms, including Hive, Spark, and other types.

[0042] Middleware: A type of software that provides connection between system software and application software, facilitating communication between various software components.

[0043] MD5 hash value: A widely used cryptographic hash function that produces a 128-bit (16-byte) hash value. Invention Overview

[0045] The inventors have discovered that when bloodline relationships are analyzed using bloodline analysis tools, the quality of the bloodline cannot be determined. Currently, the industry mainly focuses on constructing bloodlines and implementing data bloodlines, but there is still a lack of evaluation methods for the quality of bloodlines after implementation.

[0046] To address the aforementioned issues, this solution proposes a data lineage analysis method and designs a method for calculating lineage coverage. Task execution occurs at the platform layer, while lineage relationships are stored at the lineage storage layer. Since tasks typically involve data flow during runtime, and data flow naturally generates lineage relationships, the lineage storage layer is linked to the platform layer. Specifically, the number of tasks with lineage relationships is obtained from the lineage storage layer, and the number of executed tasks is obtained from the platform layer (task execution layer). The ratio of these two values ​​is used as the lineage coverage rate of the data governance platform, thereby measuring the lineage quality of the data governance platform.

[0047] After introducing the basic principles of this disclosure, various non-limiting embodiments of this disclosure will be described in detail below.

[0048] Application Scenarios Overview

[0049] First refer to Figure 1 As shown, Figure 1 The illustration shows an application scenario diagram according to an embodiment of the present disclosure, such as... Figure 1As shown, the data governance platform includes a platform layer and a lineage storage layer. In addition, it may also include middleware. The platform layer can execute tasks, and the lineage resolution tool can intercept the resolution plan after task execution, generate lineage relationships and send them to the middleware. The middleware can store the lineage relationships, and the lineage storage layer can consume and store lineage relationships from the middleware. When the lineage coverage is determined based on the method of this application, the lineage coverage can be displayed to the user so that the user can understand the lineage coverage of the data governance platform.

[0050] Exemplary methods

[0051] The following is combined with Figure 1 Application scenarios, refer to Figures 2-8 This document describes a data lineage analysis method according to exemplary embodiments of the present disclosure. It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in any way. Rather, the embodiments of the present disclosure can be applied to any applicable scenario.

[0052] refer to Figure 2 , Figure 2 A schematic flowchart illustrating a data lineage analysis method according to an embodiment of this disclosure is shown. Figure 2 As shown, this method is applied to a computing device equipped with a data governance platform. The data governance platform includes a platform layer and a lineage storage layer. The platform layer executes tasks within the data governance platform, and the lineage storage layer extracts and stores lineage relationships based on the parsing plan corresponding to the task. The lineage relationship is the relationship between the input table and the output table marked by the task. The method includes:

[0053] Step S201: Obtain the first number of tasks executed by the platform layer; obtain the second number of tasks corresponding to the bloodline relationships stored in the bloodline storage layer.

[0054] See Figure 1 A data governance platform comprises a platform layer, middleware, and a lineage storage layer. The platform layer receives and executes user-uploaded tasks. An engine is configured within the platform layer; when the platform layer receives tasks, it translates these tasks into a language the engine can understand, allowing the engine to execute them. A task may contain multiple SQL statements, and the execution process can involve transforming data tables; for example, transforming a base table (Table 1) into another table (Table 2). Lineage identifies this data transformation process, such as transforming from Table 1 to Table 2.

[0055] After the engine in the platform layer completes its task, the lineage resolution tool in the platform layer can intercept the resolution plan through hooks or listeners. The resolution plan is obtained by the engine after task execution. After intercepting the resolution plan, the lineage resolution tool can traverse the plan to collect the dependencies between data tables, thus forming the lineage relationship. The resolution plan contains the input table, output table, and task identification information corresponding to the SQL statement. For example, based on the resolution plan of an SQL statement, the input table can be determined to be data table 1, and the output table to be data table 2. Then, the lineage relationship is composed of data table 1, data table 2, and the task name or task ID to which the SQL statement belongs.

[0056] After the lineage resolution tool generates lineage relationships, it can send these relationships to middleware for temporary storage. The lineage storage layer can then consume and store these lineage relationships from the middleware. When storing lineage relationships, tasks can be used to connect various data tables. Lineage relationships can be in the form of graph data, with data tables as nodes and tasks as edges, and nodes linked together by edges. Therefore, the lineage storage layer stores tasks with defined data lineages.

[0057] For the above-mentioned lineage transfer scenario, tasks in the lineage storage layer and platform layer can be associated. Since tasks are usually accompanied by data transfer during runtime, data lineage will be generated when data transfer occurs, and thus lineage coverage can be defined.

[0058] Therefore, the initial number of tasks executed by the platform layer can be obtained. Specifically, when a task is completed, the platform layer can count the number of tasks executed. Alternatively, the platform layer can mark executed tasks, such as marking them as executed, to facilitate counting the number of executed tasks based on this "executed" mark.

[0059] It can also obtain a second number of tasks corresponding to blood relations in the bloodline storage layer. It can count the tasks contained in the bloodline storage layer and use the number of contained tasks as the second number.

[0060] When each task in the platform layer is executed once, the number of tasks executed in the platform layer can be directly determined as the first number, and the number of tasks contained in the lineage storage layer can be determined as the second number.

[0061] When one or more tasks in the platform layer have been executed more than once, the first quantity can be determined based on the number of times each task has been executed. In the lineage storage layer, the lineage relationship can be identified by the task identifier and execution ID, and the second quantity can be determined based on the number of execution IDs corresponding to each task.

[0062] Optionally, a task in the platform layer may be executed N times. Each execution of the same task will be identified by a different execution ID. In this case, when determining the first number of tasks executed in the platform layer, the first number can be determined based on the number of execution IDs corresponding to each task. For example, if there are tasks 1, 2, and 3 in the platform layer, and task 1 is executed 3 times, while task 2 and task 3 are each executed once, then the number of execution IDs corresponding to task 1 is 3, and the number of execution IDs corresponding to task 2 and task 3 is 1 each. Therefore, the first number can be determined to be 5.

[0063] When a task is executed N times, the lineage storage layer stores the lineage relationships of that task after each execution. For the same task, the lineage relationships generated after each execution are marked using the same task identifier and different execution IDs. Therefore, when determining the second number of tasks corresponding to the lineage relationships in the lineage storage layer, the second number can be determined based on the number M (M less than or equal to N) of different execution IDs corresponding to each task. For example, if the platform layer contains tasks 1, 2, and 3, and task 1 is executed 3 times, task 2 and task 3 are each executed once, and the lineage storage layer stores lineage relationship 1 marked with task 1 and execution ID 1, lineage relationship 2 marked with task 2 and execution ID 2, and lineage relationship 3 marked with task 3 and execution ID 3, and the lineage storage layer stores lineage relationships marked by task 2 and task 3 respectively, then the second number can be determined to be 5.

[0064] Optionally, the first number of tasks executed by the platform layer and the second number of tasks corresponding to lineage relationships in the lineage storage layer are the numbers under the same time metric. For example, the first number of tasks executed by the platform layer and the second number of tasks corresponding to lineage relationships in the lineage storage layer can be counted on the same day.

[0065] Step S202: Determine the lineage coverage rate of the data governance platform based on the first quantity and the second quantity.

[0066] After obtaining the first and second quantities, the ratio of the second quantity to the first quantity is calculated, and this ratio is the lineage coverage rate of the data governance platform.

[0067] The lineage coverage rate of a data governance platform, calculated using the above method, can measure the platform's lineage coverage and, to some extent, assess the quality of its lineage. For example, if a data governance platform has a lineage coverage rate of 80%, then 20% of tasks lack a data lineage, and these tasks can be further analyzed.

[0068] When determining the first and second quantities, the tasks stored in the platform layer and lineage storage layer can also be filtered according to the user's needs to count the number of filtered tasks and obtain the lineage coverage rate that meets the user's needs.

[0069] The tasks running on a data governance platform can be composed of multiple dimensions. For example, the task dimension can be divided into Jobs and Workflows, the cluster dimension into Cluster A and Cluster B, and the data source dimension into Hive and Spark. Therefore, the lineage coverage of various ETL task types can be calculated. Since the generation of data lineage mainly relies on the parsing of lineage resolution tools, the lineage coverage can be used to evaluate the support of lineage resolution tools for different types of tasks and the quality of the lineage.

[0070] When a user is concerned with the job lineage coverage in cluster A, they can count the number of jobs in the cluster A lineage storage layer on a given day, and the number of jobs executed in the cluster A platform layer on that day. Based on these two numbers, the job lineage coverage of cluster A can be calculated. Job lineage coverage of cluster A = Number of jobs counted in the cluster A lineage storage layer on a given day / Number of jobs executed in the cluster A platform layer on that day.

[0071] When a user focuses on the lineage coverage of Hive type jobs in cluster A, the Hive type job lineage coverage of cluster A is calculated as follows: Hive type job lineage coverage = Number of Hive type jobs counted in the lineage storage layer of cluster A on a given day / Number of Hive type jobs executed by the platform layer of cluster A on that day. Spark type job lineage coverage = Number of Spark type jobs counted in the lineage storage layer on a given day / Number of Spark type jobs executed by the platform layer on that day.

[0072] In this embodiment, the data governance platform includes a platform layer and a lineage storage layer. The platform layer is used to execute tasks within the data governance platform, and the lineage storage layer is used to extract and store lineage relationships based on the parsing plan corresponding to the task. The lineage relationship is the relationship between the input table and the output table marked by the task. By obtaining a first number of tasks executed by the platform layer and a second number of tasks corresponding to the lineage relationships stored in the lineage storage layer, the lineage coverage of the data governance platform is determined based on the first and second numbers. Since data flow accompanies task execution, lineage should be generated. Therefore, the lineage coverage of the data governance platform can be measured by the number of tasks executed by the platform layer and the number of tasks stored in the lineage storage layer, thereby enabling the evaluation of the lineage quality of the parsed lineage relationships.

[0073] After calculating the bloodline coverage rate, further analysis can be conducted on unrelated tasks, and adjustments can be made to improve the bloodline coverage rate.

[0074] refer to Figure 3 , Figure 3 The schematic diagram illustrates a flowchart of another data lineage analysis method provided according to an embodiment of the present disclosure, such as... Figure 3 As shown, the method includes:

[0075] Step S301: Obtain the first number of tasks executed by the platform layer; obtain the second number of tasks corresponding to the bloodline relationships stored in the bloodline storage layer;

[0076] Step S302: Determine the lineage coverage rate of the data governance platform based on the first quantity and the second quantity.

[0077] The execution process of step S301 is similar to that of step S201, and the execution process of step S302 is similar to that of step S202, so it will not be described again here.

[0078] Figure 4 The illustration schematically depicts a method for analyzing data lineage based on a basic lineage coverage set according to an embodiment of this disclosure, such as... Figure 4 As shown, tasks executed at the platform layer are compared with those stored in the lineage storage layer to identify unrelated tasks. The type of unrelated task is determined based on the basic lineage coverage set, and the corresponding target operation is then executed for each type, thereby improving lineage coverage. The basic lineage coverage set stores multiple target SQL statement types (the SQL statements corresponding to these target SQL statement types should have data flow after execution), and the SQL statement types that the lineage parsing tool can correctly parse include all target SQL statement types in the basic lineage coverage set. The following details the methods for identifying unrelated tasks and improving lineage coverage.

[0079] Step S303: In response to the bloodline coverage rate being less than the target bloodline coverage rate, the tasks stored in the bloodline storage layer and the tasks executed by the platform layer are compared according to the target attribute to obtain tasks without bloodlines.

[0080] The unrelated tasks include a first task that has been executed by the platform layer but not stored in the lineage storage layer, and a second task that has been stored in the lineage storage layer but not executed by the platform layer.

[0081] When the calculated lineage coverage is less than the target lineage coverage, it indicates that the lineage coverage needs to be improved. For example, if the target lineage coverage is 100% and the calculated lineage coverage is 80%, it means that some tasks executed in the platform layer are unrelated tasks. Therefore, it is necessary to compare the tasks stored in the lineage storage layer with the tasks stored in the platform layer to filter out unrelated tasks.

[0082] When identifying unrelated tasks, the platform layer can be compared with the tasks stored in the lineage storage layer based on target attributes. For example, target attributes can include task identifiers. For task identifier 1, if the platform layer has executed tasks corresponding to task identifier 1, and the lineage storage layer also stores tasks corresponding to task identifier 1, then the task is considered a related task. Through comparison, we can identify first tasks that have been executed by the platform layer but not stored in the lineage storage layer, and second tasks that have been stored in the lineage storage layer but not executed by the platform layer. These tasks are unrelated tasks.

[0083] In one exemplary embodiment of this disclosure, the target attribute includes the name of the source platform, the task identifier, and the execution ID; the source platform represents the source information of the task corresponding to the target attribute; the task identifier is used to distinguish different tasks; and the execution ID is used to identify each execution of the same task.

[0084] When comparing tasks in the platform layer and the lineage storage layer, the target attributes can include the name of the source platform, the task identifier, and the execution ID. The task identifier can be a task ID, representing a unique identifier for the task. If a task executed in the platform layer and a task in the lineage storage layer share the same source platform name, task identifier, and execution ID, then the task is considered a related task. The execution ID represents a unique identifier for each execution of the same task; since tasks may originate from other platforms, the name of the source platform can be used to distinguish tasks with the same task identifier.

[0085] By using the three target attributes mentioned above, two tasks with the same task identifier can be effectively distinguished, thereby improving the accuracy of identifying unrelated tasks.

[0086] Step S304: Determine the type of the unrelated task based on the unrelated task and the basic lineage coverage set; the basic lineage coverage set contains multiple target SQL statement types.

[0087] The SQL statement corresponding to the target SQL statement type has data flow.

[0088] After identifying unrelated tasks, they can be analyzed to determine their type. Optionally, the type of unrelated task can be determined based on a basic lineage coverage set. This set contains multiple target SQL statement types, each corresponding to a SQL statement that involves data flow; therefore, the SQL statement needs to generate a data lineage. Whether a SQL statement involves data flow can be determined by staff. Optionally, the SQL statement can be output to the relevant staff, and their feedback can be received to determine if data flow exists. The type of each unrelated task can be determined based on the basic lineage coverage set.

[0089] Optionally, the SQL statement types (SQL syntax) provided by the Spark and Hive official websites, as well as the SQL statement types of interest to users that involve data flow, can be stored in the basic lineage coverage set. For example, the target SQL statement types in the basic lineage coverage set can include the following:

[0090]

[0091] Once the basic lineage coverage set is determined, the unrelated tasks selected can be analyzed based on this basic lineage coverage set to determine why the unrelated tasks did not generate data lineage.

[0092] By comparing the tasks stored in the lineage storage layer with the tasks executed in the platform layer, unrelated tasks can be obtained. This allows for focused analysis of unrelated tasks, eliminating the need to analyze related tasks through the basic lineage coverage set, thus improving the efficiency of task analysis.

[0093] Step S305: Perform the corresponding target operation according to the type of the unrelated task to improve the lineage coverage of the data governance platform.

[0094] Once the type of unrelated task is determined, the reasons for the occurrence of this type of unrelated task can be analyzed to determine the corresponding target operation and execute the target operation. This target operation can improve the lineage coverage.

[0095] After identifying the type of unrelated task, the number of unrelated tasks generated by the data governance platform can be reduced by performing the corresponding target operations, thereby improving the lineage coverage of the data governance platform.

[0096] There are four types of unrelated missions, and the four types and their corresponding target operations are described below.

[0097] In one exemplary embodiment of this disclosure, determining the type of the unrelated task based on the unrelated task and the basic kinship coverage set includes:

[0098] If the type of at least one SQL statement corresponding to the first task exists in the basic lineage coverage set, then the first task is determined to be a first type of unrelated task.

[0099] Perform the corresponding target operation according to the type of the unrelated task, including:

[0100] In response to the first task being a first type of unrelated task, it is determined that there is an anomaly in the lineage resolution tool and / or the lineage consumption link; the lineage consumption link is used to represent the path through which the lineage storage layer consumes the lineage relationship; the lineage resolution tool is used to extract the lineage relationship of the task based on the resolution plan of the task.

[0101] The first task is one that has been executed at the platform layer but not stored at the lineage storage layer. For the first task, we can determine whether there exists at least one corresponding SQL statement whose type exists in the basic lineage coverage set. Since a task can contain multiple SQL statements, if the type of at least one SQL statement exists in the basic lineage coverage set, it means that the task should generate a data lineage. Optionally, "the type of at least one SQL statement exists in the basic lineage coverage set" means that the type of at least one SQL statement is the same as the type of a target SQL statement in the basic lineage coverage set.

[0102] However, if the first task is not stored in the lineage storage layer, it may indicate that the lineage resolution tool is malfunctioning and has failed to correctly parse the SQL statement in the first task, resulting in the absence of data lineage for that task. Alternatively, it may indicate that the lineage resolution tool is not malfunctioning, but the lineage consumption chain is malfunctioning. The lineage consumption chain represents the path by which the lineage storage layer consumes lineage relationships from the middleware.

[0103] If at least one SQL statement corresponding to the first task has a type that exists in the basic lineage coverage set, then the first task is determined to be a first type of unrelated task. When it is determined that a first type of unrelated task exists, it can be determined that there is an anomaly in the lineage resolution tool and / or the lineage consumption link, thereby allowing for investigation of the lineage resolution tool and the lineage consumption link to identify the problems.

[0104] Under normal circumstances, lineage resolution tools can correctly parse SQL statements of the target SQL statement type in the basic lineage storage layer. However, when the lineage resolution tool malfunctions, it may fail to correctly parse some SQL statements of the target SQL statement type, resulting in the first type of unrelated tasks. By troubleshooting and fixing the lineage resolution tool, it can be enabled to correctly parse SQL statements of any target SQL statement type, reducing the generation of unrelated tasks.

[0105] When there is an anomaly in the bloodline consumption chain, it may be due to a problem with the bloodline consumption code logic. The bloodline consumption code logic can be modified to eliminate the problem in the bloodline consumption chain, thereby reducing the generation of unrelated tasks caused by the problem in the bloodline consumption chain.

[0106] Using the methods described above, we can identify the first type of unrelated tasks and rectify the corresponding problems to reduce the generation of the first type of unrelated tasks, thereby improving the kinship coverage rate.

[0107] In one exemplary embodiment of this disclosure, determining the type of the unrelated task based on the unrelated task and the basic kinship coverage set includes:

[0108] If the types of all SQL statements corresponding to the first task do not exist in the basic lineage coverage set, and the first task has data flow, then the first task is determined to be a second type of unrelated task.

[0109] Perform the corresponding target operation according to the type of the unrelated task, including:

[0110] In response to the first task being the second type of unrelated task, a scenario enhancement operation is performed on the lineage resolution tool; the scenario enhancement is used to increase the lineage resolution tool's ability to correctly parse preset SQL statement types.

[0111] When the types of all SQL statements corresponding to the first task do not exist in the basic lineage coverage set, but the user confirms that the first task involves data flow, it indicates that the target SQL statement types in the basic lineage coverage set do not cover all SQL statement types with data flow. Lineage parsing tools typically only support parsing SQL statements corresponding to the target SQL statement types in the basic lineage coverage set. Therefore, when the types of all SQL statements corresponding to the first task do not exist in the basic lineage coverage set, and the first task involves data flow, it can be determined that there are SQL statements with data flow in each SQL statement of the first task, and the determined SQL statement type is set as the preset SQL statement type.

[0112] Once the preset SQL statement type is determined, it is added to the basic lineage coverage set, and the lineage parsing tool is enhanced for specific scenarios. Scenario enhancement means that the lineage parsing tool can support the determined SQL statement type, that is, it can correctly parse SQL statements of that type.

[0113] Using the above methods, the second type of unrelated tasks can be identified. Furthermore, the lineage analysis tool can be enhanced to reduce the generation of the second type of unrelated tasks, thereby improving the lineage coverage. By adding the preset SQL statement type to the basic lineage coverage set, the task corresponding to the SQL statement of the preset SQL statement type can be avoided from being identified as an unrelated task in the future, which will facilitate the correct identification of unrelated tasks in the future.

[0114] In one exemplary embodiment of this disclosure, determining the type of the unrelated task based on the unrelated task and the basic kinship coverage set includes:

[0115] If the types of all SQL statements corresponding to the first task do not exist in the basic lineage coverage set, and the first task does not have data flow, then the first task is determined to be a third type of unrelated task.

[0116] Perform the corresponding target operation according to the type of the unrelated task, including:

[0117] In response to the first task being the third type of unrelated task, a prompt message indicating that the third type of unrelated task will not be executed is output; and / or, the first quantity is corrected, and the lineage coverage rate is determined based on the corrected quantity; the corrected quantity is the difference between the first quantity and the number of tasks corresponding to the third type of unrelated task.

[0118] Under normal circumstances, the execution of a task involves data flow. If the types of the SQL statements corresponding to the first task do not exist in the basic lineage coverage set, and the first task does not involve data flow, then the task is a third type of unrelated task. The execution of this type of task is not very meaningful because the definition of data lineage is the relationship between data tables that involve data flow. If a task does not involve data flow after execution, then this part of the task is not very meaningful for analyzing data lineage. However, the execution of this task consumes resources such as CPU and memory. Therefore, the execution of this type of task can be stopped.

[0119] For the third type of unrelated tasks, a prompt message can be displayed to the user indicating that the third type of unrelated tasks will not be executed. Once the user confirms that the third type of unrelated tasks will be stopped, the platform layer can stop the execution of the third type of unrelated tasks, thereby improving the lineage coverage of the data governance platform after subsequent task execution.

[0120] In addition, the currently calculated lineage coverage rate can be corrected. For example, the first number is the number of tasks executed by the platform layer. The first number can be corrected by taking the difference between the first number and the number of third-category unrelated tasks as the corrected first number, and then the lineage coverage rate can be calculated based on the corrected first number, thereby improving the lineage coverage rate calculated this time.

[0121] The above methods can be used to process the third type of unrelated tasks, thereby improving the current lineage coverage rate and the subsequent lineage coverage rate of the data governance platform.

[0122] In one exemplary embodiment of this disclosure, the method further includes:

[0123] If the second task is included in the unrelated tasks, then the second task is determined to be a fourth type of unrelated task;

[0124] Perform the corresponding target operation according to the type of the unrelated task, including:

[0125] If the second task is the fourth type of unrelated task, then it is determined that the lineage analysis tool is faulty.

[0126] If a second unrelated task exists within an unrelated task list, it indicates a malfunction in the lineage analysis tool due to an incorrect parsing error. If a fourth type of unrelated task is identified, the lineage analysis tool can be repaired to resolve the malfunction.

[0127] The above methods can reduce the generation of the fourth type of unrelated tasks and improve the lineage coverage of the data governance platform.

[0128] The above embodiments provide a detailed description of bloodline coverage and methods for improving bloodline coverage. In addition, bloodline accuracy can be used to measure whether the bloodline relationship resolved by the bloodline analysis tool is correct.

[0129] Figure 5 A schematic flowchart of another data lineage analysis method according to an embodiment of the present disclosure is shown, such as... Figure 5 As shown, the method includes:

[0130] Step S501: Determine the number of third tasks in the bloodline storage layer; the third task is the task of obtaining the correct bloodline relationship through bloodline analysis tools.

[0131] Step S502: Determine the lineage accuracy of the data governance platform based on the first quantity and the number of the third task.

[0132] In addition to lineage coverage, lineage accuracy can also be calculated. Lineage accuracy measures the proportion of tasks for which the lineage analysis tool correctly identifies lineage relationships out of the total number of tasks executed at the platform layer.

[0133] When determining the correct bloodline relationship obtained by the bloodline analysis tool, a third-party bloodline analysis tool can be used for verification. Specifically, a third-party bloodline analysis tool can be introduced, and the bloodline relationship obtained by the third-party tool can be compared with that obtained by the bloodline analysis tool. For a given task, if the two analysis results match, it indicates that the bloodline relationship extracted by the bloodline analysis tool is correct, thus obtaining the number of third-party tasks.

[0134] The lineage accuracy can be obtained by calculating the ratio of the number of third tasks to the number of first tasks. Lineage accuracy reflects the reliability of the lineage analysis tool (its ability to correctly resolve lineage relationships). A low lineage accuracy indicates that the lineage analysis tool currently used by the data governance platform is unreliable and needs rectification; conversely, a high lineage accuracy indicates that the lineage analysis tool currently used by the data governance platform is highly reliable.

[0135] The reliability of kinship analysis tools can be measured by calculating the accuracy rate. When the accuracy rate is low, the tools can be rectified in a timely manner to ensure their reliability.

[0136] In one exemplary embodiment of this disclosure, determining the number of third tasks in the bloodline storage layer includes:

[0137] A third-party lineage analysis tool is used to extract the second lineage relationship corresponding to the analysis plan of the task; the third-party lineage analysis tool is used to verify the correctness of the first lineage relationship extracted by the lineage analysis tool.

[0138] A first entity is generated based on a first lineage relationship corresponding to the task, and a second entity is generated based on a second lineage relationship corresponding to the task; the first entity includes a task identifier, an input data table, and an output data table; the second entity includes a task identifier, an input data table, and an output data table.

[0139] In response to the first entity and the second entity being the same, the task corresponding to the first entity or the second entity is determined as the third task, and the number of the third tasks is determined; when the task identifier in the first entity is the same as the task identifier in the second entity, and at the same time, the input data table in the first entity is the same as the input data table in the second entity, and at the same time, the output data table in the first entity is the same as the output data table in the second entity, then the first entity and the second entity are the same.

[0140] When using a third-party lineage analysis tool to determine the number of third tasks, the tool can be used to extract the analysis plan for each task to obtain the second lineage relationship. The lineage relationship obtained by the lineage analysis tool in the current data governance platform is the first lineage relationship. The number of third tasks can be obtained by comparing the first and second lineage relationships.

[0141] Specifically, a first entity can be generated based on the first lineage relationship, and a second entity can be generated based on the second lineage relationship. Then, the first and second entities can be compared. Both the first and second entities can contain the task identifier, an input data table, and an output data table. For example, when the lineage relationships of three tasks are obtained, three first entities can be obtained.

[0142] When the first entity is the same as the second entity, the task identifiers in both entities must be the same, and the input and output data tables must be identical. This means that for this task, the lineage analysis tool and the third-party lineage analysis tool obtain the same lineage relationship. In this case, the task is the third task. Once the third task is determined, the number of third tasks can be counted.

[0143] By generating the first and second entities, it is easier to compare the blood relations obtained by the two bloodline analysis tools, thereby improving the efficiency of determining the third task.

[0144] In one exemplary embodiment of this disclosure, the method further includes:

[0145] A first hash value is determined based on the first entity corresponding to the task, and the first hash value and the task identifier corresponding to the task are stored in a first data table;

[0146] Determine a second hash value based on the second entity corresponding to the task, and store the second hash value and the task identifier corresponding to the task in a second data table;

[0147] In response to the first entity and the second entity being the same, the task corresponding to the first entity or the second entity is determined as the third task, including:

[0148] If the first hash value in the first data table is the same as the second hash value in the second data table, then the task corresponding to the first entity or the second entity is determined as the third task.

[0149] When comparing the first entity with the second entity, it is necessary to compare the task identifier, the input data table, and the output data table separately. Moreover, the input data table and the output data table can usually contain multiple data tables, which makes the comparison cumbersome and inefficient.

[0150] Based on the above issues, we consider generating hash values ​​for the first entity and the second entity respectively. The hash value is the MD5 value. When the first entity and the second entity are the same, the corresponding hash values ​​are also the same, and when the first entity and the second entity are different, the corresponding hash values ​​are also different.

[0151] Figure 6 The illustration schematically depicts a method for verifying a kinship analysis tool using a third-party kinship analysis tool, according to an embodiment of this disclosure. Figure 6 As shown, a hash value can be generated based on the entity, and the hash value and task identifier can be stored in a data table. Specifically, for the first entity (corresponding to the first lineage relationship resolved by the lineage resolution tool), a first hash value can be generated and stored in the first data table along with the task identifier. For the second entity (corresponding to the second lineage relationship resolved by the third-party lineage resolution tool), a second hash value can be generated and stored in the second data table along with the task identifier. The hash values ​​in the two data tables can then be compared.

[0152] By generating hash values, only the hash values ​​need to be compared, eliminating the need to compare the input and output data tables multiple times, which can further improve the efficiency of determining the third task.

[0153] In an exemplary embodiment of this disclosure, both the input data table and the output data table include a database name and a data table name; determining the first hash value based on the first entity corresponding to the task includes:

[0154] Based on preset rules, the database name and the data table name in the first entity are concatenated to obtain the first concatenation information;

[0155] The first hash value is determined based on the first concatenation information;

[0156] Determining the second hash value based on the second entity corresponding to the task includes:

[0157] Based on the preset rules, the database name and the data table name in the second entity are concatenated to obtain the second concatenation information;

[0158] The second hash value is determined based on the second concatenation information.

[0159] Since both the input and output data tables contain database names and table names, when determining the hash value, the database name and the table name can be concatenated first, and then the hash value can be obtained based on the concatenated information. Thus, when comparing the first hash value and the second hash value, multiple database names and multiple table names are compared simultaneously.

[0160] Obtaining the hash value by concatenating the database name and table name can improve the accuracy of the generated hash value.

[0161] like Figure 6 As shown, the first and second data tables can be associated in three ways based on the task identifier. Through inline, a third task can be obtained where the parsing results of the lineage analysis tool and the third-party lineage analysis tool are consistent. Through left and right association, a fourth and fifth task can be obtained where the parsing results of the lineage analysis tool and the third-party lineage analysis tool are inconsistent.

[0162] In one exemplary embodiment of this disclosure, in response to the first hash value in the first data table and the second hash value in the second data table being the same, the task corresponding to the first entity or the second entity is determined as the third task, including:

[0163] The first data table and the second data table are inlined according to the task identifier to obtain a first output data table, and the number of the third tasks is determined according to the first output table; the third task is a task whose first hash value and the second hash value are the same.

[0164] When comparing the first hash value in the first data table and the second hash value in the second data table, the first data table and the second data table can be inlined. The first output data table stores tasks with the same first hash value and second hash value under the same task identifier. The tasks in the first output data table are identified as the third tasks, thereby obtaining the number of third tasks.

[0165] By performing an inline operation on the two data tables, it is easy to obtain tasks where the first hash value and the second hash value are the same.

[0166] In one exemplary embodiment of this disclosure, the first data table is a left table; the second data table is a right table; the method further includes:

[0167] The first data table and the second data table are left-associated according to the task identifier to obtain the second output data table. The fourth task is selected from the second output data table if the first hash value is not empty and the second hash value is empty. The number of the fourth tasks is determined. The fourth task represents the task for which the bloodline analysis tool has obtained the bloodline relationship but the third-party bloodline analysis tool has not obtained the bloodline relationship.

[0168] The proportion of the first task to be analyzed is determined based on the number of the fourth task and the first number.

[0169] When performing a left join, the left and right tables are first determined. Here, the left table is the first data table, and the output of the left join is the second output data table. The second output data table is based on the left table and joins the two tables according to the task identifier. This means listing all query information from the left table, listing the portions of the right table where the second hash value matches the first hash value, and setting the portions where the second hash value does not match the first hash value to null. Therefore, after obtaining the second output table, a fourth task can be filtered out where the first hash value is not null and the second hash value is null. This part of the task represents tasks where the second hash value obtained by the third-party lineage analysis tool is inconsistent with the first lineage analysis tool, based on the first lineage relationship.

[0170] By determining the number of fourth tasks and comparing them to the number of first tasks, the proportion of the first tasks to be analyzed can be further determined. This proportion directly indicates the number of tasks where the second lineage relationship analyzed by the third-party lineage analysis tool is inconsistent with the first lineage relationship.

[0171] By performing a left join between the first and second data tables, the fourth task can be quickly obtained, thereby determining the proportion of the first task to be analyzed.

[0172] In one exemplary embodiment of this disclosure, the method further includes:

[0173] The first data table and the second data table are right-associated according to the task identifier to obtain a third output data table. The third output data table is used to filter out the fifth tasks where the second hash value is not empty and the first hash value is empty, and the number of the fifth tasks is determined. The fifth task represents a task where the bloodline analysis tool did not obtain the bloodline relationship but the third-party bloodline analysis tool obtained the bloodline relationship.

[0174] The proportion of the second task to be analyzed is determined based on the number of the fifth task and the first task.

[0175] Similarly, a right join can be performed on the first and second data tables to obtain the fifth task. This task represents tasks where the first and second bloodline relationships analyzed by the third-party bloodline analysis tool are inconsistent, based on the second bloodline relationship. This allows for further determination of the proportion of the second task to be analyzed. The proportion of the second task to be analyzed directly indicates the number of tasks where the first and second bloodline relationships analyzed by the bloodline analysis tool are inconsistent.

[0176] By performing a right join between the first and second data tables, the fifth task can be quickly obtained, thereby determining the proportion of the second task to be analyzed.

[0177] In one exemplary embodiment of this disclosure, the method further includes:

[0178] Obtain the task details of the fourth task corresponding to the proportion of the first task to be analyzed, and / or obtain the task details of the fifth task corresponding to the proportion of the second task to be analyzed;

[0179] The task details are analyzed to identify and fix any faults in the lineage analysis tool.

[0180] After determining the fourth and fifth tasks, the task details of the fourth and fifth tasks can be analyzed to find any faults in the lineage analysis tool and fix them accordingly, thereby improving the accuracy of lineage analysis.

[0181] By analyzing the task details of the fourth and fifth tasks, problems with the lineage analysis tool can be identified. By fixing these problems, the accuracy of lineage analysis can be improved.

[0182] In addition to the methods for determining kinship coverage and kinship accuracy mentioned above, a method for continuously evaluating kinship quality can also be designed to output the daily kinship coverage and kinship accuracy, ensuring that problems can be identified in a timely and proactive manner, and continuously and rapidly improving the kinship quality of the data governance platform. Figure 7 A schematic diagram illustrating continuous output of bloodline coverage according to an embodiment of the present disclosure is shown. Figure 8 A schematic diagram illustrating a continuous output of lineage accuracy according to an embodiment of the present disclosure is shown; Reference Figure 7 and Figure 8 The method for continuously outputting kinship coverage and kinship accuracy is explained.

[0183] In one exemplary embodiment of this disclosure, the method further includes:

[0184] Every preset time interval, the first quantity corresponding to the preset time interval is stored in the first offline table of the big data platform, and the second quantity corresponding to the preset time interval is stored in the second offline table of the big data platform;

[0185] Determining the lineage coverage of the data governance platform based on the first quantity and the second quantity includes:

[0186] The first quantity in the first offline table and the second quantity in the second offline table are obtained through the offline scheduling service. Based on the first quantity and the second quantity, the bloodline coverage rate corresponding to the preset time period is determined and stored in the third offline table.

[0187] like Figure 7As shown, in order to continuously evaluate the lineage quality of the data governance platform, the first quantity and the second quantity can be stored in the first offline table and the second offline table of the big data platform, respectively. Then, the ratio of the second quantity to the first quantity can be calculated through the offline scheduling service to obtain the lineage coverage rate.

[0188] Optionally, tasks executed at the platform layer can be stored in a first offline table via data transmission. The lineage relationships in the lineage storage layer can be stored in a MySQL table via a timed thread. The lineage relationships in the MySQL table can then be stored in a second offline table via data transmission. The offline scheduling service can then filter the tasks stored in the first and second offline tables according to the user's needs to obtain the first and second quantities, thereby calculating the lineage coverage rate and storing it in a third offline table. Finally, the calculated lineage coverage rate can be output.

[0189] By transferring tasks from the lineage storage layer and platform layer to the offline scheduling table, or by transferring the first and second quantities to the offline scheduling table, and by continuously calculating lineage coverage through the offline scheduling service, the lineage quality of the data governance platform can be continuously evaluated.

[0190] In one exemplary embodiment of this disclosure, the method further includes:

[0191] The bloodline coverage rate stored in the third offline table is uploaded to the business intelligence reporting system and / or the data dashboard system to display the bloodline coverage rate through the business intelligence reporting system and / or the data dashboard system.

[0192] After determining the lineage coverage rate and storing it in the third offline table, the lineage coverage rate can also be uploaded to the business intelligence reporting system and / or data dashboard system, which can be displayed to users or developers, allowing them to intuitively view the lineage coverage rate of the dimensions they are interested in.

[0193] In addition, such as Figure 7 As shown, while continuously outputting lineage coverage, it is also possible to identify unrelated tasks and introduce a basic lineage coverage set to obtain the type of each unrelated task, so as to rectify the lineage analysis tool or lineage consumption chain and continuously improve the lineage coverage.

[0194] In one exemplary embodiment of this disclosure, the method further includes:

[0195] Store the first entity corresponding to each task in the fourth offline table, and store the second entity corresponding to each task in the fifth offline table;

[0196] Determining the number of third tasks in the lineage storage layer includes:

[0197] The first entity in the fourth offline table and the second entity in the fifth offline table are obtained through the offline scheduling service.

[0198] If the first entity and the second entity are the same, then the task corresponding to the first entity or the second entity is determined as the third task, and the number of the third tasks is determined.

[0199] refer to Figure 8 For the parsing plan, a third-party lineage resolution tool can be used to parse it, obtaining a second entity consisting of a task identifier, an input data table, and an output data table. Furthermore, the parsing plan can be parsed according to the lineage resolution tool currently used by the data governance platform.

[0200] The first entity is obtained, and the first and second entities are stored in MySQL table 1 and MySQL table 2 respectively. The first and second entities are then transferred to the fourth and fifth offline tables respectively using data transfer methods. The number of third tasks can be calculated through offline scheduling. Alternatively, the number of fourth and fifth tasks can also be calculated to calculate the bleeding accuracy, the proportion of the first task to be analyzed, and the proportion of the second task to be analyzed.

[0201] By storing the first and second entities in the fourth and fifth offline tables respectively, and through the offline scheduling service, results such as lineage accuracy can be continuously output, ensuring the ability to continuously improve lineage accuracy and thus improve lineage quality.

[0202] This disclosure, starting from the perspectives of users and pedigree analysis tool developers, designs a pedigree coverage algorithm based on the platform layer and pedigree storage layer, which can assess the pedigree quality of the data governance platform to a certain extent.

[0203] Furthermore, by comparing tasks executed at the platform layer with those obtained from the lineage storage layer, tasks that may lack lineage can be identified. These tasks can then be analyzed manually or using automated tools to quickly and definitively determine whether lineage needs to be established. For lineage scenarios requiring lineage support, lineage analysis tool developers can proactively adapt and support them, improving lineage coverage. This solution can also be integrated into a big data platform for daily scheduling, ensuring continuous detection of scenarios requiring lineage support and continuously improving lineage coverage.

[0204] In addition, this disclosure defines a lineage accuracy metric. By obtaining the parsing plan of the platform layer, a third-party lineage parsing tool is introduced for parsing. Entities with the same structure as the lineage parsing tool used in the data governance platform are obtained for comparison. The lineage accuracy, the proportion of the first task to be analyzed, and the proportion of the second task to be analyzed are obtained and analyzed. The lineage is scheduled on a daily basis to improve the lineage accuracy.

[0205] Exemplary media

[0206] After introducing the methods of exemplary embodiments of this disclosure, the following references are made. Figure 9 The storage medium of the exemplary embodiments of this disclosure will be described.

[0207] refer to Figure 9 As shown, the storage medium 90 stores a program product for implementing the above-described method according to an embodiment of the present disclosure. This program product may be a portable compact disc read-only memory (CD-ROM) and includes program code for causing a computing device to execute the data processing method provided in this disclosure. However, the program product of this disclosure is not limited thereto.

[0208] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0209] A readable signal medium may include data signals propagated in baseband or as part of a carrier wave, carrying program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium.

[0210] Program code for performing the operations disclosed herein can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN).

[0211] Exemplary device

[0212] Having introduced the medium of exemplary embodiments of this disclosure, the following references are made to... Figure 10 The data lineage analysis apparatus of the exemplary embodiments of this disclosure will be described to implement the method in any of the above-described data lineage analysis method embodiments. The implementation principle and technical effect are similar, and will not be repeated here.

[0213] The data lineage analysis device 100 provided in this disclosure is applied to a data governance platform. The data governance platform includes a platform layer and a lineage storage layer. The platform layer is used to execute tasks within the data governance platform, and the lineage storage layer is used to extract and store lineage relationships based on the parsing plan corresponding to the task. The lineage relationship is the relationship between an input table and an output table marked by the task. The device 100 includes:

[0214] The acquisition module 1001 is used to acquire a first number of tasks executed by the platform layer; and to acquire a second number of tasks corresponding to blood relations stored in the blood relation storage layer.

[0215] The first determining module 1002 is used to determine the lineage coverage rate of the data governance platform based on the first quantity and the second quantity.

[0216] The device further includes:

[0217] The comparison module is used to, in response to the bloodline coverage rate being less than the target bloodline coverage rate, compare the tasks stored in the bloodline storage layer with the tasks executed by the platform layer according to the target attributes to obtain unrelated tasks; the unrelated tasks include a first task that has been executed by the platform layer but not stored in the bloodline storage layer, and a second task that has been stored in the bloodline storage layer but not executed by the platform layer;

[0218] The second determining module is used to determine the type of the unrelated task based on the unrelated task and the basic lineage coverage set; the basic lineage coverage set contains multiple target SQL statement types; wherein, the SQL statement corresponding to the target SQL statement type has data flow.

[0219] In one exemplary embodiment of this disclosure, the apparatus further includes:

[0220] The comparison module is used to, in response to the bloodline coverage rate being less than the target bloodline coverage rate, compare the tasks stored in the bloodline storage layer with the tasks executed by the platform layer according to the target attributes to obtain unrelated tasks; the unrelated tasks include a first task that has been executed by the platform layer but not stored in the bloodline storage layer, and a second task that has been stored in the bloodline storage layer but not executed by the platform layer;

[0221] The second determining module is used to determine the type of the unrelated task based on the unrelated task and the basic lineage coverage set; the basic lineage coverage set contains multiple target SQL statement types; wherein, the SQL statement corresponding to the target SQL statement type has data flow.

[0222] In one exemplary embodiment of this disclosure, the target attribute includes the name of the source platform, the task identifier, and the execution ID; the source platform represents the source information of the task corresponding to the target attribute; the task identifier is used to distinguish different tasks; and the execution ID is used to identify each execution of the same task.

[0223] In one exemplary embodiment of this disclosure, the apparatus further includes:

[0224] The execution module is used to perform corresponding target operations according to the type of the unrelated task, so as to improve the lineage coverage of the data governance platform.

[0225] In one exemplary embodiment of this disclosure, the second determining module is specifically used for:

[0226] If the type of at least one SQL statement corresponding to the first task exists in the basic lineage coverage set, then the first task is determined to be a first type of unrelated task.

[0227] The execution module is specifically used for:

[0228] In response to the first task being a first type of unrelated task, it is determined that there is an anomaly in the lineage resolution tool and / or the lineage consumption link; the lineage consumption link is used to represent the path through which the lineage storage layer consumes the lineage relationship; the lineage resolution tool is used to extract the lineage relationship of the task based on the resolution plan of the task.

[0229] In one exemplary embodiment of this disclosure, the second determining module is specifically used for:

[0230] If the types of all SQL statements corresponding to the first task do not exist in the basic lineage coverage set, and the first task has data flow, then the first task is determined to be a second type of unrelated task.

[0231] The execution module is specifically used for:

[0232] In response to the first task being the second type of unrelated task, a scenario enhancement operation is performed on the lineage analysis tool; the scenario enhancement is used to increase the lineage analysis tool's ability to correctly parse preset SQL statement types.

[0233] In one exemplary embodiment of this disclosure, the second determining module is specifically used for:

[0234] If the types of all SQL statements corresponding to the first task do not exist in the basic lineage coverage set, and the first task does not have data flow, then the first task is determined to be a third type of unrelated task.

[0235] The execution module is specifically used for:

[0236] In response to the first task being the third type of unrelated task, a prompt message indicating that the third type of unrelated task will not be executed is output; and / or, the first quantity is corrected, and the lineage coverage rate is determined based on the corrected quantity; the corrected quantity is the difference between the first quantity and the number of tasks corresponding to the third type of unrelated task.

[0237] In one exemplary embodiment of this disclosure, the apparatus further includes:

[0238] The third determining module is used to determine the second task as a fourth type of unrelated task in response to the fact that the unrelated tasks include the second task.

[0239] The execution module is specifically used for:

[0240] If the second task is the fourth type of unrelated task, then the kinship analysis tool is determined to be faulty.

[0241] In one exemplary embodiment of this disclosure, the apparatus further includes:

[0242] The fourth determining module is used to determine the number of third tasks in the bloodline storage layer; the third task is the task of obtaining the correct bloodline relationship through bloodline analysis tools.

[0243] The fifth determining module is used to determine the lineage accuracy of the data governance platform based on the first quantity and the number of the third task.

[0244] In an exemplary embodiment of this disclosure, the bloodline analysis tool extracts a first bloodline relationship; the fourth determining module is specifically used for:

[0245] A third-party lineage analysis tool is used to extract the second lineage relationship corresponding to the analysis plan of the task; the third-party lineage analysis tool is used to verify the correctness of the first lineage relationship extracted by the lineage analysis tool.

[0246] A first entity is generated based on a first lineage relationship corresponding to the task, and a second entity is generated based on a second lineage relationship corresponding to the task; the first entity includes a task identifier, an input data table, and an output data table; the second entity includes a task identifier, an input data table, and an output data table.

[0247] In response to the first entity and the second entity being the same, the task corresponding to the first entity or the second entity is determined as the third task, and the number of the third tasks is determined; when the task identifier in the first entity is the same as the task identifier in the second entity, and at the same time, the input data table in the first entity is the same as the input data table in the second entity, and at the same time, the output data table in the first entity is the same as the output data table in the second entity, then the first entity and the second entity are the same.

[0248] In one exemplary embodiment of this disclosure, the apparatus further includes: a storage module, configured to:

[0249] A first hash value is determined based on the first entity corresponding to the task, and the first hash value and the task identifier corresponding to the task are stored in a first data table;

[0250] Determine a second hash value based on the second entity corresponding to the task, and store the second hash value and the task identifier corresponding to the task in a second data table;

[0251] When the fourth determining module determines the task corresponding to the first entity or the second entity as the third task in response to the first entity and the second entity being the same, it is specifically used for:

[0252] If the first hash value in the first data table is the same as the second hash value in the second data table, then the task corresponding to the first entity or the second entity is determined as the third task.

[0253] In an exemplary embodiment of this disclosure, when the fourth determining module determines the third task by matching the first hash value in the first data table with the second hash value in the second data table, it is specifically used for:

[0254] The first data table and the second data table are inlined according to the task identifier to obtain a first output data table, and the number of the third tasks is determined according to the first output table; the third task is a task whose first hash value and the second hash value are the same.

[0255] In an exemplary embodiment of this disclosure, the first data table is a left table; the second data table is a right table; the apparatus further includes: a sixth determining module, configured to:

[0256] The first data table and the second data table are left-associated according to the task identifier to obtain the second output data table. The fourth task is selected from the second output data table if the first hash value is not empty and the second hash value is empty. The number of the fourth tasks is determined. The fourth task represents the task for which the bloodline analysis tool has obtained the bloodline relationship but the third-party bloodline analysis tool has not obtained the bloodline relationship.

[0257] The proportion of the first task to be analyzed is determined based on the number of the fourth task and the first number.

[0258] In one exemplary embodiment of this disclosure, the sixth determining module is further configured to:

[0259] The first data table and the second data table are right-associated according to the task identifier to obtain a third output data table. The third output data table is used to filter out the fifth tasks where the second hash value is not empty and the first hash value is empty, and the number of the fifth tasks is determined. The fifth task represents a task where the bloodline analysis tool did not obtain the bloodline relationship but the third-party bloodline analysis tool obtained the bloodline relationship.

[0260] The proportion of the second task to be analyzed is determined based on the number of the fifth task and the first task.

[0261] In an exemplary embodiment of this disclosure, both the input data table and the output data table include a database name and a data table name; when the storage module determines the first hash value based on the first entity corresponding to the task, it is specifically used for:

[0262] Based on preset rules, the database name and the data table name in the first entity are concatenated to obtain the first concatenation information;

[0263] The first hash value is determined based on the first concatenation information;

[0264] When determining the second hash value based on the second entity corresponding to the task, the storage module is specifically used for:

[0265] Based on the preset rules, the database name and the data table name in the second entity are concatenated to obtain the second concatenation information;

[0266] The second hash value is determined based on the second concatenation information.

[0267] In one exemplary embodiment of this disclosure, the apparatus further includes an analysis module for:

[0268] Obtain the task details of the fourth task corresponding to the proportion of the first task to be analyzed, and / or obtain the task details of the fifth task corresponding to the proportion of the second task to be analyzed;

[0269] The task details are analyzed to identify and fix any faults in the lineage analysis tool.

[0270] In one exemplary embodiment of this disclosure, the apparatus further includes: a first processing module, configured to:

[0271] Every preset time interval, the first quantity corresponding to the preset time interval is stored in the first offline table of the big data platform, and the second quantity corresponding to the preset time interval is stored in the second offline table of the big data platform;

[0272] The first determining module 1002 is specifically used for:

[0273] The first quantity in the first offline table and the second quantity in the second offline table are obtained through the offline scheduling service. Based on the first quantity and the second quantity, the bloodline coverage rate corresponding to the preset time period is determined and stored in the third offline table.

[0274] In one exemplary embodiment of this disclosure, the apparatus further includes: a display module, configured to:

[0275] The bloodline coverage rate stored in the third offline table is uploaded to the business intelligence reporting system and / or the data dashboard system to display the bloodline coverage rate through the business intelligence reporting system and / or the data dashboard system.

[0276] In one exemplary embodiment of this disclosure, the apparatus further includes: a second processing module, configured to:

[0277] Store the first entity corresponding to each task in the fourth offline table, and store the second entity corresponding to each task in the fifth offline table;

[0278] The fourth determining module is specifically used for:

[0279] The first entity in the fourth offline table and the second entity in the fifth offline table are obtained through the offline scheduling service.

[0280] If the first entity and the second entity are the same, then the task corresponding to the first entity or the second entity is determined as the third task, and the number of the third tasks is determined.

[0281] Exemplary computing device

[0282] Having described the methods, media, and apparatus of exemplary embodiments of this disclosure, the following references... Figure 11 A computing device according to an exemplary embodiment of the present disclosure will be described.

[0283] Figure 11 The computing device 110 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0284] like Figure 11 As shown, the computing device 110 is presented in the form of a general-purpose computing device. The components of the computing device 110 may include, but are not limited to: at least one processing unit 1101, at least one storage unit 1102, and a bus 1103 connecting different system components (including the processing unit 1101 and the storage unit 1102). The at least one storage unit 1102 stores computer-executable instructions; the at least one processing unit 1101 includes a processor that executes the computer-executable instructions to implement the methods described above.

[0285] Bus 1103 includes a data bus, a control bus, and an address bus.

[0286] Storage unit 1102 may include readable media in the form of volatile memory, such as random access memory (RAM) 11021 and / or cache memory 11022, and may further include readable media in the form of non-volatile memory, such as read-only memory (ROM) 11023.

[0287] Storage unit 1102 may also include a program / utility 11025 having a set (at least one) of program modules 11024, such program modules 11024 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0288] The computing device 110 can also communicate with one or more external devices 1104 (e.g., keyboard, pointing device, etc.). This communication can be performed via input / output (I / O) interface 1105. Furthermore, the computing device 110 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1106. Figure 11 As shown, network adapter 1106 communicates with other modules of computing device 110 via bus 1103. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with computing device 110, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0289] It should be noted that although the data processing apparatus and several units / modules or sub-units / modules of the data processing apparatus have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0290] Furthermore, although the operations of the methods disclosed herein are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0291] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A data lineage analysis method, characterized in that, This is applied to a data governance platform, which includes a platform layer and a lineage storage layer. The platform layer is used to execute tasks within the data governance platform, and the lineage storage layer is used to extract and store lineage relationships based on the parsing plan corresponding to the task. The lineage relationship is the relationship between the input table and the output table marked by the task. The method includes: Obtain the first number of tasks executed by the platform layer; obtain the second number of tasks corresponding to blood relations stored in the blood relation storage layer; The lineage coverage rate of the data governance platform is determined based on the ratio of the second quantity to the first quantity.

2. The method according to claim 1, characterized in that, The method further includes: In response to the bloodline coverage rate being less than the target bloodline coverage rate, the tasks stored in the bloodline storage layer and the tasks executed by the platform layer are compared according to the target attribute to obtain unrelated tasks; the unrelated tasks include a first task that has been executed by the platform layer but not stored in the bloodline storage layer, and a second task that has been stored in the bloodline storage layer but not executed by the platform layer. The type of the unrelated task is determined based on the unrelated task and the basic lineage coverage set; the basic lineage coverage set contains multiple target SQL statement types; wherein, the SQL statement corresponding to the target SQL statement type has data flow.

3. The method according to claim 2, characterized in that, The target attribute includes the name of the source platform, the task identifier, and the execution ID; the source platform represents the source information of the task corresponding to the target attribute; the task identifier is used to distinguish different tasks; and the execution ID is used to identify each execution of the same task.

4. The method according to claim 2, characterized in that, The method further includes: Perform corresponding target operations based on the type of the unrelated task to improve the lineage coverage of the data governance platform.

5. The method according to claim 4, characterized in that, The type of the unrelated task is determined based on the unrelated task and the basic lineage coverage set, including: If the type of at least one SQL statement corresponding to the first task exists in the basic lineage coverage set, then the first task is determined to be a first type of unrelated task. Perform the corresponding target operation according to the type of the unrelated task, including: In response to the first task being a first type of unrelated task, it is determined that there is an anomaly in the lineage resolution tool and / or the lineage consumption link; the lineage consumption link is used to represent the path through which the lineage storage layer consumes the lineage relationship; the lineage resolution tool is used to extract the lineage relationship of the task based on the resolution plan of the task.

6. The method according to claim 4, characterized in that, The type of the unrelated task is determined based on the unrelated task and the basic lineage coverage set, including: If the types of all SQL statements corresponding to the first task do not exist in the basic lineage coverage set, and the first task has data flow, then the first task is determined to be a second type of unrelated task. Perform the corresponding target operation according to the type of the unrelated task, including: In response to the first task being the second type of unrelated task, a scenario enhancement operation is performed on the lineage resolution tool; the scenario enhancement is used to increase the lineage resolution tool's ability to correctly parse preset SQL statement types.

7. The method according to claim 4, characterized in that, The type of the unrelated task is determined based on the unrelated task and the basic lineage coverage set, including: If the types of all SQL statements corresponding to the first task do not exist in the basic lineage coverage set, and the first task does not have data flow, then the first task is determined to be a third type of unrelated task. Perform the corresponding target operation according to the type of the unrelated task, including: In response to the first task being the third type of unrelated task, a prompt message indicating that the third type of unrelated task will not be executed is output; and / or, the first quantity is corrected, and the lineage coverage rate is determined based on the corrected quantity; the corrected quantity is the difference between the first quantity and the number of tasks corresponding to the third type of unrelated task.

8. The method according to claim 4, characterized in that, The method further includes: If the second task is included in the unrelated tasks, then the second task is determined to be a fourth type of unrelated task; Perform the corresponding target operation according to the type of the unrelated task, including: If the second task is the fourth type of unrelated task, then the kinship analysis tool is determined to be faulty.

9. The method according to claim 1, characterized in that, The method further includes: Determine the number of third tasks in the bloodline storage layer; the third task is the task of obtaining the correct bloodline relationship through a bloodline resolution tool. The lineage accuracy of the data governance platform is determined based on the first quantity and the third task quantity.

10. The method according to claim 9, characterized in that, The bloodline analysis tool extracts the first bloodline relationship; Determining the number of third tasks in the lineage storage layer includes: A third-party lineage analysis tool is used to extract the second lineage relationship corresponding to the analysis plan of the task; the third-party lineage analysis tool is used to verify the correctness of the first lineage relationship extracted by the lineage analysis tool. A first entity is generated based on a first lineage relationship corresponding to the task, and a second entity is generated based on a second lineage relationship corresponding to the task; the first entity includes a task identifier, an input data table, and an output data table; the second entity includes a task identifier, an input data table, and an output data table. In response to the first entity and the second entity being the same, the task corresponding to the first entity or the second entity is determined as the third task, and the number of the third tasks is determined; when the task identifier in the first entity is the same as the task identifier in the second entity, and at the same time, the input data table in the first entity is the same as the input data table in the second entity, and at the same time, the output data table in the first entity is the same as the output data table in the second entity, then the first entity and the second entity are the same.

11. The method according to claim 10, characterized in that, The method further includes: A first hash value is determined based on the first entity corresponding to the task, and the first hash value and the task identifier corresponding to the task are stored in a first data table; Determine a second hash value based on the second entity corresponding to the task, and store the second hash value and the task identifier corresponding to the task in a second data table; In response to the first entity and the second entity being the same, the task corresponding to the first entity or the second entity is determined as the third task, including: If the first hash value in the first data table is the same as the second hash value in the second data table, then the task corresponding to the first entity or the second entity is determined as the third task.

12. The method according to claim 11, characterized in that, In response to the first hash value in the first data table and the second hash value in the second data table being the same, the task corresponding to the first entity or the second entity is determined as the third task, including: The first data table and the second data table are inlined according to the task identifier to obtain a first output data table, and the number of the third tasks is determined according to the first output table; the third task is a task whose first hash value and the second hash value are the same.

13. The method according to claim 11, characterized in that, The first data table is the left table; the second data table is the right table; the method further includes: The first data table and the second data table are left-associated according to the task identifier to obtain the second output data table. The fourth task is selected from the second output data table if the first hash value is not empty and the second hash value is empty. The number of the fourth tasks is determined. The fourth task represents the task for which the bloodline analysis tool has obtained the bloodline relationship but the third-party bloodline analysis tool has not obtained the bloodline relationship. The proportion of the first task to be analyzed is determined based on the number of the fourth task and the first number.

14. The method according to claim 13, characterized in that, The method further includes: The first data table and the second data table are right-associated according to the task identifier to obtain a third output data table. The third output data table is used to filter out the fifth tasks where the second hash value is not empty and the first hash value is empty, and the number of the fifth tasks is determined. The fifth task represents a task where the bloodline analysis tool did not obtain the bloodline relationship but the third-party bloodline analysis tool obtained the bloodline relationship. The proportion of the second task to be analyzed is determined based on the number of the fifth task and the first task.

15. The method according to claim 11, characterized in that, Both the input data table and the output data table contain a database name and a data table name; Determining the first hash value based on the first entity corresponding to the task includes: Based on preset rules, the database name and the data table name in the first entity are concatenated to obtain the first concatenation information; The first hash value is determined based on the first concatenation information; Determining the second hash value based on the second entity corresponding to the task includes: Based on the preset rules, the database name and the data table name in the second entity are concatenated to obtain the second concatenation information; The second hash value is determined based on the second concatenation information.

16. The method according to claim 14, characterized in that, The method further includes: Obtain the task details of the fourth task corresponding to the proportion of the first task to be analyzed, and / or obtain the task details of the fifth task corresponding to the proportion of the second task to be analyzed; The task details are analyzed to identify and fix any faults in the lineage analysis tool.

17. The method according to any one of claims 1-8, characterized in that, The method further includes: Every preset time interval, the first quantity corresponding to the preset time interval is stored in the first offline table of the big data platform, and the second quantity corresponding to the preset time interval is stored in the second offline table of the big data platform; Determining the lineage coverage of the data governance platform based on the first quantity and the second quantity includes: The first quantity in the first offline table and the second quantity in the second offline table are obtained through the offline scheduling service. Based on the first quantity and the second quantity, the bloodline coverage rate corresponding to the preset time period is determined and stored in the third offline table.

18. The method according to claim 17, characterized in that, The method further includes: The bloodline coverage rate stored in the third offline table is uploaded to the business intelligence reporting system and / or the data dashboard system to display the bloodline coverage rate through the business intelligence reporting system and / or the data dashboard system.

19. The method according to any one of claims 10-16, characterized in that, The method further includes: Store the first entity corresponding to each task in the fourth offline table, and store the second entity corresponding to each task in the fifth offline table; Determining the number of third tasks in the lineage storage layer includes: The first entity in the fourth offline table and the second entity in the fifth offline table are obtained through the offline scheduling service. If the first entity and the second entity are the same, then the task corresponding to the first entity or the second entity is determined as the third task, and the number of the third tasks is determined.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1 to 19.

21. A data lineage analysis device, characterized in that, This is applied to a data governance platform, which includes a platform layer and a lineage storage layer. The platform layer is used to execute tasks within the data governance platform, and the lineage storage layer is used to extract and store lineage relationships based on the parsing plan corresponding to the task. The lineage relationship is the relationship between the input table and the output table marked by the task. The device includes: The acquisition module is used to acquire a first number of tasks executed by the platform layer; and to acquire a second number of tasks corresponding to blood relations stored in the blood relation storage layer. The first determining module is used to determine the lineage coverage rate of the data governance platform based on the ratio of the second quantity to the first quantity.

22. The apparatus according to claim 21, characterized in that, The device further includes: The comparison module is used to, in response to the bloodline coverage rate being less than the target bloodline coverage rate, compare the tasks stored in the bloodline storage layer with the tasks executed by the platform layer according to the target attributes to obtain unrelated tasks; the unrelated tasks include a first task that has been executed by the platform layer but not stored in the bloodline storage layer, and a second task that has been stored in the bloodline storage layer but not executed by the platform layer; The second determining module is used to determine the type of the unrelated task based on the unrelated task and the basic lineage coverage set; the basic lineage coverage set contains multiple target SQL statement types; wherein, the SQL statement corresponding to the target SQL statement type has data flow.

23. The apparatus according to claim 22, characterized in that, The target attribute includes the name of the source platform, the task identifier, and the execution ID; the source platform represents the source information of the task corresponding to the target attribute; the task identifier is used to distinguish different tasks; and the execution ID is used to identify each execution of the same task.

24. The apparatus according to claim 22, characterized in that, The device further includes: The execution module is used to perform corresponding target operations according to the type of the unrelated task, so as to improve the lineage coverage of the data governance platform.

25. The apparatus according to claim 24, characterized in that, The second determining module is specifically used for: If the type of at least one SQL statement corresponding to the first task exists in the basic lineage coverage set, then the first task is determined to be a first type of unrelated task. The execution module is specifically used for: In response to the first task being a first type of unrelated task, it is determined that there is an anomaly in the lineage resolution tool and / or the lineage consumption link; the lineage consumption link is used to represent the path through which the lineage storage layer consumes the lineage relationship; the lineage resolution tool is used to extract the lineage relationship of the task based on the resolution plan of the task.

26. The apparatus according to claim 24, characterized in that, The second determining module is specifically used for: If the types of all SQL statements corresponding to the first task do not exist in the basic lineage coverage set, and the first task has data flow, then the first task is determined to be a second type of unrelated task. The execution module is specifically used for: In response to the first task being the second type of unrelated task, a scenario enhancement operation is performed on the lineage resolution tool; the scenario enhancement is used to increase the lineage resolution tool's ability to correctly parse preset SQL statement types.

27. The apparatus according to claim 24, characterized in that, The second determining module is specifically used for: If the types of all SQL statements corresponding to the first task do not exist in the basic lineage coverage set, and the first task does not have data flow, then the first task is determined to be a third type of unrelated task. The execution module is specifically used for: In response to the first task being the third type of unrelated task, a prompt message indicating that the third type of unrelated task will not be executed is output; and / or, the first quantity is corrected, and the lineage coverage rate is determined based on the corrected quantity; the corrected quantity is the difference between the first quantity and the number of tasks corresponding to the third type of unrelated task.

28. The apparatus according to claim 24, characterized in that, The device further includes: The third determining module is used to determine the second task as a fourth type of unrelated task in response to the fact that the unrelated tasks include the second task. The execution module is specifically used for: If the second task is the fourth type of unrelated task, then the kinship analysis tool is determined to be faulty.

29. The apparatus according to claim 21, characterized in that, The device further includes: The fourth determining module is used to determine the number of third tasks in the bloodline storage layer; the third task is the task of obtaining the correct bloodline relationship through bloodline analysis tools. The fifth determining module is used to determine the lineage accuracy of the data governance platform based on the first quantity and the number of the third task.

30. The apparatus according to claim 29, characterized in that, The bloodline analysis tool extracts the first bloodline relationship; the fourth determination module is specifically used for: A third-party lineage analysis tool is used to extract the second lineage relationship corresponding to the analysis plan of the task; the third-party lineage analysis tool is used to verify the correctness of the first lineage relationship extracted by the lineage analysis tool. A first entity is generated based on a first lineage relationship corresponding to the task, and a second entity is generated based on a second lineage relationship corresponding to the task; the first entity includes a task identifier, an input data table, and an output data table; the second entity includes a task identifier, an input data table, and an output data table. In response to the first entity and the second entity being the same, the task corresponding to the first entity or the second entity is determined as the third task, and the number of the third tasks is determined; when the task identifier in the first entity is the same as the task identifier in the second entity, and at the same time, the input data table in the first entity is the same as the input data table in the second entity, and at the same time, the output data table in the first entity is the same as the output data table in the second entity, then the first entity and the second entity are the same.

31. The apparatus according to claim 30, characterized in that, The device further includes: a storage module, used for: A first hash value is determined based on the first entity corresponding to the task, and the first hash value and the task identifier corresponding to the task are stored in a first data table; Determine a second hash value based on the second entity corresponding to the task, and store the second hash value and the task identifier corresponding to the task in a second data table; When the fourth determining module determines the task corresponding to the first entity or the second entity as the third task in response to the first entity and the second entity being the same, it is specifically used for: If the first hash value in the first data table is the same as the second hash value in the second data table, then the task corresponding to the first entity or the second entity is determined as the third task.

32. The apparatus according to claim 31, characterized in that, When the fourth determining module determines the third task based on the first hash value in the first data table and the second hash value in the second data table being the same, it is specifically used for: The first data table and the second data table are inlined according to the task identifier to obtain a first output data table, and the number of the third tasks is determined according to the first output table; the third task is a task whose first hash value and the second hash value are the same.

33. The apparatus according to claim 31, characterized in that, The first data table is the left table; the second data table is the right table; the device further includes: a sixth determining module, used for: The first data table and the second data table are left-associated according to the task identifier to obtain the second output data table. The fourth task is selected from the second output data table if the first hash value is not empty and the second hash value is empty. The number of the fourth tasks is determined. The fourth task represents the task for which the bloodline analysis tool has obtained the bloodline relationship but the third-party bloodline analysis tool has not obtained the bloodline relationship. The proportion of the first task to be analyzed is determined based on the number of the fourth task and the first number.

34. The apparatus according to claim 33, characterized in that, The sixth determining module is also used for: The first data table and the second data table are right-associated according to the task identifier to obtain a third output data table. The third output data table is used to filter out the fifth tasks where the second hash value is not empty and the first hash value is empty, and the number of the fifth tasks is determined. The fifth task represents a task where the bloodline analysis tool did not obtain the bloodline relationship but the third-party bloodline analysis tool obtained the bloodline relationship. The proportion of the second task to be analyzed is determined based on the number of the fifth task and the first task.

35. The apparatus according to claim 31, characterized in that, Both the input data table and the output data table contain a database name and a data table name; when the storage module determines the first hash value based on the first entity corresponding to the task, it is specifically used for: Based on preset rules, the database name and the data table name in the first entity are concatenated to obtain the first concatenation information; The first hash value is determined based on the first concatenation information; When determining the second hash value based on the second entity corresponding to the task, the storage module is specifically used for: Based on the preset rules, the database name and the data table name in the second entity are concatenated to obtain the second concatenation information; The second hash value is determined based on the second concatenation information.

36. The apparatus according to claim 34, characterized in that, The device further includes an analysis module for: Obtain the task details of the fourth task corresponding to the proportion of the first task to be analyzed, and / or obtain the task details of the fifth task corresponding to the proportion of the second task to be analyzed; The task details are analyzed to identify and fix any faults in the lineage analysis tool.

37. The apparatus according to any one of claims 21-28, characterized in that, The device further includes: a first processing module, used for: Every preset time interval, the first quantity corresponding to the preset time interval is stored in the first offline table of the big data platform, and the second quantity corresponding to the preset time interval is stored in the second offline table of the big data platform; The first determining module is specifically used for: The first quantity in the first offline table and the second quantity in the second offline table are obtained through the offline scheduling service. Based on the first quantity and the second quantity, the bloodline coverage rate corresponding to the preset time period is determined and stored in the third offline table.

38. The apparatus according to claim 37, characterized in that, The device further includes: a display module, used for: The bloodline coverage rate stored in the third offline table is uploaded to the business intelligence reporting system and / or the data dashboard system to display the bloodline coverage rate through the business intelligence reporting system and / or the data dashboard system.

39. The apparatus according to any one of claims 30-36, characterized in that, The device further includes: a second processing module, used for: Store the first entity corresponding to each task in the fourth offline table, and store the second entity corresponding to each task in the fifth offline table; The fourth determining module is specifically used for: The first entity in the fourth offline table and the second entity in the fifth offline table are obtained through the offline scheduling service. If the first entity and the second entity are the same, then the task corresponding to the first entity or the second entity is determined as the third task, and the number of the third tasks is determined.

40. A computing device, characterized in that, include: At least one processor and memory; The memory stores computer-executed instructions; The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in any one of claims 1 to 19.

Citation Information

Patent Citations

  • Data management method and device, equipment and medium

    CN111008192A

  • Generating sufficiently sized, relatively homogeneous segments of real property transactions by clustering base geographical units

    US20080288312A1