Data blood relationship analysis method and system, terminal and medium

By generating task instances and performing data blood relationship analysis, the problem of difficult data blood relationship in big data processing is solved, and the accurate analysis and display of complex data processing logic and multi-data source fusion scenarios are achieved, which improves the efficiency and scalability of data blood relationship analysis.

CN120045614AInactive Publication Date: 2025-05-27NORTH CHINA DIGITAL HEALTH TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510517929.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the field of big data processing and analysis, data blood relationship tracking is difficult to be carried out accurately and efficiently due to the complexity of data links and the shortcomings of traditional methods.

Method used

By creating data processing tasks, generate task instances, including task parameters and task SQL, and submit them to the Blood Resolution thread pool. According to the field information and field mapping relationship in the task instance, the blood relationship entity table and model table are constructed, the task SQL is parsed to obtain the field mapping relationship, and the analysis results are stored and rendered.

Benefits of technology

It realizes accurate tracking and display of data blood relationships, can handle complex data processing logic and multi-data source fusion scenarios, improves the efficiency and scalability of data blood relationship analysis, promptly detects abnormalities in the data link, and ensures data quality and credibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045614A_ABST
    Figure CN120045614A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing and analysis, and particularly discloses a data blood relationship analysis method and system, a terminal and a medium, and the method comprises the steps: creating a data processing task, and configuring task parameters; executing the created data processing task, generating a task SQL (Structured Query Language) required for executing the task according to the task parameters, and forming a task instance by the task parameters and the task SQL; submitting the task instance which successfully executes the task to a blood relationship analysis thread pool; a to-be-analyzed task instance is extracted from the blood relationship analysis thread pool, and data blood relationship analysis is carried out according to the field information contained in the task instance and the mapping relation between the fields; and rendering and displaying the data blood relationship analysis result on a display interface. According to the invention, the efficiency and accuracy of data blood relationship tracking are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing and analysis, and particularly to a method, system, terminal and medium for data lineage parsing. Background Art

[0002] In the field of big data processing and analysis, data often goes through multiple transformation and synchronization steps from a source database to a target database. These steps involve multiple data sources, complex data processing logics, and different data storage systems, thus intertwining to form an intricate data link. For example, in a risk assessment system in the financial industry, data needs to be extracted from multiple different business databases (such as customer information databases, transaction record databases, etc.) first, and then the extracted data is directly synchronized to another database or stored in another database after passing through certain transformation rules. In such a process, the complexity of the data link poses a huge challenge to the tracking of data lineage.

[0003] Traditional data lineage analysis methods usually rely on adding additional metadata tags at the database level. That is, during the generation, transformation, and flow of data, specific fields are artificially added to the database tables to record the source and flow information of the data. For example, a "source_table" field is added to each data table to record which table the data originally comes from; a "transformation_step" field is added to record which processing steps the data has gone through. However, this method has many defects. On the one hand, the scalability of this method is poor. As the data scale continues to expand and the data processing process becomes increasingly complex, manually adding and maintaining these metadata tags becomes extremely cumbersome. When new data processing steps or data sources are added, a large number of data tables need to be modified, which is prone to omissions or errors. On the other hand, this method lacks an understanding of complex data processing logics. For some data transformation processes involving the fusion of multiple data sources and complex function calculations, simple metadata tags cannot accurately record the true data lineage. For example, after performing a union query on multiple tables and passing through complex aggregation calculations, it is difficult to clearly show the association between the final data and the original data through simple metadata tags. Furthermore, this complexity makes it difficult to track data lineage, and it is difficult to detect anomalies and problems in the data link in a timely manner, thereby affecting the quality and credibility of the data. Summary of the Invention

[0004] To solve the above problems, the present invention provides a method, system, terminal and medium for data lineage parsing, which improves the efficiency and accuracy of data lineage tracking.

[0005] In a first aspect, the technical solution of the present invention provides a data lineage parsing method, including the following steps: Create a data processing task and configure task parameters; Execute the created data processing task, generate the task SQL required for task execution according to the task parameters, and form a task instance with the task parameters and the task SQL; Submit the task instance of the successfully executed task to the lineage parsing thread pool; Extract the task instances to be parsed from the lineage parsing thread pool, and perform data lineage parsing according to the field information contained in the task instances and the mapping relationship between the fields; Render and display the data lineage parsing result on the display interface.

[0006] In an optional implementation, when submitting the task instance of the successfully executed task to the lineage parsing thread pool, initialize the status of the task instance to be parsed; Extract the task instances to be parsed from the lineage parsing thread pool, and perform data lineage parsing according to the field information contained in the task instances and the mapping relationship between the fields, specifically including: Extract all task instances in the to-be-parsed status from the lineage parsing thread pool; Group all the extracted task instances by task. If a task contains multiple task instances, sort the multiple task instances according to the task operation sequence; the data processing task involves at least one target table, and each target table corresponds to a task instance; Perform data lineage analysis on each group of task instances in parallel, and process each task instance in sequence according to the sorting of the task instances in the group containing multiple task instances.

[0007] In an optional implementation, the data processing task includes a data synchronization task; For the data synchronization task, perform data lineage parsing according to the field information contained in the task instance and the mapping relationship between the fields, specifically including: Construct a lineage relationship entity table and a lineage relationship model table; Parse the task SQL to obtain the names of each field it contains; the field names include the source field name and the target field name; Store each field name, together with the first target parameter in the task parameters, as an entity in the lineage relationship entity table, and configure a unique index identifier for the entity; According to the mapping relationship between the source field name and the target field name in the task SQL, store the corresponding mapping relationship between any two entities in the lineage relationship entity table as a lineage relationship in the lineage relationship model table, and configure a unique index identifier for the lineage relationship; the lineage relationship includes the entity unique index identifier and the second target parameter in the task parameters.

[0008] In an optional implementation, the data processing task includes a data conversion task. When creating a data conversion task, a parsing plug-in for calling an SQL script is pre-embedded in each conversion node, and the upper and lower hierarchical relationships, marking sources and targets of each node in the conversion task are customized; For data conversion tasks, data lineage analysis is performed based on the field information contained in the task instance and the mapping relationship between the fields, including: Construct blood relationship entity table and blood relationship model table; Find the first SQL script based on the tag source and the upper and lower hierarchical relationship, and parse all source field names from the first SQL script; Find the last SQL script based on the marked target and the upper and lower hierarchical relationship, and parse all target field names from the last SQL script; Each field name, together with the first target parameter in the task parameters, is stored as an entity in the blood relationship entity table, and a unique index identifier is configured for the entity; the field name includes the source field name and the target field name; Starting from the source, the search is continuously conducted downwards based on the upper and lower layer relationships. During the process, the mapping relationship between each target field name and the corresponding source field name is found based on the mapping relationship between the field names in each SQL script. The corresponding mapping relationship between any two entities in the blood relationship entity table is stored as a blood relationship in the blood relationship model table, and a unique index identifier is configured for the blood relationship. The blood relationship includes the entity unique index identifier and the second target parameter in the task parameters.

[0009] In an optional implementation, the first target parameter of the task parameter includes a data source unique index identifier, a database name, a schema name, and a table name; the second target parameter of the task parameter includes a single-table task unique index identifier and a task instance unique index identifier.

[0010] In an optional implementation, after performing data lineage analysis based on the field information included in the task instance and the mapping relationship between the fields, the following steps are also included: Configure the status of the successfully parsed task instance to be parsed successfully; Configure the status of the task instance that failed to resolve to resolution failed.

[0011] In an optional implementation, rendering and displaying the data lineage analysis result on a display interface specifically includes: The data lineage relationship is integrated into the lineage database according to the relationship between the source and the target, and the metadata of each relationship node is collected; Select the target relationship node to be displayed, render and display the data lineage relationship related to the target relationship node on the display interface according to the data lineage relationship in the lineage database, and render metadata at the relationship node.

[0012] In a second aspect, the technical solution of the present invention provides a data lineage parsing system, including: A data processing task creation module, configured to create a data processing task and configure task parameters; A task instance generation module, configured to execute the created data processing task, generate a task SQL required for executing the task according to the task parameters, and form a task instance with the task parameters and the task SQL; A task instance storage module, configured to submit the task instance of the successfully executed task to a lineage parsing thread pool; A data lineage parsing module, configured to extract a task instance to be parsed from the lineage parsing thread pool, and perform data lineage parsing according to the field information included in the task instance and the mapping relationship between the fields; A data lineage rendering module, configured to render and display the data lineage parsing result on a display interface.

[0013] In a third aspect, the technical solution of the present invention provides a terminal, including: A memory, configured to store a data lineage parsing program; A processor, configured to implement the steps of the data lineage parsing method as described in any one of the above when executing the data lineage parsing program.

[0014] In a fourth aspect, the technical solution of the present invention provides a computer-readable storage medium, on which a data lineage parsing program is stored, and when the data lineage parsing program is executed by a processor, the steps of the data lineage parsing method as described in any one of the above are implemented.

[0015] As can be seen from the above technical solutions, the present application has the following advantages: When performing a data processing task, a task instance is generated, which contains task parameters and a task SQL. Then, data lineage parsing is performed based on the information contained in the task parameters and the task SQL. By generating a task SQL according to the task parameters and forming a task instance with the task parameters and the task SQL, and performing data lineage parsing from the task instance information, the flow path of data under complex processing logic can be accurately grasped, and thus the true lineage relationship of data from the source database through multiple transformation and synchronization steps to the target database can be accurately recorded. Even in the face of complex scenarios such as multi-data source integration and complex function calculations, the association between the final data and the original data can be clearly displayed, effectively solving the problem that it is difficult to trace data lineage. Performing lineage parsing based on task instances, when new data processing steps or data sources are added, only operations need to be performed in the link of creating data processing tasks and configuring task parameters, without having to modify a large number of data tables one by one, greatly improving scalability and adapting to the dynamic changes of the big data processing environment. Since the data lineage relationship can be accurately traced, when an anomaly or problem occurs in the data link, the link where the problem occurs and the data sources involved can be quickly located based on the parsing results, and the problem can be discovered and processed in a timely manner, effectively shortening the problem troubleshooting time and ensuring data quality and credibility. Rendering and displaying the data lineage parsing results on the display interface facilitates users to intuitively understand the data processing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions of the present application, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0017] Figure 1 It is a schematic flowchart of a data lineage parsing method provided by an embodiment of the present invention.

[0018] Figure 2 It is a schematic diagram of the interface for creating a data conversion task.

[0019] Figure 3 It is a schematic diagram of the conversion node configuration interface.

[0020] Figure 4 It is a schematic diagram of an entity in the lineage relationship entity table.

[0021] Figure 5 It is a schematic diagram of a lineage relationship in the lineage relationship model table.

[0022] Figure 6 It is a schematic diagram of table lineage rendering.

[0023] Figure 7 Schematic diagram for field blood relationship rendering

[0024] Figure 8 Schematic block diagram of the structure of a data blood relationship parsing system provided by an embodiment of the present invention

[0025] Figure 9 Schematic diagram of the structure of a terminal provided by an embodiment of the present invention Detailed implementation manners

[0026] To make the application purpose, features, and advantages of the present application more obvious and understandable, the technical solutions protected by the present application will be clearly and completely described below by using specific embodiments and the accompanying drawings. Obviously, the embodiments described below are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention

[0028] The following explains the key terms that appear in the present invention

[0029] SQL: Structured Query Language, a database language with various functions such as data manipulation and data definition

[0030] Figure 1 Schematic flowchart of a data blood relationship parsing method provided by an embodiment of the present invention. Among them Figure 1 The execution subject can be a data blood relationship parsing system. The data blood relationship parsing method provided by the embodiment of the present invention is executed by a computer device. Correspondingly, the data blood relationship parsing system runs in the computer device. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted

[0031] Such as Figure 1 shown, the method includes the following steps

[0032] S1, create a data processing task and configure task parameters

[0033] In the data synchronization management platform, users can create data synchronization tasks (single table synchronization or whole database synchronization) or data conversion tasks, and configure parameters such as data source connection, data filtering, incremental conditions, field mapping, etc. This step clarifies the detailed requirements of the data processing task and provides basic information for subsequent task execution and data lineage analysis, so as to determine the source, destination, and processing rules of the data.

[0034] S2, execute the created data processing task, generate the task SQL required to execute the task according to the task parameters, and form a task instance with the task parameters and the task SQL.

[0035] After the task configuration is completed, the system generates task SQL based on the parameters and combines the parameters and SQL into a task instance. It should be noted that the data processing task involves at least one target table, and each target table corresponds to a task instance, that is, an instance is generated separately for each target table. The task instance of this step contains complete task execution information, which is independent of the task template, making it easy to trace and repeat execution. The generated task SQL is used for actual data processing to ensure the recordability and repeatability of data processing operations, and provide a basis for blood relationship analysis.

[0036] S3 submits the task instance of the successfully executed task to the lineage resolution thread pool.

[0037] After the task is successfully executed, the task instance is submitted to the lineage resolution thread pool and its status is set to pending resolution. This step centrally manages the task instances to be resolved and provides a unified entry for lineage resolution. The use of the thread pool facilitates the subsequent concurrent processing of task instances, improves the efficiency of lineage resolution, ensures that task instances are not missed, and performs lineage resolution in an orderly manner.

[0038] S4, extract the task instance to be analyzed from the lineage analysis thread pool, and perform data lineage analysis based on the field information contained in the task instance and the mapping relationship between the fields.

[0039] Extract the task instances to be parsed from the thread pool, sort them by task group, perform lineage parsing in parallel, build corresponding tables for different task types (synchronization or conversion), and store the mapping relationships obtained from the parsed fields. This step sorts out the data flow path from source to target, records the data lineage relationship, and can clearly display data associations for complex data processing scenarios. Lineage relationship storage can provide support for data quality management, problem troubleshooting, and process optimization.

[0040] S5, rendering and displaying the data lineage analysis results on the display interface.

[0041] Integrate data lineage into the lineage database and collect metadata, and render relevant lineage relationships and metadata on the display interface according to user selection. This step presents data lineage in an intuitive and visual way, facilitating users to understand the data processing flow and quickly locate the source and destination of data.

[0042] Furthermore, as a refinement and extension of the specific implementation manner of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another data lineage parsing method is provided, and this method includes the following steps.

[0043] SS1. Create a data processing task and configure task parameters.

[0044] In some optional implementation manners, create and execute data processing tasks on the data synchronization management platform. First, log in to the data synchronization management platform and create a task. The task database types include, but are not limited to, common database types such as oracle, mysql, KingbaseES, and DM.

[0045] Data processing tasks all converge data from one database to another, including data synchronization tasks and data conversion tasks. Data synchronization tasks include single-table synchronization and full-database synchronization tasks.

[0046] The task parameters configured for creating a single-table synchronization task include: (1) Data source connection parameters Related to the source data source: including the source data source (oracle_uz7), the source database (NCTU97), the source schema (ORA_US1), and the source table (TEST_J1). These parameters determine from which database and schema to obtain the source table data; Related to the target data source: the target data source (OpenGauss1), the target database (omm), the target schema (zyg_test), and the target table (TEST_J1), which clarify the target location of data synchronization; (2) Data filtering parameters: There are corresponding SQL conditions (DATE_SUB(System_datestr, INTERVAL 10 DAY)) in the data filtering column, which are used to filter out data that meets specific conditions for synchronization and can control the range of synchronized data; (3) Incremental condition parameters: Incremental synchronization rules can be set to determine which data needs to be synchronized after being newly added or updated. For example, it can be judged according to conditions such as timestamps; (4)Field mapping parameters: In the field mapping configuration area, there are source table fields (such as ID, COLUMN1) and their field types, length precision, etc. information, and a corresponding filling area for the target table fields is also reserved. These parameters are used to determine the mapping relationship between the source table fields and the target table fields to ensure that data is correctly synchronized to the corresponding fields of the target table; (5)Other parameters: The target table creation method (automatic table creation) determines whether the target table is automatically created by the system according to the source table structure or needs to be created manually in advance; DDL alarm is used to configure the alarm rules for DDL operations on the target table.

[0047] The task parameters of the full database synchronization task are similar to those of the single table synchronization task, including data source connection parameters, table-related parameters, data filtering parameters, incremental condition parameters, field mapping parameters, and other parameters. Among them, the table-related parameters include: Source table: The source table name specifies the source table from which data is to be synchronized.

[0048] Target table: The target table name, and the target table creation method can be automatic table creation, which means that the system will automatically create the table in the target database according to the source table structure.

[0049] Figure 2 To create a schematic diagram of the data conversion task interface, since the conversion synchronization task involves multiple conversion links, parsing plugins need to be pre-embedded for each conversion node link, and the upper and lower hierarchical order within the instance, as well as the source and target, are customized. As Figure 2 shown, the upper and lower node relationships have been defined in the process of creating each task. The middle processes are all conversion nodes. Only the starting point of the task is the source table information, which can be multiple, and the ending node is the target table information, which is only one. The pre-embedded parsing plugin is used to retrieve the SQL script. Specifically, by introducing the jar package method. For example, the SQL script is both an introduced jar package and a front-end program package. Therefore, the conversion node is a flexible plugin and can be extended. The plugins will all implement the code for parsing the upper and lower level relationships. No matter how many conversion nodes are accessed, the information association from the source to the target is ultimately achieved.

[0050] Figure 3 For the schematic diagram of the conversion node configuration interface, the corresponding configured task parameters include data source connection parameters and conversion node parameters. Among them, the conversion node parameters include node name and whether to customize SQL.

[0051] SS2 executes the created data processing task, generates the task SQL required for executing the task according to the task parameters, and forms a task instance with the task parameters and the task SQL.

[0052] After the task parameters are configured, click the execution button to execute the created data processing task. During the execution process, the background will generate a task SQL based on the task parameters. At the same time, the task parameters and the task SQL will form a task instance, and a unique index identifier will be configured for the task instance. In some alternative implementation manners, the task instance also includes the task execution time and the unique index identifier of the task instance. It should be noted that the task instance is actually a copy of the task at the current moment. Subsequently, if the task template is modified (modifying specific parameter values or parameter types), the information of the task instance will not be modified, and the task instance can be executed again according to the parameters retained at that time.

[0053] It should be noted that each target table corresponds to a task instance. For a single-table synchronization task, one task instance is generated. For a full-database synchronization task, generally multiple target tables are involved. At this time, multiple task instances will be generated for the full-database synchronization task. For a transformation task, also one target table corresponds to one task instance.

[0054] For a single-table synchronization task, the generation of the task SQL can be achieved through the following steps.

[0055] Step 1, determine the source table and the target table.

[0056] Based on the source table (TEST_J1) and the target table (TEST_J1), as well as the corresponding database and schema information, determine the table names operated in the SQL. For example, query the data of the source table in the Oracle data source and insert the data into the target table in the OpenGauss target database.

[0057] Step 2, construct the SELECT clause.

[0058] According to the source table field information in the field mapping configuration, determine the field list after SELECT. If there is no special requirement for field conversion, the source table fields can be directly selected, such as SELECT ID, COLUMN1 FROM ORA_US1.TEST_J1. If there are requirements such as data type conversion, corresponding processing needs to be performed in the SELECT clause, such as converting a numeric field to a character type for display, etc.

[0059] Step 3, add data filtering conditions.

[0060] According to the SQL condition in the data filtering column (DATE_SUB(System_datestr, INTERVAL 10 DAY)), add it to the WHERE clause of the SELECT statement to further filter the data in the source table, resulting in a statement like SELECT ID, COLUMN1 FROM ORA_US1.TEST_J1 WHERE DATE_SUB(System_datestr, INTERVAL 10 DAY).

[0061] Step 4, construct the INSERT INTO clause.

[0062] According to the target table information (zyg_test.TEST_J1), construct the INSERT INTO statement part to insert the data filtered by SELECT into the target table. Combining the previous SELECT statement, a complete synchronization SQL statement may finally be formed, such as INSERT INTO zyg_test.TEST_J1 (ID, COLUMN1) SELECT ID, COLUMN1 FROM ORA_US1.TEST_J1 WHERE DATE_SUB(System_datestr, INTERVAL 10 DAY).

[0063] Step 5, consider the incremental condition.

[0064] If the incremental condition is configured, the incremental condition logic also needs to be incorporated into the SQL statement. For example, for incremental synchronization based on timestamps, a judgment condition for the timestamp field may need to be added to the WHERE clause to ensure that only newly added or updated data is synchronized.

[0065] For the full database synchronization task, the task SQL for generating a single table is similar to that for generating the task SQL in the above single table synchronization task, and will not be elaborated here.

[0066] For the transformation task, the task SQL can be generated through the following steps.

[0067] Step 1, read and obtain node information.

[0068] Obtain the source data information from the read node. Determine the data source type (OpenGauss), data source (OpenGauss1), database (omm), schema (wy_test), and data table (wy_city). Based on this information, the basic SQL statement structure for reading data from the source table can be constructed, such as SELECT * FROM wy_test.wy_city (assuming full table reading, and fields may actually be selected as needed).

[0069] Step 2, process the SQL script node.

[0070] Pass the data obtained by the reading node into the SQL script node. In this node, write an SQL script according to specific business requirements to perform data conversion operations. For example, it may perform field calculations (such as SELECT column1 + column2 AS new_column, other_column FROM...), data filtering (SELECT * FROM... WHERE some_condition), etc. Process the data read from the source table through the SQL script to form the converted data.

[0071] Step 3, the writing node determines the target.

[0072] Clarify the target table information of the writing node. According to the target table structure and the previously converted data, construct an SQL statement to write the data into the target table, such as INSERT INTO target_table (column1, column2,...) SELECT processed_column1, processed_column2,... FROM..., where target_table is the name of the target table and processed_column is the field after being processed by the SQL script. Integrate this series of operations to form a complete synchronization SQL statement from the source table through conversion to the target table.

[0073] SS3, submit the task instances of the successfully executed tasks to the lineage parsing thread pool.

[0074] After the data processing task is successfully executed, submit all task instances of the task to the lineage parsing thread pool, and at the same time initialize the status of the task instances to be parsed. In some embodiments, configure the status of the successfully parsed task instances as parsed successfully; configure the status of the failed parsed task instances as parsed failed.

[0075] SS4, extract the task instances to be parsed from the lineage parsing thread pool, and perform data lineage parsing according to the field information contained in the task instances and the mapping relationship between the fields.

[0076] The lineage parsing thread pool stores all task instances, and there may be multiple task instances to be parsed. In some optional real-time modes, the lineage parsing thread pool can group the task instances and process the task instances orderly and concurrently, ensuring that while accurately parsing the lineage of the task instances, it can also be efficient without backlog. Specifically, it includes the following steps.

[0077] SS4.1, extract all task instances in the to-be-parsed status from the lineage parsing thread pool.

[0078] SS4.2, group all the extracted task instances by task. If a task contains multiple task instances, sort the multiple task instances according to the task operation sequence.

[0079] SS4.3, process each group of task instances in parallel for data lineage analysis. In a group containing multiple task instances, process each task instance in sequence according to the sorting of the task instances.

[0080] Specifically, the grouping is based on the tasks corresponding to all task instances, and it will be in the form of: (key, List <values>)Stored in Redis, the key is the task ID, and the values are a list of task instances to be parsed corresponding to the task. The list is added in chronological order. Concurrent processing is handled using a thread pool. Concurrency is targeted at the key, while the instance list List corresponding to the task <values>To ensure that the order is processed in a single-threaded manner, early task instances are processed first, followed by later task instances. If one instance gets stuck in the middle, the subsequent ones will no longer be processed. After parsing an instance, the corresponding values will be deleted.

[0081] For example, the full database synchronization task contains multiple task instances for single tables. First, a group will be defined, and within the group, it will be split into multiple single-table task instances to parse the lineage. The array here is the same as the task key above. The full database task includes tasks for multiple tables, as long as List <values>It only needs to be ordered. To ensure order, a List is generated by comparing the times at the moment when the task is successfully executed and the blood relationship to be parsed status is generated <values>. Specifically, after clicking the execution button for the full-library task, multi-threaded concurrent technology is used to generate multiple instances and submit them to DataX. The data synchronization instances are also executed concurrently, and this is determined by the server resources and configurations of the middleware DataX.

[0082] In some alternative embodiments, for the data synchronization task, data lineage parsing is performed based on the field information included in the task instance and the mapping relationships between the fields, which specifically includes the following steps.

[0083] SS410, construct the lineage entity table and the lineage model table.

[0084] Two tables are involved, namely the lineage entity table and the lineage model table. The lineage entity table is used to store lineage entities, and the lineage model table is used to store the relationships between entities.

[0085] SS411, parse the task SQL to obtain each field name included therein; the field names include the source field name and the target field name.

[0086] Use an SQL parsing tool to parse the task SQL in the task instance to obtain the field names. Exemplarily, the select statement of the task SQL is as follows: (select MEDICAL_ORG_ID as medical_org_id, ID as id, MEMBER_ID as member_id, LINKMAN_TEL as linkman_tel, LINKMAN_RELATION_CODE as linkman_relation_code, LINKMAN_RELATION_NAME as linkman_relation_name, LINKMAN_NAME as linkman_name, CREATE_TIME as create_time, UPDATE_TIME as update_time, VALID_FLAG as valid_flag from AID_MEMBER_CONTACT ) Through this statement, the source table, source table fields, and target fields can be obtained. On the left side of "as" in the above statement is the source table field name, and on the right side of "as" is the target field name. AID_MEMBER_CONTACT is the table name.

[0087] SS412. Store each field name, together with the first target parameter in the task parameters, as an entity in the lineage entity table, and configure a unique index identifier for the entity.

[0088] In some alternative embodiments, the first target parameter includes a data source unique index identifier, a database name, a schema name, and a table name.

[0089] Store each field name, together with the first target parameter in the task parameters, as an entity in the lineage entity table, and configure a unique index identifier for the entity. Figure 4 A schematic diagram of an entity in the lineage entity table, including an entity unique index identifier, a data source unique index identifier, a database name, a schema name, a table name, a field name, and an operation time. The operation time refers to the task execution time and is stored in the task instance. At the same time, metadata of each data item is also extracted, including data type, length, etc.

[0090] SS413. According to the mapping relationship between the source field name and the target field name in the task SQL, store the corresponding mapping relationship between any two entities in the lineage entity table as a lineage in the lineage model table, and configure a unique index identifier for the lineage; the lineage includes an entity unique index identifier and the second target parameter in the task parameters.

[0091] All entities are stored in the lineage entity table. Then, according to the mapping relationship between the source field name and the target field name in the task SQL, store the corresponding mapping relationship between any two entities in the lineage entity table as a lineage in the lineage model table. For example, if there is a mapping relationship between the first source field and the first target field, then there is a mapping relationship between the entity where the first source field is located and the entity where the first target field is located.

[0092] In some alternative embodiments, the second target parameter includes a single-table task unique index identifier and a task instance unique index identifier.

[0093] Figure 5 It is a schematic diagram of a blood relationship in the blood relationship model table, including a unique index identifier for the blood relationship, a unique index identifier for the entity (including the source field entity ID and the target field entity ID), a unique index identifier for the single-table task, a unique index identifier for the task instance (corresponding to the latest subtask ID in the figure), and the operation time. The operation time refers to the task execution time and is stored in the task instance. At the same time, the metadata of each data item is also extracted, including data type, length, etc.

[0094] In some alternative embodiments, for the data conversion task, data lineage parsing is performed according to the field information and the mapping relationship between fields included in the task instance, which specifically includes the following steps.

[0095] S420, construct a blood relationship entity table and a blood relationship model table.

[0096] S421, find the first SQL script according to the marked source and the hierarchical relationship, and parse out all source field names from the first SQL script.

[0097] S422, find the last SQL script according to the marked target and the hierarchical relationship, and parse out all target field names from the last SQL script.

[0098] S423, store each field name, together with the first target parameter in the task parameters, as an entity in the blood relationship entity table, and configure a unique index identifier for the entity; the field names include source field names and target field names.

[0099] S424, continuously search downward from the source according to the hierarchical relationship. During the process, find the mapping relationship between each target field name and the corresponding source field name according to the field name mapping relationship in each SQL script, and store the corresponding mapping relationship between any two entities in the blood relationship entity table as a blood relationship in the blood relationship model table, and configure a unique index identifier for the blood relationship; the blood relationship includes the unique index identifier of the entity and the second target parameter in the task parameters.

[0100] In some alternative embodiments, the synchronization task and the conversion task share a blood relationship entity table and a blood relationship model table. The process is similar to data synchronization. Different from the synchronization task, the conversion instance blood lineage parsing first finds the source table and collects the source table metadata information, and parses out the source table fields; there can be multiple source tables. Continuously search downward from the source according to the pre-buried points. During the process, continuously generate the hierarchical relationship, and finally find the final target according to the target table mark, and determine the blood relationship between the final source and the final target. That is, compared with the synchronization task, there is more process information of the intermediate nodes, which is equivalent to generating the synchronization SQL many times, and finally parsing the synchronization SQL to the target.

[0101] Exemplarily, the select statement of the task SQL is as follows: (select AA.ID as id, AA.MEMBER_ID as member_id, AA.LINKMAN_TEL as linkman_tel, BB.CREATE_TIME as create_time, BB.UPDATE_TIME as update_time, BB.VALID_FLAG as valid_flag from AID_MEMBER_CONTACT AA INNER JOIN USER BB ON AA.ID = BB.ID ) There are two source tables involved in this statement, AID_MEMBER_CONTACT (source fields: ID, MEMBER_ID, LINKMAN_TEL) and USER (CREATE_TIME, UPDATE_TIME, VALID_FLAG), and 6 target fields in lowercase, namely id...valid_flag. So there are a total of 12 entity data in the lineage entity table and 6 relationships in the lineage model table.

[0102] SS5, render and display the data lineage parsing result on the display interface.

[0103] SS5.1, integrate the data lineage relationship into the lineage database according to the source and target relationships, and collect the metadata of each relationship node.

[0104] SS5.2, select the target relationship nodes to be displayed, render and display the data lineage relationship related to the target relationship nodes on the display interface according to the data lineage relationship in the lineage database, and render the metadata at the relationship nodes.

[0105] According to the source_entity_id in the lineage model table, the target_entity_id can be found. Then the target_entity_id may also be the source_entity_id of another record. Keep looking down and up to form a chain relationship or even a cyclic relationship. Eventually, it may also be a star relationship. Figure 6 For the schematic diagram of table lineage rendering, Figure 7 It is a schematic diagram for field lineage rendering, assembling a specific data format. Clicking the plus sign can trace up and down. Specifically, select any node in the vector diagram (a table can be regarded as an arbitrary node, or a table field can be regarded as an arbitrary node), and realize tracing the paths of other nodes affected by this node downward (upward) starting from (ending at) this node.

[0106] In some alternative embodiments, the node metadata includes database, data type, data source, storage capacity, and lineage update time.

[0107] In the above, embodiments of a data lineage parsing method have been described in detail. Based on the data lineage parsing method described in the above embodiments, an embodiment of the present invention also provides a data lineage parsing device corresponding to this method.

[0108] Figure 8 It is a schematic block diagram of the structure of a data lineage parsing system provided by an embodiment of the present invention. In this embodiment, the data lineage parsing system 800 can be divided into multiple functional modules according to the functions it performs. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory.

[0109] The data processing task creation module 810 is used to create a data processing task and configure task parameters.

[0110] The task instance generation module 820 is used to execute the created data processing task, generate the task SQL required for executing the task according to the task parameters, and form a task instance with the task parameters and the task SQL.

[0111] The task instance storage module 830 is used to submit the task instances of successfully executed tasks to the lineage parsing thread pool.

[0112] The data lineage parsing module 840 is used to extract the task instances to be parsed from the lineage parsing thread pool and perform data lineage parsing according to the field information and the mapping relationship between fields included in the task instances.

[0113] The data lineage rendering module 850 is used to render and display the data lineage parsing results on the display interface.

[0114] The data lineage parsing system in this embodiment is used to implement the foregoing data lineage parsing method. Therefore, the specific implementation manners in this system can be seen in the embodiment part of the data lineage parsing method in the foregoing text. Therefore, its specific implementation manners can refer to the descriptions of the corresponding various part embodiments and will not be elaborated here.

[0115] In addition, since the data lineage parsing system of this embodiment is used to implement the aforementioned data lineage parsing method, its functions correspond to those of the above method, and will not be elaborated here.

[0116] Figure 9 FIG. 4 is a schematic structural diagram of a terminal 900 provided by an embodiment of the present invention, including: a processor 910, a memory 920, and a communication unit 930. The processor 910 is used to implement the method steps in the above-mentioned data lineage parsing method embodiment when implementing the data lineage parsing program stored in the memory 920.

[0117] The terminal 900 includes a processor 910, a memory 920, and a communication unit 930. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and can also include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0118] Among them, the memory 920 can be used to store the execution instructions of the processor 910. The memory 920 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. When the execution instructions in the memory 920 are executed by the processor 910, the terminal 900 can execute some or all of the steps in the above method embodiments.

[0119] The processor 910 is the control center of the storage terminal, connecting various parts of the entire electronic terminal through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 920, and calling the data stored in the memory, it executes various functions of the electronic terminal and / or processes data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or can be composed of multiple packaged ICs with the same or different functions connected. For example, the processor 910 can only include a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single operation core or can include multiple operation cores.

[0120] The communication unit 930 is used to establish a communication channel, so that the storage terminal can communicate with other terminals. Receive user data sent by other terminals or send user data to other terminals.

[0121] The present invention also provides a computer storage medium, which may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0122] The computer storage medium stores a data lineage parsing program, and when the data lineage parsing program is executed by a processor, it implements the method steps in the above-mentioned method embodiments of the data lineage parsing method.

[0123] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes, and includes several instructions to enable a computer terminal (which can be a personal computer, a server, or a second terminal, a network terminal, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0124] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0125] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0126] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0127] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.< / values> < / values> < / values> < / values>

Claims

1. A data lineage analysis method, characterized in that: The following steps are involved: Create data processing tasks and configure task parameters; Execute the created data processing task, generate the task SQL required to execute the task according to the task parameters, and form a task instance with the task parameters and task SQL; Submit the task instance that successfully executes the task to the lineage resolution thread pool; Extract the task instance to be parsed from the lineage parsing thread pool, and perform data lineage parsing based on the field information contained in the task instance and the mapping relationship between the fields; The data lineage analysis results are rendered and displayed on the display interface.

2. The data lineage analysis method according to claim 1, characterized in that: When submitting a task instance of a successfully executed task to the lineage resolution thread pool, the state of the task instance is initialized to pending resolution; Extract the task instance to be parsed from the lineage parsing thread pool, and perform data lineage parsing based on the field information contained in the task instance and the mapping relationship between the fields, specifically including: Extract all task instances to be parsed from the lineage parsing thread pool; All the extracted task instances are grouped according to the tasks. If the task contains multiple task instances, the multiple task instances are sorted according to the task operation order. The data processing task involves at least one target table, and each target table corresponds to a task instance. Each group of task instances is processed in parallel to perform data lineage analysis. In a group containing multiple task instances, each task instance is processed in sequence according to the order of the task instances.

3. The data lineage analysis method according to claim 1, characterized in that: Data processing tasks include data synchronization tasks; For data synchronization tasks, data lineage analysis is performed based on the field information contained in the task instance and the mapping relationship between the fields, including: Construct blood relationship entity table and blood relationship model table; Parse the task SQL to obtain the names of the fields it contains; the field names include the source field name and the target field name; Store each field name, together with the first target parameter in the task parameters, as an entity in the blood relationship entity table, and configure a unique index identifier for the entity; According to the mapping relationship between the source field name and the target field name in the task SQL, the corresponding mapping relationship between any two entities in the blood relationship entity table is stored as a blood relationship in the blood relationship model table, and a unique index identifier is configured for the blood relationship; the blood relationship includes the entity unique index identifier and the second target parameter in the task parameters.

4. The data lineage analysis method according to claim 1, characterized in that: Data processing tasks include data conversion tasks. When creating a data conversion task, a parsing plug-in for calling SQL scripts is embedded in each conversion node, and the upper and lower hierarchical relationships, source and target tags of each node in the conversion task are customized. For data conversion tasks, data lineage analysis is performed based on the field information contained in the task instance and the mapping relationship between the fields, including: Construct blood relationship entity table and blood relationship model table; Find the first SQL script based on the tag source and the upper and lower hierarchical relationship, and parse all source field names from the first SQL script; Find the last SQL script based on the marked target and the upper and lower hierarchical relationship, and parse all target field names from the last SQL script; Each field name, together with the first target parameter in the task parameters, is stored as an entity in the blood relationship entity table, and a unique index identifier is configured for the entity; the field name includes the source field name and the target field name; Starting from the source, the search is continuously conducted downwards based on the upper and lower layer relationships. During the process, the mapping relationship between each target field name and the corresponding source field name is found based on the mapping relationship between the field names in each SQL script. The corresponding mapping relationship between any two entities in the blood relationship entity table is stored as a blood relationship in the blood relationship model table, and a unique index identifier is configured for the blood relationship. The blood relationship includes the entity unique index identifier and the second target parameter in the task parameters.

5. The data lineage analysis method according to claim 3 or 4, characterized in that: The first target parameter of the task parameter includes the data source unique index identifier, database name, mode name, and table name; the second target parameter of the task parameter includes the single-table task unique index identifier and the task instance unique index identifier.

6. The data lineage analysis method according to any one of claims 1 to 4, characterized in that: After performing data lineage analysis based on the field information contained in the task instance and the mapping relationship between the fields, the following steps are also included: Configure the status of the successfully parsed task instance to be parsed successfully; Configure the status of the task instance that failed to resolve to resolution failed.

7. The data lineage analysis method according to any one of claims 1 to 4, characterized in that: The data lineage analysis results are rendered and displayed on the display interface, including: The data lineage relationship is integrated into the lineage database according to the relationship between the source and the target, and the metadata of each relationship node is collected; Select the target relationship node to be displayed, render and display the data lineage relationship related to the target relationship node on the display interface according to the data lineage relationship in the lineage database, and render metadata at the relationship node.

8. A data lineage analysis system, characterized in that: include: Data processing task creation module, used to create data processing tasks and configure task parameters; The task instance generation module is used to execute the created data processing task, generate the task SQL required to execute the task according to the task parameters, and form a task instance with the task parameters and the task SQL; The task instance storage module is used to submit the task instances of successfully executed tasks to the lineage resolution thread pool; The data lineage analysis module is used to extract the task instance to be analyzed from the lineage analysis thread pool, and perform data lineage analysis based on the field information contained in the task instance and the mapping relationship between the fields; The data lineage rendering module is used to render and display the data lineage analysis results on the display interface.

9. A terminal, characterized in that: include: A memory for storing a data lineage analysis program; A processor, used to implement the steps of the data lineage analysis method as described in any one of claims 1 to 7 when executing the data lineage analysis program.

10. A computer-readable storage medium, characterized in that: The readable storage medium stores a data lineage analysis program, which, when executed by a processor, implements the steps of the data lineage analysis method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data management full link-based field-level blood relationship analysis method

    CN114116856A

  • Data blood relationship analysis method and system

    CN117216034A

  • Data blood relationship determination method, device and equipment and readable storage medium

    CN117238398A

  • Data integration method and device, server, medium and program

    CN118626496A

  • Data blood relationship full-link monitoring method and system, terminal and storage medium

    CN118939839A