Field Blood Relationship Analysis Method, Device, Electronic Device and Storage Medium

By constructing abstract syntax trees and directed acyclic graphs, the problem of low accuracy of field blood relationship analysis in the existing technology is solved, and more flexible and highly accurate blood relationship analysis is achieved.

CN113961584BActive Publication Date: 2025-06-24PING AN BANK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111219787.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-20
Publication Date
2025-06-24
Estimated Expiration
2041-10-20

AI Technical Summary

Technical Problem

The existing field blood relationship analysis methods have low accuracy, rely too much on pre-acquisitioned analysis templates, and are not flexible enough.

Method used

By obtaining the source data table and job set, extracting SQL statements, generating target data tables, building an abstract syntax tree, performing blood relationship analysis, building a directed acyclic graph, and obtaining the blood relationship of the fields to be queryed.

Benefits of technology

Improve the accuracy of field blood relationship analysis, enhance the flexibility of analysis, and can more effectively process and analyze blood relationships between data tables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113961584B_ABST
    Figure CN113961584B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence and digital medical technologies, and discloses a method for field lineage analysis, including: obtaining a source data table and a job set, extracting the SQL statement of any job in the job set, generating a target data table according to the SQL statement and the job corresponding to the SQL statement, constructing an abstract syntax tree based on the SQL statement, using the abstract syntax tree to perform lineage parsing on the source data table and the target data table to obtain a lineage parsing result, constructing a directed acyclic graph according to the lineage parsing result and a preset single correspondence relationship, obtaining a field to be queried, and obtaining the lineage relationship corresponding to the field to be queried based on the directed acyclic graph. In addition, the present invention also relates to blockchain technology, and the SQL statement can be stored in the nodes of the blockchain. The present invention also provides a field lineage analysis device, an electronic device, and a storage medium. The present invention can solve the problem of low accuracy in performing field lineage analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device and computer-readable storage medium for field blood relationship analysis. Background Art

[0002] With the advent of the information age, enterprises, employees, devices, etc. will continuously generate and consume data. Facing the vast amount of data, it is particularly important for enterprises to manage and use the data to give play to its value. Among them, as one of the applications of metadata, field blood relationship analysis can allow users to view the processing path of table fields on the metadata platform, and mark whether downstream fields need to be desensitized in the desensitization system according to the field blood relationship information. It is used to represent a logical relationship formed during the generation, fusion, transfer, and extinction of data, and more attention needs to be paid to and utilized. Therefore, there is an urgent need to propose a method for field blood relationship analysis.

[0003] Existing field blood relationship analysis methods usually rely on pre-obtained analysis templates as references for blood relationship analysis. This method is too dependent on existing analysis templates, not flexible enough, and the accuracy of blood relationship analysis is not high enough. Summary of the Invention

[0004] The present invention provides a method, device, electronic device and computer-readable storage medium for field blood relationship analysis, and its main purpose is to solve the problem of low accuracy in field blood relationship analysis.

[0005] To achieve the above object, a method for field blood relationship analysis provided by the present invention includes:

[0006] Obtain the source data table and the job set, and extract the SQL statement of any job in the job set;

[0007] Generate a target data table according to the SQL statement and the job corresponding to the SQL statement;

[0008] Construct an abstract syntax tree based on the SQL statement;

[0009] Use the abstract syntax tree to perform blood relationship parsing on the source data table and the target data table to obtain a blood relationship parsing result;

[0010] Construct a directed acyclic graph according to the blood relationship parsing result and a preset single correspondence relationship;

[0011] Obtain the field to be queried, and obtain the blood relationship corresponding to the field to be queried based on the directed acyclic graph.

[0012] Optionally, the generating a target data table according to the SQL statement and the job corresponding to the SQL statement includes:

[0013] Extract the job name in the job corresponding to the SQL statement, and extract the query fields in the SQL statement;

[0014] Use the job name as the table name of the initial data table, and use the query fields as the table fields of the initial data table to generate the target data table.

[0015] Optionally, constructing the abstract syntax tree based on the SQL statement includes:

[0016] Obtain a preset keyword set, filter out the words in the SQL statement that are the same as the keywords in the keyword set to obtain standard keywords;

[0017] Use the standard keywords to split the SQL statement to obtain multiple SQL sub-statements;

[0018] Convert multiple SQL sub-statements into an abstract syntax tree.

[0019] Optionally, using the standard keywords to split the SQL statement to obtain multiple SQL sub-statements includes:

[0020] Use the standard keywords as split nodes, and split the SQL statement to the left and right based on the split nodes to obtain multiple SQL sub-statements; or

[0021] Based on a preset random split length, split the SQL statement to obtain multiple split statements, and use the standard keywords as positioning points to perform secondary splitting on the multiple split statements to obtain multiple SQL sub-statements.

[0022] Optionally, converting multiple SQL sub-statements into an abstract syntax tree includes:

[0023] Use a preset lexical analyzer to analyze the SQL sub-statements to obtain multiple tokens;

[0024] Construct a syntax tree for multiple tokens according to a preset syntax analysis method to obtain the abstract syntax tree.

[0025] Optionally, using the abstract syntax tree to perform lineage parsing on the source data table and the target data table to obtain a lineage parsing result includes:

[0026] Traverse the abstract syntax tree, and mark the nodes in the abstract syntax tree that are the same as the nodes of the preset fields as target nodes;

[0027] Extract the data value in the target node, compare and query based on the data value and the data table in the database, and obtain the job table corresponding to the data value;

[0028] Determine the type of the job table according to the label corresponding to the job table, determine the relationship between different types of job tables, and summarize the job table and the relationship between job tables to obtain the blood relationship parsing result.

[0029] Optionally, constructing a directed acyclic graph according to the blood relationship parsing result and a preset single correspondence relationship includes:

[0030] Use the starting data table in the blood relationship parsing result as the first node and the target data table in the blood relationship parsing result as the second node;

[0031] Construct an initial node graph according to the relationship that the first node points to the second node;

[0032] According to the single correspondence relationship, replace the job data table corresponding to each node in the initial node graph with the job corresponding to the job set to obtain the directed acyclic graph.

[0033] To solve the above problems, the present invention also provides a field blood relationship analysis device, and the device includes:

[0034] A statement extraction module, configured to obtain a source data table and a job set, and extract the SQL statement of any job in the job set;

[0035] A data table generation module, configured to generate a target data table according to the SQL statement and the job corresponding to the SQL statement;

[0036] A syntax tree construction module, configured to construct an abstract syntax tree based on the SQL statement;

[0037] A blood relationship analysis module, configured to perform blood relationship parsing on the source data table and the target data table by using the abstract syntax tree to obtain a blood relationship parsing result, construct a directed acyclic graph according to the blood relationship parsing result and a preset single correspondence relationship, obtain a field to be queried, and obtain a blood relationship corresponding to the field to be queried based on the directed acyclic graph.

[0038] To solve the above problems, the present invention also provides an electronic device, and the electronic device includes:

[0039] At least one processor; and,

[0040] A memory communicatively connected to the at least one processor; wherein,

[0041] The memory stores a computer program executable by the at least one processor. When executed by the at least one processor, the computer program enables the at least one processor to execute the field blood relationship analysis method described above.

[0042] To solve the above problems, the present invention further provides a computer-readable storage medium storing at least one computer program. When the at least one computer program is executed by a processor in an electronic device, the field blood relationship analysis method described above is implemented.

[0043] In an embodiment of the present invention, a target data table is generated through an SQL statement and a job corresponding to the SQL statement, and an abstract syntax tree is constructed based on the SQL statement; the abstract syntax tree can intuitively express the relationship between fields in the statement. The source data table and the target data table are subjected to blood relationship parsing using the abstract syntax tree to obtain a blood relationship parsing result. A directed acyclic graph is constructed according to the blood relationship parsing result and a preset one-to-one correspondence. The directed acyclic graph can asynchronously and concurrently write a lot of transactions and finally form a topological tree structure, which can greatly improve the scalability. The blood relationship corresponding to the field to be queried is obtained based on the directed acyclic graph. Therefore, the field blood relationship analysis method, device, electronic device, and computer-readable storage medium proposed by the present invention can solve the problem of low accuracy in field blood relationship analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a schematic flowchart of a field blood relationship analysis method provided by an embodiment of the present invention;

[0045] Figure 2 It is a functional module diagram of a field blood relationship analysis device provided by an embodiment of the present invention;

[0046] Figure 3 It is a schematic structural diagram of an electronic device for implementing the field blood relationship analysis method provided by an embodiment of the present invention.

[0047] The realization, functional characteristics, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] An embodiment of the present application provides a method for field lineage analysis. The execution subject of the field lineage analysis method includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the field lineage analysis method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0050] Referring to Figure 1 As shown, it is a flowchart of the field lineage analysis method provided by an embodiment of the present invention. In this embodiment, the field lineage analysis method includes:

[0051] S1. Obtain the source data table and the job set, and extract the SQL statement of any one job in the job set.

[0052] In the embodiment of the present invention, the source data table refers to the data table of the starting source and is also the basis for the data source of the downstream table. The job set is a set containing multiple SQL jobs, where the SQL jobs include: scheduled time and SQL statements.

[0053] For example, in the embodiment of the present invention, the source data table can be cust_tran, where the source data table cust_tran contains the customer number CUST_NO, the ATM transaction amount atm_am, and the POS transaction amount pos_am. The SQL statement of any one job can be select cust_no,atm-am,pos-am from cust-tran or selectcust_no,atm-am+pos-am as zz_am from TMPt-CUST.

[0054] S2. Generate a target data table according to the SQL statement and the job corresponding to the SQL statement.

[0055] In the embodiment of the present invention, generating the target data table according to the SQL statement and the job includes:

[0056] Extract the job name in the job corresponding to the SQL statement, and extract the query fields in the SQL statement;

[0057] Generate the target data table by using the job name as the table name of the initial data table and the query fields as the table fields of the initial data table.

[0058] Specifically, the target table is automatically generated by the platform. When the user creates a job, the user will input the job name. The name of the table comes from the job name, and the fields of the table come from the fields queried in the SQL written by the user.

[0059] S3. Construct an abstract syntax tree based on the SQL statement.

[0060] In the embodiment of the present invention, the constructing an abstract syntax tree based on the SQL statement includes:

[0061] Obtain a preset keyword set, and screen out the words in the SQL statement that are the same as the keywords in the keyword set to obtain standard keywords;

[0062] Use the standard keywords to segment the SQL statement to obtain multiple SQL sub-statements;

[0063] Convert the multiple SQL sub-statements into an abstract syntax tree.

[0064] Specifically, the keyword set includes a starting keyword set and a target keyword set. Among them, the starting keyword set includes 'LEFT JOIN', 'RIGHT JOIN', 'LEFT OUTER JOIN', 'RIGHT OUTER JOIN', etc., and the target keywords include "CREATE", "INSERT", "SELECT", 'INTO', 'OVERWRITE'.

[0065] Specifically, the using the standard keywords to segment the SQL statement to obtain multiple SQL sub-statements includes:

[0066] Use the standard keywords as segmentation nodes to segment the SQL statement to the left and right based on the segmentation nodes to obtain multiple SQL sub-statements; or

[0067] Segment the SQL statement based on a preset random segmentation length to obtain multiple segmented statements, and use the standard keywords as positioning points to perform secondary segmentation on the multiple segmented statements to obtain multiple SQL sub-statements.

[0068] For example, the SQL statement is "select col_a from A", and the standard keywords are "select" and "from". Therefore, the standard keywords "select" and "from" are used as splitting nodes for splitting, resulting in two SQL sub-statements: "select col_a" and "from A".

[0069] Specifically, after splitting the SQL statement using the standard keywords to obtain multiple SQL sub-statements, the method further includes:

[0070] Performing label marking on the multiple SQL sub-statements.

[0071] For example, the keyword corresponding to the SQL sub-statement "select col_a" is a keyword in the target keyword set. Therefore, the label corresponding to the SQL sub-statement "select col_a" is the target label.

[0072] Further, converting the multiple SQL sub-statements into an abstract syntax tree includes:

[0073] Analyzing the SQL sub-statement using a preset lexical analyzer to obtain multiple tokens;

[0074] Constructing a syntax tree for the multiple tokens according to a preset syntax analysis method to obtain the abstract syntax tree.

[0075] Among them, the preset syntax analysis method includes a top-down analysis method and a bottom-up analysis method.

[0076] In another embodiment of the present invention, an ANTLR tool can be used to construct the SQL statement to obtain an abstract syntax tree.

[0077] Among them, the ANTLR tool refers to an open-source syntax analyzer that can automatically generate a syntax tree based on the input and visually display it. It provides a framework for automatically constructing recognizers, compilers, and interpreters for languages including Java, C++, and C# through syntax descriptions.

[0078] S4. Using the abstract syntax tree to perform lineage parsing on the source data table and the target data table to obtain a lineage parsing result.

[0079] In the embodiment of the present invention, using the abstract syntax tree to perform lineage parsing on the source data table and the target data table to obtain a lineage parsing result includes:

[0080] Traversing the abstract syntax tree and marking the nodes in the abstract syntax tree that are consistent with the nodes of the preset fields as target nodes;

[0081] Extract the data value in the target node, compare and query based on the data value and the data table in the database, and obtain the job table corresponding to the data value;

[0082] Determine the type of the job table according to the label corresponding to the job table, determine the relationship between different types of job tables, and summarize the job table and the relationship between job tables to obtain the blood relationship analysis result.

[0083] Specifically, in the embodiment of the present invention, the preset field is table, and the value corresponding to the preset field in the target node is extracted. For example, if the table is table A, the corresponding value is A, and then the data table with the table name A in the preset database is queried to obtain the corresponding job table.

[0084] Specifically, determine the type of the job table according to the type to which the label corresponding to the job table belongs. For example, if the label type corresponding to the job table is the starting data table label, the job table is determined as the starting data table; if the label type corresponding to the data table is the target data table label, the job table is determined as the target data table. The blood relationship between the starting data table and the target data table is that the starting data table is the upstream data table of the target data table. Summarize all the starting data tables and the target data tables to obtain the blood relationship analysis result.

[0085] S5. Construct a directed acyclic graph according to the blood relationship analysis result and the preset one-to-one correspondence.

[0086] In the embodiment of the present invention, constructing a directed acyclic graph according to the blood relationship analysis result and the preset one-to-one correspondence includes:

[0087] Use the starting data table in the blood relationship analysis result as the first node and the target data table in the blood relationship analysis result as the second node;

[0088] Construct an initial node graph according to the relationship that the first node points to the second node;

[0089] According to the one-to-one correspondence, replace the job data table corresponding to each node in the initial node graph with the corresponding job in the job set to obtain the directed acyclic graph.

[0090] Specifically, the directed acyclic graph (DAG) refers to a directed graph without loops. If there is a non-directed acyclic graph, and starting from point A, going to B via C and then back to A forms a loop. Changing the direction of the edge from C to A to from A to C will turn it into a directed acyclic graph. The number of spanning trees of a directed acyclic graph is equal to the product of the in-degrees of the nodes with non-zero in-degrees. Among them, the single correspondence relationship means that since there is a one-to-one correspondence between the data table and the job, according to the single correspondence relationship, each node corresponding job data table in the initial node graph is replaced with the corresponding job in the job set to obtain the directed acyclic graph.

[0091] S6. Obtain the fields to be queried, and based on the directed acyclic graph, obtain the lineage relationship corresponding to the fields to be queried.

[0092] In an embodiment of the present invention, the directed acyclic graph is constructed from the lineage analysis result including the lineage relationship between fields. Based on the directed acyclic graph, the lineage relationship corresponding to the fields to be queried can be obtained.

[0093] Specifically, there are many usage scenarios of field lineage. Data can view the influence scope after modifying the job by analyzing the field lineage. The desensitization system can mark the desensitization types of downstream fields through the lineage system. It reduces the labor cost of viewing code to analyze lineage.

[0094] In an embodiment of the present invention, a target data table is generated through an SQL statement and the job corresponding to the SQL statement, and an abstract syntax tree is constructed based on the SQL statement; the abstract syntax tree can intuitively express the relationship between fields in the statement. The source data table and the target data table are subjected to lineage analysis using the abstract syntax tree to obtain a lineage analysis result. A directed acyclic graph is constructed according to the lineage analysis result and a preset single correspondence relationship. The directed acyclic graph can asynchronously and concurrently write a lot of transactions and finally form a topological tree structure, which can greatly improve the scalability. The lineage relationship corresponding to the fields to be queried is obtained based on the directed acyclic graph. Therefore, the field lineage analysis method proposed by the present invention can solve the problem of low accuracy of field lineage analysis.

[0095] As Figure 2 shown, it is a functional module diagram of a field lineage analysis device provided by an embodiment of the present invention.

[0096] The field lineage analysis device 100 of the present invention can be installed in an electronic device. According to the functions implemented, the field lineage analysis device 100 may include a statement extraction module 101, a data table generation module 102, a syntax tree construction module 103, and a lineage analysis module 104. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0097] In this embodiment, the functions of each module / unit are as follows:

[0098] The statement extraction module 101 is configured to obtain a source data table and a job set, and extract the SQL statement of any one job in the job set;

[0099] The data table generation module 102 is configured to generate a target data table according to the SQL statement and the job corresponding to the SQL statement;

[0100] The syntax tree construction module 103 is configured to construct an abstract syntax tree based on the SQL statement;

[0101] The lineage analysis module 104 is configured to perform lineage parsing on the source data table and the target data table by using the abstract syntax tree to obtain a lineage parsing result, construct a directed acyclic graph according to the lineage parsing result and a preset one-to-one correspondence, obtain a field to be queried, and obtain a lineage relationship corresponding to the field to be queried based on the directed acyclic graph.

[0102] Specifically, the specific implementation manners of each module of the field lineage analysis device 100 are as follows:

[0103] Step 1: Obtain a source data table and a job set, and extract the SQL statement of any one job in the job set.

[0104] In the embodiment of the present invention, the source data table refers to the data table of the starting source, and is also the basis for the data source of the downstream table. The job set is a set including multiple SQL jobs, where the SQL jobs include: scheduled time and SQL statements.

[0105] For example, in the embodiment of the present invention, the source data table may be cust_tran, where the source data table cust_tran includes a customer number CUST_NO, an ATM transaction amount atm_am, and a POS transaction amount pos_am. The SQL statement of any one job may be select cust_no,atm-am,pos-am from cust-tran or select cust_no,atm-am+pos-am as zz_am from TMPt-CUST.

[0106] Step 2: Generate a target data table according to the SQL statement and the job corresponding to the SQL statement.

[0107] In the embodiment of the present invention, generating a target data table according to the SQL statement and the job includes:

[0108] Extract the job name in the job corresponding to the SQL statement, and extract the query fields in the SQL statement;

[0109] Use the job name as the table name of the initial data table, and use the query fields as the table fields of the initial data table to generate the target data table.

[0110] Specifically, the target table is automatically generated by the platform. When the user creates a job, the user will input the job name. The name of the table comes from the job name, and the fields of the table come from the query fields in the SQL written by the user.

[0111] Step 3: Construct an abstract syntax tree based on the SQL statement.

[0112] In the embodiment of the present invention, the constructing an abstract syntax tree based on the SQL statement includes:

[0113] Obtain a preset keyword set, and filter out the words in the SQL statement that are the same as the keywords in the keyword set to obtain standard keywords;

[0114] Use the standard keywords to split the SQL statement to obtain multiple SQL sub-statements;

[0115] Convert multiple SQL sub-statements into an abstract syntax tree.

[0116] Specifically, the keyword set includes a starting keyword set and a target keyword set. Among them, the starting keyword set includes 'LEFT JOIN', 'RIGHT JOIN', 'LEFT OUTER JOIN', 'RIGHT OUTER JOIN', etc., and the target keywords include "CREATE", "INSERT", "SELECT", 'INTO', 'OVERWRITE'.

[0117] Specifically, the using the standard keywords to split the SQL statement to obtain multiple SQL sub-statements includes:

[0118] Use the standard keywords as split nodes to split the SQL statement to the left and right based on the split nodes to obtain multiple SQL sub-statements; or

[0119] Based on a preset random split length, split the SQL statement to obtain multiple split statements, and use the standard keywords as positioning points to perform secondary splitting on the multiple split statements to obtain multiple SQL sub-statements.

[0120] For example, the SQL statement is "select col_a from A", and the standard keywords are "select" and "from". Therefore, the standard keywords "select" and "from" are used as splitting nodes for splitting, resulting in two SQL sub-statements: "select col_a" and "from A".

[0121] Specifically, after splitting the SQL statement using the standard keywords to obtain multiple SQL sub-statements, the method further includes:

[0122] Performing label marking on the multiple SQL sub-statements.

[0123] For example, the keyword corresponding to the SQL sub-statement "select col_a" is a keyword in the target keyword set. Therefore, the label corresponding to the SQL sub-statement "select col_a" is the target label.

[0124] Further, the conversion of the multiple SQL sub-statements into an abstract syntax tree includes:

[0125] Analyzing the SQL sub-statements using a preset lexical analyzer to obtain multiple tokens;

[0126] Constructing a syntax tree for the multiple tokens according to a preset syntax analysis method to obtain the abstract syntax tree.

[0127] Among them, the preset syntax analysis method includes a top-down analysis method and a bottom-up analysis method.

[0128] In another embodiment of the present invention, an ANTLR tool can be used to construct the SQL statement to obtain an abstract syntax tree.

[0129] Among them, the ANTLR tool refers to an open-source syntax analyzer that can automatically generate a syntax tree based on the input and visually display it. It provides a framework for automatically constructing recognizers, compilers, and interpreters for languages including Java, C++, and C# through syntax descriptions.

[0130] Step Four: Using the abstract syntax tree to perform lineage parsing on the source data table and the target data table to obtain a lineage parsing result.

[0131] In the embodiment of the present invention, the use of the abstract syntax tree to perform lineage parsing on the source data table and the target data table to obtain a lineage parsing result includes:

[0132] Traversing the abstract syntax tree and marking the nodes that are the same as the nodes of the preset fields in the abstract syntax tree as target nodes;

[0133] Extract the data value in the target node, compare and query based on the data value and the data table in the database to obtain the job table corresponding to the data value;

[0134] Determine the type of the job table according to the label corresponding to the job table, determine the relationship between different types of job tables, and summarize the job table and the relationship between job tables to obtain the blood relationship analysis result.

[0135] Specifically, in the embodiment of the present invention, the preset field is table, extract the value corresponding to the preset field in the target node, for example: table A, then the corresponding value is A, and then query the data table with the table name A in the preset database to obtain the corresponding job table.

[0136] Specifically, determine the type of the job table according to the type to which the label corresponding to the job table belongs. For example, if the label type corresponding to the job table is the starting data table label, then determine the job table as the starting data table; if the label type corresponding to the data table is the target data table label, then determine the job table as the target data table. The blood relationship between the starting data table and the target data table is that the starting data table is the upstream data table of the target data table. Summarize all the starting data tables and the target data tables to obtain the blood relationship analysis result.

[0137] Step Five: Construct a directed acyclic graph according to the blood relationship analysis result and the preset one-to-one correspondence.

[0138] In the embodiment of the present invention, constructing a directed acyclic graph according to the blood relationship analysis result and the preset one-to-one correspondence includes:

[0139] Use the starting data table in the blood relationship analysis result as the first node and the target data table in the blood relationship analysis result as the second node;

[0140] Construct an initial node graph according to the relationship that the first node points to the second node;

[0141] According to the one-to-one correspondence, replace the job data table corresponding to each node in the initial node graph with the corresponding job in the job set to obtain the directed acyclic graph.

[0142] Specifically, the directed acyclic graph (DAG) refers to a directed graph without loops. If there is a non-directed acyclic graph, and starting from point A, going to B via C and then back to A forms a loop. Changing the direction of the edge from C to A to from A to C will turn it into a directed acyclic graph. The number of spanning trees of a directed acyclic graph is equal to the product of the in-degrees of the nodes with non-zero in-degrees. Among them, the single correspondence relationship means that since there is a one-to-one correspondence between the data table and the job, according to the single correspondence relationship, each node corresponding job data table in the initial node graph is replaced with the corresponding job in the job set to obtain the directed acyclic graph.

[0143] Step 6: Obtain the fields to be queried, and based on the directed acyclic graph, obtain the lineage relationship corresponding to the fields to be queried.

[0144] In the embodiment of the present invention, the directed acyclic graph is constructed from the lineage analysis result including the lineage relationship between fields. Based on the directed acyclic graph, the lineage relationship corresponding to the fields to be queried can be obtained.

[0145] Specifically, there are many usage scenarios for field lineage. Data can view the impact scope after modifying the job by analyzing the field lineage. The desensitization system can mark the desensitization types of downstream fields through the lineage system. Reduce the labor cost of viewing code to analyze lineage.

[0146] In the embodiment of the present invention, a target data table is generated through an SQL statement and the job corresponding to the SQL statement, and an abstract syntax tree is constructed based on the SQL statement; the abstract syntax tree can intuitively express the relationship between fields in the statement, and the source data table and the target data table are subjected to lineage analysis using the abstract syntax tree to obtain a lineage analysis result. According to the lineage analysis result and the preset single correspondence relationship, a directed acyclic graph is constructed. The directed acyclic graph can asynchronously and concurrently write a lot of transactions and finally form a topological tree structure, which can greatly improve the scalability. Based on the directed acyclic graph, the lineage relationship corresponding to the fields to be queried is obtained. Therefore, the field lineage analysis device proposed by the present invention can solve the problem of low accuracy in field lineage analysis.

[0147] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the field lineage analysis method provided by an embodiment of the present invention.

[0148] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a field lineage analysis program.

[0149] Among them, in some embodiments, the processor 10 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. By running or executing programs or modules stored in the memory 11 (such as executing a field blood relationship analysis program, etc.), and calling data stored in the memory 11, it performs various functions of the electronic device and processes data.

[0150] The memory 11 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical discs, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In some other embodiments, the memory 11 may also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can not only be used to store application software installed on the electronic device and various types of data, such as the code of the field blood relationship analysis program, etc., but also be used to temporarily store data that has been output or will be output.

[0151] The communication bus 12 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0152] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.

[0153] Figure 3 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0154] For example, although not shown, the electronic device may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device may also include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0155] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0156] The field blood relationship analysis program stored in the memory 11 in the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:

[0157] Obtain the source data table and the job set, and extract the SQL statement of any job in the job set;

[0158] Generate a target data table according to the SQL statement and the job corresponding to the SQL statement;

[0159] Construct an abstract syntax tree based on the SQL statement;

[0160] Perform lineage parsing on the source data table and the target data table by using the abstract syntax tree to obtain a lineage parsing result;

[0161] Construct a directed acyclic graph according to the lineage parsing result and a preset one-to-one correspondence;

[0162] Obtain a field to be queried, and obtain the lineage relationship corresponding to the field to be queried based on the directed acyclic graph.

[0163] Specifically, for the specific implementation method of the above instructions by the processor 10, reference may be made to the description of the relevant steps in the corresponding embodiments of the accompanying drawings, which will not be elaborated herein.

[0164] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0165] The present invention also provides a computer-readable storage medium, where the readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, it can implement:

[0166] Obtain a source data table and a job set, and extract the SQL statement of any job in the job set;

[0167] Generate a target data table according to the SQL statement and the job corresponding to the SQL statement;

[0168] Construct an abstract syntax tree based on the SQL statement;

[0169] Perform lineage parsing on the source data table and the target data table by using the abstract syntax tree to obtain a lineage parsing result;

[0170] Construct a directed acyclic graph according to the lineage parsing result and a preset one-to-one correspondence;

[0171] Obtain a field to be queried, and obtain the lineage relationship corresponding to the field to be queried based on the directed acyclic graph.

[0172] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0173] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0174] In addition, in each embodiment of the present invention, the functional modules can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0175] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0176] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

[0177] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0178] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems.

[0179] In addition, it is obvious that the term "comprising" does not exclude other units or steps, and the singular does not exclude the plural. A plurality of units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Terms such as first, second, etc. are used to denote names and do not denote any particular order.

[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for field lineage analysis, characterized in that, The method includes: Obtain the source data table and the job set, and extract the SQL statement of any one job in the job set; Generate a target data table according to the SQL statement and the job corresponding to the SQL statement; Construct an abstract syntax tree based on the SQL statement; Use the abstract syntax tree to perform lineage parsing on the source data table and the target data table to obtain a lineage parsing result, including: marking the nodes in the abstract syntax tree that are the same as the nodes of the preset fields as target nodes, determining the job table corresponding to the data value according to the data value in the target node, and summarizing the relationships between different types of job tables to obtain the lineage parsing result; Construct an initial node graph according to the starting data table in the lineage parsing result and the target data table in the lineage parsing result, and replace each node corresponding job data table in the initial node graph with the corresponding job in the job set to obtain a directed acyclic graph; Obtain the field to be queried, and obtain the lineage relationship corresponding to the field to be queried based on the directed acyclic graph.

2. The field blood relationship analysis method according to claim 1, wherein The generating the target data table according to the SQL statement and the job corresponding to the SQL statement includes: Extract the job name in the job corresponding to the SQL statement, and extract the query fields in the SQL statement; Use the job name as the table name of the initial data table and the query fields as the table fields of the initial data table to generate the target data table.

3. The field blood relationship analysis method according to claim 1, characterized in that, The constructing the abstract syntax tree based on the SQL statement includes: Obtain a preset keyword set, and filter out the words in the SQL statement that are the same as the keywords in the keyword set to obtain standard keywords; Use the standard keywords to split the SQL statement to obtain multiple SQL sub-statements; Convert the multiple SQL sub-statements into an abstract syntax tree.

4. The field blood relationship analysis method according to claim 3, characterized in that The using the standard keywords to split the SQL statement to obtain multiple SQL sub-statements includes: Using the standard keywords as splitting nodes to split the SQL statement to the left and to the right based on the splitting nodes to obtain multiple SQL sub-statements; or Based on a preset random splitting length, split the SQL statement to obtain multiple split statements, and use the standard keywords as positioning points to perform secondary splitting on the multiple split statements to obtain multiple SQL sub-statements.

5. The field blood relationship analysis method according to claim 3, wherein The converting the multiple SQL sub-statements into an abstract syntax tree includes: Use a preset lexical analyzer to analyze the SQL sub-statements to obtain multiple tokens; Construct a syntax tree for the multiple tokens according to a preset syntax analysis method to obtain the abstract syntax tree.

6. The field blood relationship analysis method according to claim 1, characterized in that The marking the nodes in the abstract syntax tree that are the same as the nodes of the preset fields as target nodes, determining the job table corresponding to the data value according to the data value in the target node, and summarizing the relationships between different types of job tables to obtain the lineage parsing result includes: Extract the data value in the target node, and perform a comparison query based on the data value and the data tables in the database to obtain the job table corresponding to the data value; Determine the type of the worksheet according to the label corresponding to the worksheet, determine the relationship between different types of worksheets, and summarize the worksheets and the relationship between the worksheets to obtain a blood relationship parsing result.

7. The field blood relationship analysis method according to claim 1, characterized in that, Construct an initial node graph according to the starting data table in the blood relationship parsing result and the target data table in the blood relationship parsing result, including: Use the starting data table in the blood relationship parsing result as the first node, and use the target data table in the blood relationship parsing result as the second node; Construct an initial node graph according to the relationship that the first node points to the second node.

8. A field blood relationship analysis device, characterized in that, The device includes: A statement extraction module, configured to obtain a source data table and a job set, and extract the SQL statement of any job in the job set; A data table generation module, configured to generate a target data table according to the SQL statement and the job corresponding to the SQL statement; A syntax tree construction module, configured to construct an abstract syntax tree based on the SQL statement; A blood relationship analysis module, configured to perform blood relationship parsing on the source data table and the target data table by using the abstract syntax tree to obtain a blood relationship parsing result, including: marking the nodes in the abstract syntax tree that are consistent with the nodes of the preset fields as target nodes, determining the worksheet corresponding to the data value according to the data value in the target node, and summarizing the relationship between different types of worksheets to obtain a blood relationship parsing result; Construct an initial node graph according to the starting data table in the blood relationship parsing result and the target data table in the blood relationship parsing result, replace each node corresponding job data table in the initial node graph with the corresponding job in the job set according to the preset single correspondence relationship to obtain a directed acyclic graph, obtain the field to be queried, and obtain the blood relationship corresponding to the field to be queried based on the directed acyclic graph.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the field blood relationship analysis method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the field blood relationship analysis method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • SQL-based data blood relationship analysis method and system

    CN111538743A