An efficient data lineage parsing method supporting multiple data dialects

By building parsers and listeners that support multiple data dialects, analyzing and generating data blood relationships, the problem of only supporting a single data dialect in the existing technology is solved, and efficient and accurate data blood relationship analysis is achieved, adapting to complex scenarios and reducing costs.

CN118964384BActive Publication Date: 2025-05-02北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411027968.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-05-02
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

The existing technology can only support a single data dialect, which is inefficient and has low accuracy, making it difficult to track and analyze data ties between multiple data sources and tools.

Method used

By obtaining the syntax rule files of multiple data dialects, a domain-specific language generation parser is built, analyzing the query statement to determine the type of data dialect, determining the analysis method based on the number of types, and optimizing the analysis process through the listener to generate data blood relationship.

Benefits of technology

It has achieved support for a variety of data dialects, improved the efficiency and accuracy of data blood relationship analysis, adapted to complex real-life scenarios, provided more comprehensive data blood relationship information, reduced development and maintenance costs, and enhanced the competitiveness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118964384B_ABST
    Figure CN118964384B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of database management technology, and in particular to an efficient data lineage parsing method that supports multiple data dialects, the method comprising obtaining several types of target data dialects, and compiling grammar rule files of each target data dialect based on the characteristics of each target data dialect; constructing a domain-specific language to parse the grammar rule files of each target data dialect to generate a parser; obtaining a query statement of a user's query request, and parsing the query statement through the parser to obtain a parsing result of the query statement; extracting the dependency relationship of several elements corresponding to the query statement in the parsing result to generate a data lineage relationship of the query statement; and outputting the data lineage relationship to a storage medium. The present invention solves the problem that the prior art can only support a single data dialect, is inefficient, and has low precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of database management, and in particular to an efficient data lineage analysis method supporting multiple data dialects. Background Art

[0002] In today's data management and analysis field, there are a variety of database and data warehouse technologies, including traditional relational databases, Hadoop ecosystem, NoSQL databases, and various cloud computing platforms. These different technical solutions are designed to solve various data management and analysis challenges and have been widely used in different scenarios.

[0003] However, in practical applications, these technical solutions usually use different data storage dialects and query languages, such as SQL, HiveQL, and Spark SQL, etc. This makes data lineage tracking more complicated and difficult, especially when analyzing large-scale data sets.

[0004] At present, there are some technologies and methods for data lineage tracking, such as collecting and analyzing data lineage information through code scanning, static analysis, and dynamic tracing. However, these methods have many defects, such as only supporting a single data dialect, low efficiency, and low accuracy. Summary of the invention

[0005] To this end, the present invention provides an efficient data lineage analysis method that supports multiple data dialects, so as to overcome the problems in the prior art that only a single data dialect can be supported, the efficiency is low, and the accuracy is low.

[0006] To achieve the above object, the present invention provides an efficient data lineage parsing method supporting multiple data dialects, comprising:

[0007] Acquire several types of target data dialects, and compile grammar rule files for each target data dialect based on the characteristics of each target data dialect;

[0008] Constructing a domain specific language to parse the grammar rule file of each of the target data dialects to generate a parser;

[0009] Obtaining a query statement of a user's query request, and analyzing the query statement to determine the number of types of data dialects in the query statement, so as to determine whether to parse the query statement in a direct parsing manner or to monitor the parsing process by setting a listener according to the number of types;

[0010] Determine the proportion of the number of listening nodes monitored by the listener in the parsing process based on the difference between the number of types and the preset number of types to determine the listening strength of the listener, and determine the eligibility of the listening strength based on the abnormality rate of several element positions under the corresponding listening strength to adjust the listening strength;

[0011] Extracting dependency relationships of several elements corresponding to the query statement in the parsing result, and generating data lineage relationships of the query statement based on the accuracy of the dependency relationships;

[0012] Determine the accuracy of data source relationships based on the tracking and recording results of data sources and data flows in the analysis results to optimize the analysis process;

[0013] Output the data lineage relationship to a storage medium.

[0014] Furthermore, the process of parsing the query statement by the parser includes:

[0015] Obtaining the type of data dialect in the query statement;

[0016] Determining the parsing of the query statement by the parser according to the number of types of the data dialect;

[0017] The parsing result of the query statement by the parser is output.

[0018] Further, determining the parsing of the query statement by the parser according to the number of types of data dialects includes determining to directly parse the query statement based on a comparison result that the number of types is less than or equal to a preset number of types.

[0019] Further, determining the parsing of the query statement by the parser according to the number of types of data dialects includes determining, based on a comparison result that the number of types is greater than a preset number of types, to set a listener to monitor the parsing process during the parsing of the query statement by the parser.

[0020] Further, under the condition of monitoring the parsing process, it is determined according to the comparison result that the quantity difference is less than or equal to the preset quantity difference that the proportion of the number of monitoring nodes in the parsing process by the listener is set to be the first proportion.

[0021] Further, under the condition of monitoring the parsing process, it is determined according to the comparison result that the quantity difference is greater than a preset quantity difference that the proportion of the number of nodes monitored by the listener in the parsing process is set to the second proportion.

[0022] Further, the determining of the parsing of the query statement by the parser according to the number of types of data dialects also includes determining that the monitoring strength does not meet the standard according to a comparison result that the abnormality rate of a number of element positions under the corresponding monitoring strength is greater than a preset abnormality rate, and adjusting the proportion of the number of monitoring nodes of the monitoring strength by an adjustment coefficient based on a comparison result that the abnormality difference is less than a preset abnormality difference;

[0023] Or, based on the comparison result that the abnormal difference value is greater than or equal to a preset abnormal difference value, it is determined to add an accessor to traverse the parsing process.

[0024] Furthermore, determining the parsing of the query statement by the parser based on the number of types of the data dialect also includes judging that the dependency does not meet the standard based on a comparison result that the accuracy of the element dependency is less than the accuracy standard, and determining a correction coefficient to correct it with the preset number of types.

[0025] Furthermore, when the parsing result of the parser on the query statement is output, the accuracy evaluation value of the data source relationship is determined according to the tracking and recording results of the data source and the data flow, so that the adjustment coefficient is optimized with the optimization coefficient under the condition that the accuracy evaluation value is less than a preset accuracy evaluation value.

[0026] Furthermore, when the parsing result of the query statement by the parser is output, the accuracy evaluation value of the data source relationship is determined according to the tracking and recording results of the data source and the data flow, so as to optimize the data amount of the target data dialect with the optimization coefficient under the condition that the accuracy evaluation value is greater than or equal to a preset accuracy evaluation value.

[0027] Compared with the prior art, the present invention has the advantage that it can better adapt to complex real-world scenarios: in actual data management and analysis scenarios, different data sources and tools usually use different dialects and query statements, which makes it more difficult to track and analyze data lineage relationships. Our data lineage analysis method that supports multiple dialects can automatically adapt to different query statements and data storage formats, improving adaptability to complex real-world scenarios.

[0028] Furthermore, the present invention can provide more comprehensive data lineage information: in the case of supporting multiple dialects, our data lineage parsing method can process more types of query statements, thereby providing more comprehensive data lineage information. This information can help users better understand the source, flow and impact of data, and provide more in-depth support for data management and analysis.

[0029] Furthermore, the present invention can reduce development and maintenance costs: by implementing the data lineage parsing method as a technical solution that supports multiple dialects, it is possible to avoid developing and maintaining independent parsing methods for each dialect. This can save development and maintenance costs and improve the scalability and adaptability of the system.

[0030] Furthermore, the present invention can enhance the competitiveness of the system: since the data lineage analysis method supporting multiple dialects can better adapt to real-life scenarios and provide more comprehensive data lineage information, it has higher competitiveness. This can help enterprises gain advantages in the fierce market competition and enhance their commercial value and reputation.

[0031] Furthermore, the present invention determines the number of types of data dialects and determines whether the parser parses the query statement directly or by setting a listener according to the number of types, so as to determine the parsing method of the data dialect according to the complexity of the data dialect in the query statement, and determines the proportion of the listener to the number of listening nodes of the parser according to the difference between the number of types and the preset number of types when the listener is set to analyze the data dialect. This improves the control accuracy of the parsing process and further improves the efficiency of data lineage analysis.

[0032] Furthermore, the present invention performs an exception analysis on the position of each element in the stored source code, that is, on the positions of several elements of the data dialect corresponding to the query statement, under the corresponding monitoring intensity. If there is an incorrect element at the current position, it is determined that the current position is abnormal, thereby calculating the abnormality rate of several element positions of the data dialect to determine the accuracy of the set monitoring intensity, and then determining the processing of the parsing process based on the abnormality rate, and adjusting the monitoring intensity of the parsing process, or adding an accessor to traverse the parsing process, thereby further improving the accuracy of the control of the parsing process, thereby improving the efficiency of Shuju blood relationship analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A flowchart of an efficient data lineage parsing method supporting multiple data dialects according to an embodiment of the present invention;

[0034] Figure 2 A flow chart of re-parsing a user query request by an efficient data lineage parsing method supporting multiple data dialects according to an embodiment of the present invention;

[0035] Figure 3 This is a flowchart of determining the parsing method in the efficient data lineage parsing method supporting multiple data dialects according to an embodiment of the present invention. DETAILED DESCRIPTION

[0036] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0037] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0038] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is merely for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0039] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0040] See also Figure 1-Figure 3 As shown, Figure 1 A flowchart of an efficient data lineage parsing method supporting multiple data dialects according to an embodiment of the present invention; Figure 2 A flow chart of re-parsing a user query request by an efficient data lineage parsing method supporting multiple data dialects according to an embodiment of the present invention; Figure 3 This is a flowchart of determining the parsing method in the efficient data lineage parsing method supporting multiple data dialects according to an embodiment of the present invention.

[0041] The embodiment of the present invention supports an efficient data lineage parsing method for multiple data dialects, including:

[0042] Step S1, obtaining several types of target data dialects, and compiling a grammar rule file for each target data dialect based on the characteristics of each target data dialect;

[0043] Step S2: constructing a domain specific language to parse the grammar rule files of each target data dialect to generate a parser;

[0044] Step S3: obtaining a query statement of the user's query request, and parsing the query statement through the parser to obtain a parsing result of the query statement;

[0045] Step S4, extracting the dependency relationship of several elements corresponding to the query statement in the parsing result to generate the data lineage relationship of the query statement;

[0046] Step S5: output the data blood relationship to a storage medium.

[0047] In an embodiment of the present invention, the process of writing a grammar rule file for each target data dialect based on the characteristics of each target data dialect includes defining common grammar rules and analysis tree nodes for representing query statements in different databases and data warehouses to generate a common abstract syntax tree, which is used to convert the user's query request into a specific query language.

[0048] In the embodiment of the present invention, the domain specific language is constructed by Ant l r4 to support the parsing of multiple data dialects, and Ant l r4 is enabled to understand and parse various query statements by constructing specific grammars, text rules and vocabularies. In Ant l r4, the file with the suffix ".g4" is a grammar rule file.

[0049] In an embodiment of the present invention, grammar rules and analysis tree nodes are input into Ant lr4 so that Ant lr4 parses the grammar rules and analysis tree nodes to generate a general abstract syntax tree, and a query request is input into the general abstract syntax tree to obtain a parsing result, which is an abstract syntax tree corresponding to the query statement of the user's query request, and the abstract syntax tree represents the structure and content of the query statement.

[0050] Specifically, the process of parsing the query statement by the parser includes:

[0051] Step S31, obtaining the type of data dialect in the query statement;

[0052] Step S32: determining the parsing of the query statement by the parser according to the number of types of the data dialect;

[0053] Step S33: output the parsing result of the query statement by the parser.

[0054] Specifically, determining the parsing of the query statement by the parser according to the number of types of data dialects includes determining the parsing method of the query statement by the parser based on a comparison result of the number of types with a preset number of types;

[0055] If the number of categories is less than or equal to the preset number of categories, the parser parses the query statement in a first parsing manner;

[0056] If the number of categories is greater than the preset number of categories, the parser parses the query statement in a second parsing manner.

[0057] In the embodiment of the present invention, the value of the preset number of types is 5, but the above value is not limited to this, and those skilled in the art may also limit the value according to actual conditions.

[0058] Specifically, the parser parsing the query statement in the first parsing manner includes directly parsing the query statement.

[0059] Specifically, the parser parsing the query statement in the second parsing mode includes setting a listener to monitor the parsing process during the parsing of the query statement by the parser, and determining the monitoring strength of the listener according to the comparison result of the quantity difference and the preset quantity difference;

[0060] If the quantity difference is less than or equal to a preset quantity difference, the analysis process is monitored at a first monitoring intensity;

[0061] If the quantity difference is greater than a preset quantity difference, a second monitoring intensity is set to monitor the analysis process.

[0062] In the embodiment of the present invention, the value of the preset data difference is 2, but the above value is not limited to this, and those skilled in the art may also limit the value according to actual conditions.

[0063] Specifically, under the first monitoring intensity, the proportion of the number of monitoring nodes in the parsing process monitored by the listener is set to a first proportion, and under the second monitoring intensity, the proportion of the number of monitoring nodes in the parsing process monitored by the listener is set to a second proportion.

[0064] In the embodiment of the present invention, the value of the first proportion is 0.3, and the value of the second proportion is 0.5, but the above values ​​are not limited thereto, and those skilled in the art may also limit the values ​​according to actual conditions.

[0065] Specifically, the determining of the parsing of the query statement by the parser according to the number of types of data dialects further includes determining whether the monitoring intensity meets the standard according to the comparison result of the abnormality rate of several element positions under the corresponding monitoring intensity and the preset abnormality rate;

[0066] If the abnormality rate is less than or equal to the preset abnormality rate, it is determined that the monitoring intensity meets the standard;

[0067] If the abnormality rate is greater than a preset abnormality rate, it is determined that the monitoring intensity does not meet the standard.

[0068] In the embodiment of the present invention, the preset abnormality rate is set to 0.01, that is, if there are 100 element positions in the parsing process, the number of abnormal element positions should not exceed 1, but the value is not limited to this. For example, when the data dialect complexity in the query statement is high, the value can be adjusted according to the actual situation.

[0069] Specifically, when it is determined that the monitoring intensity does not meet the standard, the abnormality difference between the abnormality rate and the preset abnormality rate is calculated, and based on the comparison result of the abnormality rate difference and the preset abnormality difference, the adjustment method of the monitoring intensity is determined or an accessor is added to traverse the analysis process;

[0070] If the abnormal difference is less than the preset abnormal difference, it is determined to adjust the monitoring intensity, and an adjustment coefficient of the proportion of the number of monitoring nodes to the monitoring intensity is determined.

[0071] If the abnormal difference is greater than or equal to the preset abnormal difference, it is determined to add an accessor to traverse the parsing process.

[0072] In the embodiment of the present invention, the adjustment coefficient of the monitoring intensity is calculated according to the following formula:

[0073] K=1+[(W0-W) / W]

[0074] Wherein, K represents the adjustment coefficient, W represents the abnormal difference, and W0 represents the preset abnormal difference.

[0075] In an embodiment of the present invention, the preset abnormality difference is the difference between the abnormality rate and the preset abnormality rate, the value of the preset abnormality difference is 0.1, and the adjusted proportion of the number of monitoring nodes is set to Bi×K, where Bi is the first proportion or the second proportion, and the value of i is 1 or 2.

[0076] Specifically, determining the parsing of the query statement by the parser according to the number of types of the data dialect further includes determining a correction to the preset number of types according to a comparison result of the accuracy of the dependency relationship of the elements with the accuracy standard;

[0077] If the accuracy is greater than or equal to the accuracy standard, it is determined that the dependency relationship meets the standard;

[0078] If the accuracy is less than the accuracy standard, it is determined that the dependency relationship does not meet the standard and a correction coefficient is determined to be corrected with the preset number of types.

[0079] In an embodiment of the present invention, the accuracy is the ratio of the number of correct dependencies to the total number of dependencies, the value of the accuracy standard is not less than 0.95, and the correction coefficient is determined to be 0.6-0.9. Preferably, this embodiment selects 0.75 as the correction coefficient for the preset number of types.

[0080] Specifically, when the parsing result of the parser on the query statement is output, the accuracy evaluation value of the data source relationship is determined according to the tracking and recording results of the data source and the data flow, so as to determine the optimization method of the parsing process according to the comparison result of the accuracy evaluation value and the preset accuracy evaluation value;

[0081] If the accuracy evaluation value is less than the preset accuracy evaluation value, optimizing the parsing process in a first optimization manner;

[0082] If the accuracy evaluation value is greater than or equal to the preset accuracy evaluation value, the parsing process is optimized in a second optimization manner.

[0083] In the embodiment of the present invention, the accuracy evaluation value is the sum of the accuracy of several elements of the query statement and the accuracy of the element relationship, and the preset accuracy evaluation value is 0.95.

[0084] Specifically, when the parsing process is optimized in a first optimization manner, the adjustment coefficient is optimized with the optimization coefficient;

[0085] When the parsing process is optimized in the second optimization manner, the data volume of the target data dialect is optimized with the optimization coefficient.

[0086] In the embodiment of the present invention, the value of the optimization coefficient is 1.2, but the value is not limited thereto, and those skilled in the art may also adjust the value according to actual conditions.

[0087] In the embodiment of the present invention, the dependency relationship between the elements in the query statement can be extracted by using the information in the AST to determine the data lineage relationship. Specifically, a visitor can be added to the custom parser to traverse the AST and record information such as table names and field names to infer the source and flow of the data.

[0088] In the embodiment of the present invention, in the process of determining the accuracy evaluation value of the data source relationship based on the tracking and recording results of the data source and data flow, the data source and data flow are tracked and recorded by utilizing reverse engineering and metadata management functions to further verify the correctness and effectiveness of our solution.

[0089] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

[0090] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An efficient data lineage parsing method supporting multiple data dialects, characterized in that: include: Acquire several types of target data dialects, and compile grammar rule files for each target data dialect based on the characteristics of each target data dialect; Constructing a domain specific language to parse the grammar rule file of each target data dialect to generate a parser, wherein the constructed domain specific language is constructed by Antlr4 to support the parsing of multiple data dialects; Obtaining a query statement of a user's query request, and analyzing the query statement to determine the number of types of data dialects in the query statement, so as to determine whether to parse the query statement in a direct parsing manner or to monitor the parsing process by setting a listener according to the number of types; Determining the parser to parse the query statement according to the number of types of data dialects includes determining to directly parse the query statement based on a comparison result that the number of types is less than or equal to a preset number of types; Determining the parsing of the query statement by the parser according to the number of types of data dialects includes determining, based on a comparison result that the number of types is greater than a preset number of types, to set a listener to monitor the parsing process during the parsing of the query statement by the parser; Determine the proportion of the number of listening nodes monitored by the listener in the parsing process based on the difference between the number of types and the preset number of types to determine the listening strength of the listener, and determine the eligibility of the listening strength based on the abnormality rate of several element positions under the corresponding listening strength to adjust the listening strength; Under the condition of monitoring the parsing process, determining, according to the comparison result that the quantity difference is less than or equal to the preset quantity difference, to set the proportion of the number of monitoring nodes of the listener in the parsing process to be the first proportion; Under the condition of monitoring the parsing process, determining, according to the comparison result that the quantity difference is greater than a preset quantity difference, to set the proportion of the number of nodes monitored by the listener in the parsing process to a second proportion; Determining the parsing of the query statement by the parser according to the number of types of data dialects also includes determining that the monitoring strength does not meet the standard according to a comparison result that the abnormality rate of several element positions under the corresponding monitoring strength is greater than a preset abnormality rate, calculating the abnormality difference between the abnormality rate and the preset abnormality rate, and determining to adjust the proportion of the number of monitoring nodes of the monitoring strength by an adjustment coefficient based on the comparison result that the abnormality difference is less than the preset abnormality difference; Or, based on the comparison result that the abnormal difference is greater than or equal to a preset abnormal difference, determining to add an accessor to traverse the parsing process; Extracting dependencies of several elements corresponding to the query statement in the parsing result, and generating data lineage relationships of the query statement based on the accuracy of the dependencies; Determine the accuracy of the data source relationship based on the tracking and recording results of the data source and the tracking and recording results of the data flow in the analysis results to optimize the analysis process; Output the data lineage relationship to a storage medium.

2. The efficient data lineage parsing method supporting multiple data dialects according to claim 1 is characterized in that: The process of parsing the query statement by the parser includes: Obtaining the type of data dialect in the query statement; Determining the parsing of the query statement by the parser according to the number of types of the data dialect; The parsing result of the query statement by the parser is output.

3. The efficient data lineage parsing method supporting multiple data dialects according to claim 1 is characterized in that: Determining the parsing of the query statement by the parser according to the number of types of the data dialect also includes judging that the dependency does not meet the standard based on a comparison result that the accuracy of the element dependency is less than the accuracy standard, and determining a correction coefficient to correct the preset number of types.

4. The efficient data lineage parsing method supporting multiple data dialects according to claim 2, characterized in that: When the parsing result of the parser on the query statement is output, the accuracy evaluation value of the data source relationship is determined according to the tracking and recording results of the data source and the tracking and recording results of the data flow, so as to optimize the adjustment coefficient with the optimization coefficient under the condition that the accuracy evaluation value is less than the preset accuracy evaluation value.

5. The efficient data lineage parsing method supporting multiple data dialects according to claim 2, characterized in that: When the parsing result of the parser on the query statement is output, the accuracy evaluation value of the data source relationship is determined according to the tracking and recording results of the data source and the tracking and recording results of the data flow, so as to optimize the data volume of the target data dialect with the optimization coefficient under the condition that the accuracy evaluation value is greater than or equal to the preset accuracy evaluation value.

Citation Information

Patent Citations

  • Data blood relationship acquisition method and device

    CN116166718A

  • Data processing method and device, equipment and medium

    CN116483850A