Method and system for querying database source based on blood relationship

By generating and parsing target production logs, obtaining and inserting blood ties, the problem of difficult-to-regress data origin and change process caused by the diversity of data processing engines in the prior art is solved, and fast and accurate data query is achieved.

CN120469983AInactive Publication Date: 2025-08-12ZHEJIANG MEIRI HUDONG NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510976496.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-08-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, due to the large amount of data and the diverse data processing engines used by enterprises, it is difficult to obtain data blood relationships, and it is impossible to quickly and accurately trace the origin and change process of data.

Method used

By receiving instructions from the initial database, the target production log is generated, the blood relationship between the initial database and the final database is parsed, and the blood relationship between the initial database is inserted into the target blood relationship database, and the blood relationship between the data to be queried is obtained based on the client query information.

Benefits of technology

The origin and change process of data acquisition is realized quickly and accurately, avoiding the acquisition difficulties caused by the diversity of data processing engines, and improving the efficiency of data query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469983A_ABST
    Figure CN120469983A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of database query, in particular to a method and system for querying a database source based on a blood relationship, the method is used for a calculation process of a data processing engine, and the method comprises the following steps: receiving an instruction input to a target data processing engine by an initial database; analyzing a target production log generated in the calculation process of the initial database by the target data processing engine to obtain a target blood relationship database; according to the to-be-queried data information needing to be queried by the target client, querying the blood relationship of the to-be-queried data from the target blood relationship database so as to obtain the target query information of the to-be-queried data; it can be known that according to the preset rule of the target data processing engine, the target production logs with different contents are generated, analysis is performed according to the contents of the target production logs to generate the blood relationship of the data, and the problems that the blood relationship is difficult to obtain and the origin and transition process of the data cannot be rapidly and accurately traced due to the fact that the used data processing engines are diversified are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of database query, and in particular to a method and system for querying database sources based on blood relationship. Background Art

[0002] Data lineage, also known as data lineage, data provenance, or data genealogy, refers to the relationships that naturally form between data throughout its lifecycle, from generation, processing, manipulation, integration, circulation, and eventual extinction. It records the links that generate data, which are similar to human blood relationships, hence the term "data lineage." Data lineage clearly demonstrates the upstream and downstream dependencies of data production. Data lineage is particularly crucial in enterprise-level data management and analysis, tracing the origins and evolution of data.

[0003] In the existing technology, due to the large amount of enterprise data and the diverse data processing engines used, it is difficult to obtain blood relationships, and it is impossible to quickly and accurately trace the origin and change process of the data. Summary of the Invention

[0004] In response to the above technical problems, the present invention provides a method for querying a database source based on blood relationship, which is used in the calculation process of a data processing engine and includes the following steps: S10, receiving an instruction of inputting the initial database into the target data processing engine, wherein the target data processing engine processes the initial database to obtain a final database corresponding to the initial database.

[0005] S20, obtaining the log of the initial database generated in the process of the target data processing engine calculating the initial database and using the log of the initial database as the target production log.

[0006] S30, parse the target production log, obtain the blood relationship between each initial data in the initial database and each final data in the final database, and insert the blood relationship between each initial data in the initial database and each final data in the final database into the initial blood relationship database to obtain the target blood relationship database.

[0007] S40 , according to the data information to be queried that the target client needs to query, query the blood relationship of the data to be queried from the target blood relationship database, and obtain target query information of the data to be queried according to the blood relationship of the data to be queried.

[0008] The present invention also provides a system for querying a database source based on blood relationship, the system being used in a calculation process of a data processing engine, the system comprising: The first execution module is configured to receive an instruction from an initial database and input it into a target data processing engine, wherein the target data processing engine processes the initial database to obtain a final database corresponding to the initial database.

[0009] The second execution module is configured to obtain a log of the initial database generated during the calculation of the initial database by the target data processing engine and use the log of the initial database as a target production log.

[0010] The third execution module is used to parse the target production log, obtain the blood relationship between each initial data in the initial database and each final data in the final database, and insert the blood relationship between each initial data in the initial database and each final data in the final database into the initial blood relationship database to obtain the target blood relationship database.

[0011] The fourth execution module is used to query the blood relationship of the data to be queried from the target blood relationship database according to the data information to be queried required by the target client, so as to obtain the target query information of the data to be queried according to the blood relationship of the data to be queried.

[0012] The present invention has at least the following beneficial effects: In summary, a method for querying a database source based on blood relationship is provided, and the method is used in the calculation process of a data processing engine. The method comprises the following steps: receiving an instruction from an initial database to input into a target data processing engine, wherein the target data processing engine processes the initial database to obtain a final database corresponding to the initial database; obtaining a log of the initial database generated during the calculation process of the initial database by the target data processing engine and using the log of the initial database as a target production log; parsing the target production log to obtain the blood relationship between each initial data in the initial database and each final data in the final database, and using each initial data in the initial database and each final data in the final database as a target production log; parsing the target production log to obtain the blood relationship between each initial data in the initial database and each final data in the final database, and The blood relationship between each final data is inserted into the initial blood relationship database to obtain the target blood relationship database; according to the data information to be queried that the target client needs to query, the blood relationship of the data to be queried is queried from the target blood relationship database, so as to obtain the target query information of the data to be queried based on the blood relationship of the data to be queried; it can be seen that according to the preset rules of the target data processing engine, target production logs with different contents are generated, and the data are parsed according to the content of the target production log to generate the blood relationship of the data, avoiding the use of a variety of data processing engines, which makes it difficult to obtain the blood relationship and cannot quickly and accurately trace the origin and change process of the data; at the same time, based on the blood relationship, the data source and change process of the data to be queried can be quickly and accurately obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0014] Figure 1 A flowchart of a method for querying a database source based on blood relationship provided by an embodiment of the present invention; Figure 2 A schematic diagram of the structure of a system for querying a database source based on blood relationship provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0016] Example 1 like Figure 1 As shown, the first embodiment of the present invention provides a method for querying a database source based on blood relationship, which is used in the calculation process of a data processing engine and includes the following steps: S10, receiving an instruction of inputting the initial database into the target data processing engine, wherein the target data processing engine processes the initial database to obtain a final database corresponding to the initial database.

[0017] Specifically, step S10 also includes the following steps: S101, the target data processing engine obtains the initial database and metadata of the initial database; S102, the target data processing engine processes each initial data of the initial database according to data rules preset in the target data processing engine to obtain final data corresponding to the initial data, so as to generate a final database corresponding to the initial database according to the final data corresponding to the initial data.

[0018] Furthermore, the metadata of the initial database includes each initial field name of the initial database and the metadata of each initial field name. For example, the metadata of the initial field name includes: the data type of the initial field name, the size of the initial field name, the timestamp of the initial field name, etc. For example, the data type of the initial field name is an integer or a floating point type; for example, the size of the initial field name is the size of the memory occupied.

[0019] Furthermore, the target data processing engine includes a plurality of the preset data rules, and the preset data rules are pre-set calculation rules for any initial data, wherein the calculation rule type is a calculation code type or an SQL statement type.

[0020] In a specific embodiment, step S102 also includes the following steps: obtaining the initial data ID and the initial field name corresponding to the initial data ID and the metadata corresponding to the initial field name; then executing the initial data, the initial field name corresponding to the initial data ID and the metadata corresponding to the initial field name through preset rules, and when a prompt message is issued through the preset rules, the initial data is sent to other data processing engines for calculation; wherein the prompt message is information that the final data cannot be obtained by executing the preset rules; when no prompt message is issued through the preset rules, the final data corresponding to the initial data is generated; those skilled in the art are aware of the method of processing data according to any data processing engine in the prior art, and will not repeat it here; for example, the data processing engine is SparkSQL, Flink, Hive, etc.

[0021] S20, obtaining the log of the initial database generated in the process of the target data processing engine calculating the initial database and using the log of the initial database as the target production log.

[0022] Specifically, step S20 also includes the following steps: S201: When the calculation rule type corresponding to the preset data rule of the target data engine is a calculation code type, the log of the initial database is used as the target production log. Those skilled in the art are aware of the method of obtaining the generated log in the prior art and will not be described in detail here. S202, when the calculation rule type corresponding to the preset data rule of the target data engine is an SQL statement type, the initial data uses the preset data rule of the target data processing engine, including: performing logical parsing, semantic analysis, query optimization, physical plan generation, execution and result return on the SQL statement corresponding to the preset rule of the target data engine, so as to generate a target production log according to the result of logical parsing, semantic analysis, query optimization, physical plan generation, execution and result return of the SQL statement; it is further understood that: if the preset rule of the data processing engine is an SQL statement type, the preset rule of the SQL statement type is to perform operations on a table, and the output is also a table, and the preset rule of the SQL statement type is written in a variety of ways during the writing process, resulting in it being impossible to directly extract it like the log generated by the code class model; therefore, the present invention performs logical parsing / semantic analysis / query optimization / physical plan generation / execution and result return on the SQL statement, so that the SQL statement is converted into an understandable result for execution; those skilled in the art know that any method of parsing SQL statements in the prior art falls within the scope of protection of the present invention and will not be repeated here.

[0023] Furthermore, if the preset rule of the target data engine is the SQL statement type, obtaining the parsing result of the SQL statement includes: obtaining the parsing result of the SQL statement after the physical plan is generated and before execution; further understanding is: after the physical plan is generated, a callback is performed to obtain the full amount of data processed this time, and the parsing result is obtained based on the popularity of the data.

[0024] In summary, if the preset rule of the target data engine is the calculation code type, the target production log of the target data processing engine during the calculation process is obtained, and the target blood relationship is obtained based on the target production log; if the preset rule of the target data engine is the SQL statement type, the parsing result of the SQL statement is obtained, and the target blood relationship is obtained based on the parsing result of the SQL statement; the present invention sets the preset rule of the target data engine as the calculation code type and the SQL statement type, extracts the target production log generated by the preset rule of the calculation code type to obtain the target blood relationship, and obtains the parsing result by parsing the SQL statement, thereby obtaining the target blood relationship more comprehensively.

[0025] S30, parse the target production log, obtain the blood relationship between each initial data in the initial database and each final data in the final database, and insert the blood relationship between each initial data in the initial database and each final data in the final database into the initial blood relationship database to obtain the target blood relationship database.

[0026] Specifically, the initial bloodline database is an empty set database.

[0027] Specifically, in step S30, the blood relationship data includes: initial data ID, initial data corresponding to the initial data ID, final data ID corresponding to the initial data ID, final data corresponding to the final data ID, and relevant features parsed from the target production log corresponding to the initial data ID, where the relevant features are used to represent the final data obtained by passing the initial data through the preset rules of the target data engine; the initial data ID is the unique identity identifier of the initial data; and the final data ID is the unique identity identifier of the final data.

[0028] S40 , according to the data information to be queried that the target client needs to query, query the blood relationship of the data to be queried from the target blood relationship database, and obtain target query information of the data to be queried according to the blood relationship of the data to be queried.

[0029] In a specific embodiment, step S40 further includes the following steps: S1. Obtaining information of data to be queried that a target client needs to query, wherein the information of the data to be queried includes: a starting data ID of the data to be queried and / or an ending data ID of the data to be queried.

[0030] Furthermore, the starting data ID of the data to be queried is a unique identifier of the starting data in the process of obtaining another data from a certain data through the data engine; that is, it can be understood that the ending data is the data input by the data engine.

[0031] The termination data ID of the data to be queried is a unique identifier of the termination data in the process of obtaining another data from a certain data through the data engine; that is, it can be understood that the termination data is the data output by the data engine.

[0032] S2. When the data information to be queried is the starting data ID of the data to be queried and the ending data ID of the data to be queried, the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID are queried from the target lineage database, so as to determine the target processing path between the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID based on the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID, and use the path information of the target processing path as the target query information, wherein the target query information serves as the data processing source of the data to be queried; it is further understood as: the initial data ID consistent with the starting data ID is queried from the target lineage database, and the final data ID consistent with the ending data ID is queried from the target lineage database at the same time.

[0033] Furthermore, step S2 also includes the following steps: S21, marking the initial data ID corresponding to the starting data ID as A and marking the final data ID corresponding to the ending data ID as B; S22, A and B are passed through a preset path algorithm to obtain a target processing path; those skilled in the art are aware of any path algorithm in the prior art, which will not be described here; for example, ant colony algorithm, A star algorithm, Floyd-Warshall algorithm, etc.

[0034] S3, when the data information to be queried is the starting data ID of the data to be queried or the ending data ID of the data to be queried, a preset path algorithm is used to obtain several key processing paths corresponding to the data to be queried; a person skilled in the art is aware of any path algorithm in the prior art and will not elaborate on it here; for example, the ant colony algorithm, the A star algorithm, the Floyd-Warshall algorithm, etc.; further understanding: when the data information to be queried is the starting data ID of the data to be queried, all the final data IDs are used as the ending data IDs, and the intermediate processing paths corresponding to all the final data IDs are constructed, and the intermediate processing paths with the number of path nodes less than the preset node number threshold are selected from all the intermediate processing paths as the key processing paths; or, when the data information to be queried is the ending data ID of the data to be queried, all the initial data IDs are used as the starting data IDs, and the intermediate processing paths with the number of path nodes less than the preset node number threshold are selected from all the intermediate processing paths as the key processing paths.

[0035] S4, query the endpoint data ID with the greatest frequency from all key processing paths as the data source corresponding to the data to be queried; further understood as: the endpoint data ID with the greatest frequency of appearance in all key processing paths, the endpoint data ID is the starting data ID or the ending data ID, the path information of the shortest key processing path corresponding to the endpoint data ID is used as the target query information, and the target query information and the endpoint data ID are used as the data processing sources of the data to be queried.

[0036] The first embodiment of the present invention provides a method for querying a database source based on a blood relationship, the method being used in a calculation process of a data processing engine, the method comprising the following steps: receiving an instruction from an initial database to input into a target data processing engine, wherein the target data processing engine processes the initial database to obtain a final database corresponding to the initial database; obtaining a log of the initial database generated during the calculation process of the initial database by the target data processing engine and using the log of the initial database as a target production log; parsing the target production log to obtain a blood relationship between each piece of initial data in the initial database and each piece of final data in the final database, and comparing each piece of initial data in the initial database and each piece of final data in the final database. The blood relationship between the data is inserted into the initial blood relationship database to obtain the target blood relationship database; according to the data information to be queried that the target client needs to query, the blood relationship of the data to be queried is queried from the target blood relationship database, so as to obtain the target query information of the data to be queried based on the blood relationship of the data to be queried; it can be seen that according to the preset rules of the target data processing engine, target production logs with different contents are generated, and the data are parsed according to the content of the target production log to generate the blood relationship of the data, avoiding the use of a variety of data processing engines, which makes it difficult to obtain the blood relationship and cannot quickly and accurately trace the origin and change process of the data; at the same time, based on the blood relationship, the data source and change process of the data to be queried can be quickly and accurately obtained.

[0037] Example 2 like Figure 2 As shown, the second embodiment of the present invention provides a system for querying a database source based on blood relationship, the system is used in the calculation process of a data processing engine, and the system includes: The first execution module is configured to receive an instruction from an initial database and input it into a target data processing engine, wherein the target data processing engine processes the initial database to obtain a final database corresponding to the initial database.

[0038] Specifically, the first execution module includes: A first acquisition module is used for the target data processing engine to acquire the initial database and metadata of the initial database; The second acquisition module is used for the target data processing engine to process each initial data of the initial database according to the data rules preset in the target data processing engine to obtain the final data corresponding to the initial data, so as to generate the final database corresponding to the initial database according to the final data corresponding to the initial data.

[0039] Furthermore, the metadata of the initial database includes each initial field name of the initial database and the metadata of each initial field name. For example, the metadata of the initial field name includes: the data type of the initial field name, the size of the initial field name, the timestamp of the initial field name, etc. For example, the data type of the initial field name is an integer or a floating point type; for example, the size of the initial field name is the size of the memory occupied.

[0040] Furthermore, the target data processing engine includes several preset data rules, wherein the preset data rules are pre-set calculation rules for any initial data, wherein the calculation rule type is a calculation code type or an SQL statement type.

[0041] In a specific embodiment, the second acquisition module executes the steps of: obtaining the initial data ID and the initial field name corresponding to the initial data ID and the metadata corresponding to the initial field name; then executing the initial data, the initial field name corresponding to the initial data ID and the metadata corresponding to the initial field name through preset rules, and when a prompt message is issued through the preset rules, the initial data is sent to other data processing engines for calculation; wherein, the prompt message is information that the final data cannot be obtained by executing the preset rules; when the prompt message is not issued through the preset rules, the final data corresponding to the initial data is generated; those skilled in the art are aware of the method of data processing according to any data processing engine in the prior art, and will not repeat it here; for example, the data processing engine is SparkSQL, Flink, Hive, etc.

[0042] The second execution module is configured to obtain a log of the initial database generated during the calculation of the initial database by the target data processing engine and use the log of the initial database as a target production log.

[0043] Specifically, the second execution module includes: A first generation module is configured to use the log of the initial database as the target production log when the calculation rule type corresponding to the preset data rule of the target data engine is a calculation code type. Persons skilled in the art are familiar with the method of obtaining the generated log in the prior art and will not be described in detail here. The second generation module is used for when the calculation rule type corresponding to the preset data rule of the target data engine is an SQL statement type. The initial data uses the preset data rule of the target data processing engine, including: logical parsing, semantic analysis, query optimization, physical plan generation, execution and result return of the SQL statement corresponding to the preset rule, so as to generate a target production log according to the result of logical parsing, semantic analysis, query optimization, physical plan generation, execution and result return of the SQL statement; it is further understood that: if the preset rule of the data processing engine is an SQL statement type, the preset rule of the SQL statement type is to perform operations on a table, and the output is also a table, and the preset rule of the SQL statement type is written in a variety of ways during the writing process, resulting in it being impossible to directly extract it like the log generated by the code class model; therefore, the present invention performs logical parsing / semantic analysis / query optimization / physical plan generation / execution and result return on the SQL statement, so that the SQL statement is converted into an understandable result for execution; those skilled in the art know that any method of parsing SQL statements in the prior art falls within the scope of protection of the present invention and will not be repeated here.

[0044] Furthermore, if the preset rule of the target data engine is the SQL statement type, obtaining the parsing result of the SQL statement includes: obtaining the parsing result of the SQL statement after the physical plan is generated and before execution; further understanding is: after the physical plan is generated, a callback is performed to obtain the full amount of data processed this time, and the parsing result is obtained based on the popularity of the data.

[0045] In summary, if the preset rule of the target data engine is the calculation code type, the target production log of the target data processing engine during the calculation process is obtained, and the target blood relationship is obtained based on the target production log; if the preset rule of the target data engine is the SQL statement type, the parsing result of the SQL statement is obtained, and the target blood relationship is obtained based on the parsing result of the SQL statement; the present invention sets the preset rule of the target data engine as the calculation code type and the SQL statement type, extracts the target production log generated by the preset rule of the calculation code type to obtain the target blood relationship, and obtains the parsing result by parsing the SQL statement, thereby obtaining the target blood relationship more comprehensively.

[0046] The third execution module is used to parse the target production log, obtain the blood relationship between each initial data in the initial database and each final data in the final database, and insert the blood relationship between each initial data in the initial database and each final data in the final database into the initial blood relationship database to obtain the target blood relationship database.

[0047] Specifically, the initial bloodline database is an empty set database.

[0048] Specifically, the blood relationship data in the third execution module includes: initial data ID, initial data corresponding to the initial data ID, final data ID corresponding to the initial data ID, final data corresponding to the final data ID, and relevant features parsed from the target production log corresponding to the initial data ID. The relevant features are used to represent the final data obtained by passing the initial data through the preset rules of the target data engine; the initial data ID is the unique identity identifier of the initial data; and the final data ID is the unique identity identifier of the final data.

[0049] The fourth execution module is used to query the blood relationship of the data to be queried from the target blood relationship database according to the data information to be queried required by the target client, so as to obtain the target query information of the data to be queried according to the blood relationship of the data to be queried.

[0050] In a specific embodiment, the fourth execution module includes: The module for obtaining the data information to be queried is used to obtain the data information to be queried that the target client needs to query, wherein the data information to be queried includes: the starting data ID of the data to be queried and / or the ending data ID of the data to be queried.

[0051] Furthermore, the starting data ID of the data to be queried is a unique identifier of the starting data in the process of obtaining another data from a certain data through the data engine; that is, it can be understood that the ending data is the data input by the data engine.

[0052] The termination data ID of the data to be queried is a unique identifier of the termination data in the process of obtaining another data from a certain data through the data engine; that is, it can be understood that the termination data is the data output by the data engine.

[0053] The target query information acquisition module is used to query the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID from the target lineage database when the data information to be queried is the starting data ID of the data to be queried and the ending data ID of the data to be queried, so as to determine the target processing path between the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID according to the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID, and use the path information of the target processing path as the target query information, wherein the target query information serves as the data processing source of the data to be queried; it is further understood as: querying the initial data ID consistent with the starting data ID from the target lineage database, and at the same time querying the final data ID consistent with the ending data ID from the target lineage database.

[0054] Furthermore, the target query information acquisition module includes: a marking module, configured to mark an initial data ID corresponding to the start data ID as A and mark a final data ID corresponding to the end data ID as B; The target processing path acquisition module is used to obtain the target processing path by passing A and B through a preset path algorithm; those skilled in the art are aware of any path algorithm in the prior art, which will not be described here; for example, the ant colony algorithm, the A star algorithm, the Floyd-Warshall algorithm, etc.

[0055] The key processing path acquisition module is used to obtain several key processing paths corresponding to the data to be queried through a preset path algorithm when the data information to be queried is the starting data ID of the data to be queried or the ending data ID of the data to be queried; those skilled in the art are aware of any path algorithm in the prior art and will not repeat them here; for example, the ant colony algorithm, the A star algorithm, the Floyd-Warshall algorithm, etc.; further understanding: when the data information to be queried is the starting data ID of the data to be queried, all the final data IDs are used as the ending data IDs, and the intermediate processing paths corresponding to all the final data IDs are constructed, and the intermediate processing paths whose path node number is less than the preset node number threshold are selected from all the intermediate processing paths as the key processing paths.

[0056] The data source acquisition module is used to query the endpoint data ID with the greatest frequency from all key processing paths as the data source corresponding to the data to be queried; this is further understood as: the endpoint data ID with the greatest frequency of appearance in all key processing paths, where the endpoint data ID is the starting data ID or the ending data ID, and the path information of the shortest key processing path corresponding to the endpoint data ID is used as the target query information, and the target query information and the endpoint data ID are used as the data processing source for the data to be queried.

[0057] The second embodiment of the present invention provides a system for querying a database source based on a blood relationship, the system being used in a calculation process of a data processing engine, the system comprising: a first execution module for receiving an instruction from an initial database and inputting it into a target data processing engine, wherein the target data processing engine processes the initial database to obtain a final database corresponding to the initial database; a second execution module for obtaining a log of the initial database generated during the calculation process of the initial database by the target data processing engine and using the log of the initial database as a target production log; a third execution module for parsing the target production log, obtaining the blood relationship between each piece of initial data in the initial database and each piece of final data in the final database, and using each piece of initial data in the initial database and each piece of final data in the final database as a target production log; The blood relationship between each final data in the library is inserted into the initial blood relationship database to obtain the target blood relationship database; the fourth execution module is used to query the blood relationship of the data to be queried from the target blood relationship database according to the data information to be queried that the target client needs to query, so as to obtain the target query information of the data to be queried based on the blood relationship of the data to be queried; it can be seen that according to the preset rules of the target data processing engine, target production logs with different contents are generated, and the data are parsed according to the content of the target production log to generate the blood relationship of the data, avoiding the use of a variety of data processing engines, which makes it difficult to obtain the blood relationship and cannot quickly and accurately trace the origin and change process of the data; at the same time, based on the blood relationship, the data source and change process of the data to be queried can be quickly and accurately obtained.

[0058] Although some specific embodiments of the present invention have been described in detail by way of example, it will be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It will also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. A method for querying a database source based on blood relationship, the method being used in a calculation process of a data processing engine, characterized in that: The method comprises the following steps: S10, receiving an instruction for inputting the initial database into the target data processing engine, wherein the target data processing engine processes the initial database to obtain a final database corresponding to the initial database; S20, obtaining the log of the initial database generated during the calculation of the initial database by the target data processing engine and using the log of the initial database as the target production log; S30, parsing the target production log to obtain the lineage relationship between each piece of initial data in the initial database and each piece of final data in the final database, and inserting the lineage relationship between each piece of initial data in the initial database and each piece of final data in the final database into the initial lineage database to obtain the target lineage database; S40 , according to the data information to be queried that the target client needs to query, query the blood relationship of the data to be queried from the target blood relationship database, and obtain target query information of the data to be queried according to the blood relationship of the data to be queried.

2. The method for querying a database source based on blood relationship according to claim 1, characterized in that: The step S10 also includes the following steps: S101, the target data processing engine obtains the initial database and metadata of the initial database; S102, the target data processing engine processes each initial data of the initial database according to data rules preset in the target data processing engine to obtain final data corresponding to the initial data, so as to generate a final database corresponding to the initial database according to the final data corresponding to the initial data.

3. The method for querying a database source based on blood relationship according to claim 2, characterized in that: The target data processing engine includes a plurality of preset data rules, and the preset data rules are pre-set calculation rules for any initial data, wherein the calculation rule type is a calculation code type or an SQL statement type.

4. The method for querying a database source based on blood relationship according to claim 3, characterized in that: The following steps are also included in step S20: S201, when the calculation rule type corresponding to the preset data rule of the target data engine is a calculation code type, the log of the initial database is used as the target production log; S202, when the calculation rule type corresponding to the preset data rule of the target data engine is an SQL statement type, the initial data uses the preset data rule of the target data processing engine to include: performing logical parsing, semantic analysis, query optimization, physical plan generation, execution and result return on the SQL statement corresponding to the preset data rule, so as to generate a target production log based on the results of logical parsing, semantic analysis, query optimization, physical plan generation, execution and result return of the SQL statement.

5. The method for querying a database source based on blood relationship according to claim 1, characterized in that: S40 includes the following steps: S1, obtaining the data information to be queried that the target client needs to query, wherein the data information to be queried includes: the starting data ID of the data to be queried and / or the ending data ID of the data to be queried; S2, when the data information to be queried is the starting data ID of the data to be queried and the ending data ID of the data to be queried, the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID are queried from the target lineage database, so as to determine the target processing path between the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID according to the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID, and use the path information of the target processing path as the target query information, wherein the target query information serves as the data processing source of the data to be queried; S3, when the data to be queried is the starting data ID of the data to be queried or the ending data ID of the data to be queried, a preset path algorithm is used to obtain several key processing paths corresponding to the data to be queried; S4, querying the endpoint data ID with the highest frequency from all key processing paths as the data source corresponding to the data to be queried.

6. A system for querying database sources based on blood relationship, the system being used in the calculation process of a data processing engine, characterized in that: The system comprises: A first execution module is configured to receive an instruction from an initial database and input it into a target data processing engine, wherein the target data processing engine processes the initial database to obtain a final database corresponding to the initial database; The second execution module is configured to obtain the log of the initial database generated during the calculation of the initial database by the target data processing engine and use the log of the initial database as the target production log; The third execution module is used to parse the target production log, obtain the blood relationship between each initial data in the initial database and each final data in the final database, and insert the blood relationship between each initial data in the initial database and each final data in the final database into the initial blood relationship database to obtain the target blood relationship database; The fourth execution module is used to query the blood relationship of the data to be queried from the target blood relationship database according to the data information to be queried required by the target client, so as to obtain the target query information of the data to be queried according to the blood relationship of the data to be queried.

7. The system for querying database sources based on blood relationship according to claim 6, characterized in that: The first execution module includes: A first acquisition module is used for the target data processing engine to acquire the initial database and metadata of the initial database; The second acquisition module is used for the target data processing engine to process each initial data of the initial database according to the data rules preset in the target data processing engine to obtain the final data corresponding to the initial data, so as to generate the final database corresponding to the initial database according to the final data corresponding to the initial data.

8. The system for querying database sources based on blood relationship according to claim 7, characterized in that: The target data processing engine includes a plurality of preset data rules, and the preset data rules are pre-set calculation rules for any initial data, wherein the calculation rule type is a calculation code type or an SQL statement type.

9. The system for querying database sources based on blood relationship according to claim 8, characterized in that: The second execution module includes: A first generating module is configured to use the log of the initial database as the target production log when the calculation rule type corresponding to the preset data rule of the target data engine is a calculation code type; The second generation module is used for, when the calculation rule type corresponding to the preset data rule of the target data engine is the SQL statement type, the initial data uses the preset data rule of the target data processing engine, including: performing logical parsing, semantic analysis, query optimization, physical plan generation, execution and result return on the SQL statement corresponding to the preset data rule, so as to generate a target production log according to the results of logical parsing, semantic analysis, query optimization, physical plan generation, execution and result return of the SQL statement.

10. The system for querying database sources based on blood relationship according to claim 6, characterized in that: The fourth execution module includes: The module for obtaining the data information to be queried is used to obtain the data information to be queried that the target client needs to query, wherein the data information to be queried includes: the starting data ID of the data to be queried and / or the ending data ID of the data to be queried; a target query information acquisition module, for, when the data information to be queried is the starting data ID of the data to be queried and the ending data ID of the data to be queried, querying the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID from the target lineage database, so as to determine the target processing path between the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID according to the initial data ID corresponding to the starting data ID and the final data ID corresponding to the ending data ID, and using the path information of the target processing path as the target query information, wherein the target query information serves as the data processing source of the data to be queried; The key processing path acquisition module is used to obtain several key processing paths corresponding to the data to be queried by using a preset path algorithm when the data information to be queried is the starting data ID of the data to be queried or the ending data ID of the data to be queried; The data source acquisition module is used to query the endpoint data ID with the highest frequency from all key processing paths as the data source corresponding to the data to be queried.

Citation Information

Patent Citations

  • MySQL-based blood relationship analysis method, electronic equipment and medium

    CN117951230A

  • Method for analyzing blood relationship between subtables in data warehouse system

    CN118113718A

  • Query method and device for blood relationship data

    CN118585544A