A method and apparatus for determining data lineage, a business processing method, and a device and equipment
By determining the hierarchical relationship in the data to be analyzed and gradually deleting irrelevant fields, the problem of difficulty in determining data blood relationship is solved, and rapid and accurate data blood relationship determination and business processing efficiency are achieved.
Patent Information
- Application Number
- CN202310787345.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-06-29
AI Technical Summary
In data asset management, it is difficult for the prior art to effectively determine the link relationship between data, resulting in low data call efficiency, complex business processing and inability to clearly converge, affecting the business processing process and user experience in actual applications.
By obtaining the data to be analyzed, determining the hierarchical relationship between multiple fields, selecting the top-level target field, deleting irrelevant fields, and repeating the above process until all fields are traversed to determine the blood relationship of the data.
It realizes the rapid and accurate determination of data blood relationships, reduces the amount of data processing, improves business processing efficiency, and ensures business processing results.
Smart Images

Figure CN116821229B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of big data technology, and particularly to a method for determining data lineage relationships, a business processing method, an apparatus, and a device. Background Art
[0002] In data asset management, it is necessary to determine the link relationships between data, and then effectively call the data according to the association relationships of the data. Specifically, the method of data lineage analysis can be used to effectively determine the relationships between data, and then realize metadata management. Data lineage, also known as data pedigree, data origin, and data lineage, refers to a relationship that naturally forms between data during the entire life cycle of data, from generation, processing, processing, fusion, transfer to final extinction. It records the link relationships of data generation, and these relationships are similar to human blood relationships, so they are called data lineage relationships. Based on data lineage, it can effectively assist in business processing. For example, when processing a specific business, it can directly call the data based on the business and find the associated data corresponding to the data lineage, so as to conveniently and effectively complete the business processing.
[0003] However, there are often great difficulties in applying data lineage to determine data relationships at present. In actual applications, due to factors such as user writing habits and different specification methods, there is a lack of clear corresponding relationships between data. When using the direct tracing method to determine the lineage relationship, it is very easy to lose key lineage information, making the tracing process of data lineage not only complex but also unable to converge clearly, and it cannot achieve a good data relationship determination effect. Correspondingly, it is also impossible to use data lineage to ensure the business processing effect, affecting the business processing process in actual applications and even affecting the user experience. Therefore, there is an urgent need for a technical solution that can effectively determine data lineage to assist in effective business processing. Summary of the Invention
[0004] The purpose of the embodiments of this specification is to provide a method for determining data lineage relationships, a business processing method, an apparatus, and a device to solve the problem of how to effectively determine data lineage to assist in effective business processing.
[0005] To solve the above technical problems, an embodiment of this specification provides a method for determining data lineage, including: obtaining data to be analyzed; the data to be analyzed contains multiple fields; determining the hierarchical relationship between the multiple fields; the hierarchical relationship includes the one-way association relationship between each field; based on the hierarchical relationship, selecting at least one top-level target field from the multiple fields; determining irrelevant fields that have no association relationship with each top-level target field; deleting fields that have a one-way association relationship with the irrelevant fields from the data to be analyzed; based on the hierarchical relationship, selecting lower-level target fields corresponding to the top-level target fields from the data to be analyzed after deletion as new top-level target fields, and repeating the operations of determining irrelevant fields, deleting fields that have a one-way association relationship with the irrelevant fields, and selecting new top-level target fields until all fields are traversed; determining the hierarchical relationship in the data to be analyzed obtained after traversing all fields and deleting as the data lineage corresponding to the top-level target fields.
[0006] In some embodiments, the data to be analyzed includes SQL statements; determining the hierarchical relationship between the multiple fields includes: converting the data to be analyzed into an abstract syntax tree; the abstract syntax tree is used to represent the hierarchical relationship between fields.
[0007] Based on the above embodiment, deleting fields that have a one-way association relationship with the irrelevant fields from the data to be analyzed includes: deleting fields associated with the irrelevant fields based on the hierarchical relationship in the abstract syntax tree.
[0008] In some embodiments, determining irrelevant fields that have no association relationship with each top-level target field includes: respectively screening out irrelevant fields that have no association relationship with the top-level target field from the lower-level target fields corresponding to each top-level target field; correspondingly, deleting fields that have a one-way association relationship with the irrelevant fields from the data to be analyzed includes: pruning fields that have a one-way association relationship with the irrelevant fields from the data to be analyzed; using the remaining lower-level target fields in the data to be analyzed as new top-level target fields, and repeating the operations of screening out irrelevant fields and pruning fields until all levels of fields are traversed.
[0009] In some embodiments, determining irrelevant fields that have no association relationship with each top-level target field includes: determining the association relationship between each lower-level target field and the top-level target field based on a preset inference logic; the preset inference logic is used to determine the association relationship based on the semantic conflict relationship.
[0010] In some embodiments, determining the hierarchical relationship in the to-be-analyzed data after deletion as the data lineage corresponding to the top-level target field includes: when the fields in the to-be-analyzed data after deletion are from a preset data table, determining that the fields in the to-be-analyzed data after deletion correspond to lineage information; and taking the hierarchical relationship between the fields corresponding to the lineage information as the data lineage corresponding to the top-level target field.
[0011] An embodiment of this specification also provides a service processing method based on data lineage, including: determining the service field data involved in the target service; obtaining the data lineage corresponding to the service field data; the data lineage is obtained by the following method: obtaining the to-be-analyzed data; the to-be-analyzed data contains multiple fields; determining the hierarchical relationship between the multiple fields; the hierarchical relationship includes the unidirectional association relationship between each field; selecting the service field data from the multiple fields based on the hierarchical relationship; determining the irrelevant fields that have no association relationship with the service field data; deleting the fields in the to-be-analyzed data that have a unidirectional association relationship with the irrelevant fields; selecting the lower-level target field corresponding to the service field data from the to-be-analyzed data after deletion based on the hierarchical relationship as the new top-level target field, and repeating the operations of determining the irrelevant fields, deleting the fields that have a unidirectional association relationship with the irrelevant fields, and selecting the new top-level target field until all fields are traversed; taking the hierarchical relationship in the to-be-analyzed data obtained after deleting all fields as the data lineage corresponding to the service field data; extracting the associated field data corresponding to the service field data based on the data lineage; and processing the target service using the associated field data.
[0012] The embodiment of this specification also proposes a data lineage determination device, including: a to-be-analyzed data acquisition module, configured to acquire to-be-analyzed data; the to-be-analyzed data contains multiple fields; a hierarchical relationship determination module, configured to determine the hierarchical relationship between the multiple fields; the hierarchical relationship includes the one-way association relationship between each field; a top-level target field selection module, configured to select at least one top-level target field from the multiple fields based on the hierarchical relationship; an irrelevant field determination module, configured to determine the irrelevant fields that have no association relationship with each top-level target field; a field deletion module, configured to delete the fields that have a one-way association relationship with the irrelevant fields from the to-be-analyzed data; a traversal module, configured to select the lower-level target fields corresponding to the top-level target fields from the to-be-analyzed data after deletion based on the hierarchical relationship as new top-level target fields, and repeatedly execute the operations of determining irrelevant fields, deleting the fields that have a one-way association relationship with the irrelevant fields, and selecting new top-level target fields until all fields are traversed; a data lineage determination module, configured to determine the hierarchical relationship in the to-be-analyzed data obtained after deleting all fields as the data lineage corresponding to the top-level target fields.
[0013] The embodiment of this specification also proposes a service processing device based on data lineage, including: a service field data determination module, configured to determine the service field data involved in the target service; a data lineage acquisition module, configured to acquire the data lineage corresponding to the service field data; the data lineage is acquired through the following method: acquire to-be-analyzed data; the to-be-analyzed data contains multiple fields; determine the hierarchical relationship between the multiple fields; the hierarchical relationship includes the one-way association relationship between each field; select the service field data from the multiple fields based on the hierarchical relationship; determine the irrelevant fields that have no association relationship with the service field data; delete the fields that have a one-way association relationship with the irrelevant fields from the to-be-analyzed data; select the lower-level target fields corresponding to the service field data from the to-be-analyzed data after deletion based on the hierarchical relationship as new top-level target fields, and repeatedly execute the operations of determining irrelevant fields, deleting the fields that have a one-way association relationship with the irrelevant fields, and selecting new top-level target fields until all fields are traversed; use the hierarchical relationship in the to-be-analyzed data obtained after deleting all fields as the data lineage corresponding to the service field data; an associated field extraction module, configured to extract the associated field data corresponding to the service field data based on the data lineage; a target service processing module, configured to process the target service by using the associated field data.
[0014] An embodiment of this specification also provides an electronic device, including a memory and a processor; the memory is used to store computer programs / instructions; the processor is used to execute the computer programs / instructions to implement the steps of the above-mentioned data lineage determination method and / or the service processing method based on data lineage.
[0015] An embodiment of this specification also provides a computer-readable storage medium, on which computer programs / instructions are stored, and when the computer programs / instructions are executed by a processor, the steps of the above-mentioned data lineage determination method and / or the service processing method based on data lineage are implemented.
[0016] An embodiment of this specification also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the above-mentioned data lineage determination method and / or the service processing method based on data lineage are implemented.
[0017] As can be seen from the technical solutions provided by the embodiments of this specification above, the embodiments of this specification first determine the hierarchical relationship between fields based on the fields included in the data to be analyzed. Then, according to the top-level target fields in these fields, irrelevant fields that have no association relationship with the top-level target fields are determined based on the hierarchical relationship, and then the fields that have a one-way association relationship with the irrelevant fields are deleted from the data to be analyzed. After the above processing, the hierarchical relationship reflected in the remaining data to be analyzed can be reflected as the data lineage corresponding to the top-level target fields. The above method, by deleting irrelevant fields, while gradually obtaining the data lineage, also gradually reduces the amount of data that needs to be processed, thereby ensuring that the data lineage can be determined quickly and accurately, and then using the data lineage to quickly and effectively assist in the processing of related services, ensuring the processing effect of the services. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a flowchart of a data lineage determination method according to an embodiment of this specification;
[0020] Figure 2 It is a flowchart of a service processing method based on data lineage according to an embodiment of this specification;
[0021] Figure 3It is a module diagram of a data lineage determination device according to an embodiment of this specification;
[0022] Figure 4 It is a module diagram of a service processing device based on data lineage according to an embodiment of this specification. Specific implementation manners
[0023] Next, the technical solutions in the embodiments of this specification will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this specification without creative efforts shall fall within the protection scope of this specification.
[0024] To solve the above technical problems, an embodiment of this specification proposes a data lineage determination method. The execution subject of the data lineage determination method is the corresponding electronic device, and the electronic device includes but is not limited to a server, a distributed processing system, an industrial control computer, a PC, etc. As Figure 1 shown, the data lineage determination method may include the following specific implementation steps.
[0025] S110: Obtain data to be analyzed; the data to be analyzed contains multiple fields.
[0026] The data to be analyzed is the data for which the hierarchical relationship needs to be determined. The data to be analyzed may be data such as text data and code data. The association relationship between the data is not directly reflected in the data to be analyzed, so the data to be analyzed needs to be processed to obtain the relevant data lineage.
[0027] The data to be analyzed contains multiple fields, and generally, a field is the smallest analysis unit in the data to be analyzed. Therefore, the corresponding fields can be extracted respectively, and the data lineage between these fields can be determined.
[0028] In some implementation manners, the data to be analyzed is an SQL statement. Since in actual applications, SQL statements are not written based on the data lineage method, it is necessary to actively sort out the data lineage from the SQL statements. At the same time, since SQL statements may be written by different programmers respectively, the SQL statements may have different writing specifications, and it is also difficult to directly extract the data lineage from the SQL statements based on a fixed format.
[0029] S120: Determine the hierarchical relationship between the multiple fields; the hierarchical relationship includes the one-way association relationship between each field.
[0030] The hierarchical relationship is used to describe the association relationship between fields. Generally, the association relationship is a unidirectional association relationship, that is, there is only an association relationship from field A to field B. Through the hierarchical relationship, the overall relationship between fields in the data can be reflected. However, this hierarchical relationship cannot accurately and effectively reflect the data lineage relationship corresponding to a certain field. Therefore, it is also necessary to screen out the data lineage relationship from the hierarchical relationship.
[0031] Based on the implementation manner in step S110, when the data to be analyzed is an SQL statement, to determine the hierarchical relationship of these SQL statements, the SQL statements can be converted into an abstract syntax tree. The abstract syntax tree is the tree-like representation result of the abstract syntax structure of the source code of the SQL statement, and each node on the tree can be used to represent a structure of the source code. In the embodiments of this specification, each node on the abstract syntax tree can correspond to different fields in the SQL statement. Through the abstract syntax tree, the syntactic relationship between different fields in the SQL statement can be clearly and intuitively determined, that is, the hierarchical relationship corresponding to the fields in the data to be analyzed.
[0032] The specific process of converting the SQL statement into an abstract syntax tree can be set according to the actual application situation and will not be elaborated here.
[0033] S130: Select at least one top-level target field from the multiple fields based on the hierarchical relationship.
[0034] Based on the hierarchical relationship, at least one top-level target field can be selected. The top-level target field is the field at the topmost level in the hierarchical relationship, that is, this field only has an indication relationship and no pointing relationship. In some cases, multiple top-level target fields can be selected, and then the data lineage relationships of each top-level field can be determined respectively by successively following the following process for each top-level field.
[0035] S140: Determine the irrelevant fields that have no association relationship with each top-level target field.
[0036] To determine the top-level target field, a top-level target field can be selected first, and the irrelevant fields that have no association relationship with this top-level target field can be determined.
[0037] Specifically, the irrelevant fields that have no association relationship with the top-level target field can be screened out from the lower-level target fields corresponding to each top-level target field respectively. Since the hierarchical relationship can reflect the association relationship between fields, the irrelevant fields can be determined according to the hierarchical relationship, that is, the fields with no pointing relationship are the irrelevant fields.
[0038] In some embodiments, determining the irrelevant fields may also be based on a preset inference logic to determine the association relationship between each lower-layer target field and the top-layer target field; the preset inference logic is used to determine the association relationship based on the semantic conflict relationship. For example, if the reference name does not directly appear in field C, which may originate from field A or field B, then the result set of field A or B can be used to determine whether there is a semantic conflict, and then the source relationship between field C and field A or B can be deduced inversely. The specific preset inference logic can be set according to the actual application situation, and no limitation is imposed thereon.
[0039] S150: Delete the fields in the data to be analyzed that have a one-way association relationship with the irrelevant fields.
[0040] After determining the irrelevant fields, the fields that have a one-way association relationship with the irrelevant fields can be deleted. These deleted fields can be determined according to the hierarchical relationship. If there is a one-way association relationship between the irrelevant field and the corresponding field, the subsequent fields must also be irrelevant fields of the top-layer target field, and thus can be directly deleted.
[0041] In addition, when deleting the irrelevant fields, if there are fields in the lower-layer fields that have a one-way association relationship with the deleted fields, they can be deleted together to streamline the data as much as possible and ensure the quick and effective search for the data lineage corresponding to the top-layer target field.
[0042] S160: Based on the hierarchical relationship, select the lower-layer target fields corresponding to the top-layer target field from the data to be analyzed after deletion as the new top-layer target fields, and repeat the operations of determining the irrelevant fields, deleting the fields that have a one-way association relationship with the irrelevant fields, and selecting the new top-layer target fields until all fields are traversed.
[0043] In addition, if, based on the method in step S140, after determining the irrelevant fields in the lower-layer target fields, the remaining lower-layer target fields in the data to be analyzed can also be used as the new top-layer target fields, and the operations of determining the irrelevant fields, deleting the fields that have a one-way association relationship with the irrelevant fields, and selecting the new top-layer target fields are repeated until all levels of fields are traversed. Since among the remaining data after deletion, there may also be fields that are not related among the fields at other levels, by traversing and deleting the fields at other levels, the irrelevant fields can be further reduced, and thus the data lineage relationship reflected in the hierarchical relationship can be further streamlined in this way.
[0044] Specifically, when the hierarchical relationship is an abstract syntax tree, deleting the fields that have a one-way association with the irrelevant fields can be to delete the fields associated with the irrelevant fields based on the hierarchical relationship in the abstract syntax tree. The specific operation process can be set according to the actual application situation and will not be elaborated here.
[0045] S170: Determine the hierarchical relationship in the data to be analyzed obtained after deleting all fields as the data lineage corresponding to the top-level target field.
[0046] After deleting the irrelevant fields in the data to be analyzed, the hierarchical relationship reflected in the data to be analyzed at this time can be determined as the data lineage. Different top-level target fields can have different data lineages and are respectively saved corresponding to different fields for subsequent invocation.
[0047] In some embodiments, to determine the data lineage, it can also be determined whether the fields in the data to be analyzed after deletion come from a preset data table. The preset data table can be a table preset to store certain fields. If the fields in the data to be analyzed after deletion come from this table, it can be determined that these fields correspond to the lineage relationship, and then the hierarchical relationship between the fields corresponding to the lineage information is used as the data lineage corresponding to the top-level target field.
[0048] Illustrated with a specific example, assume the data to be analyzed is an SQL statement: insert into d(id,name) select A.cid,A.ename from (select cid,ename,row_number() over(partition by s1.cid order by s1.vno desc) AS rn from s1) A left join s2 on A.cid = s2.cid where A.rn = 1 and ename = 'N'.
[0049] First, convert the SQL statement into an abstract syntax tree and take the first target field at the first layer. In the example, it is insert into d(id); Query the next-layer fields. In the example, it is select A.cid,A.ename from A left join s2 on A.cid = s2.cid where A.rn = 1.
[0050] Fields that are irrelevant to the fields in the previous layer are pruned through SELECT. In the example, after pruning, it becomes "select A.cid from A left join s2 on A.cid = s2.cid where A.rn = 1". Then, all the fields in the current layer's query are traversed. In the example, they are A.cid, A.rn, A.ename, and s2.cid. Based on the above query, if the field is from a preset data table, it is the blood relationship information, which is s2.cid in the example.
[0051] Repeat the above steps until all the fields in the current layer are from the preset data table. In the example, repeat the operations of querying and pruning to obtain that the data blood relationship information comes from s2.cid, s1.cid, s1.ename, and s1.vno. Among them, s1.cid is the direct data blood relationship, and s2.cid, s1.ename, and s1.vno are the indirect data blood relationships.
[0052] Based on this process, the data blood relationship of all target fields can be gradually determined.
[0053] Based on the introduction of the above embodiments and specific examples, the above method for determining data blood relationship first determines the hierarchical relationship between fields based on the fields included in the data to be analyzed. Then, according to the top-level target fields among these fields, the irrelevant fields that have no association relationship with the top-level target fields are determined based on the hierarchical relationship. Furthermore, the fields that have a one-way association relationship with the irrelevant fields are deleted from the data to be analyzed. After the above processing, the hierarchical relationship reflected in the remaining data to be analyzed can be reflected as the data blood relationship corresponding to the top-level target fields. Through the above method, by deleting irrelevant fields, while gradually obtaining the data blood relationship, the amount of data that needs to be processed is also gradually reduced, thereby ensuring that the data blood relationship can be determined quickly and accurately. Furthermore, the data blood relationship is used to assist the processing of related services quickly and effectively, ensuring the processing effect of the services.
[0054] Based on Figure 1 the corresponding method for determining data blood relationship, an embodiment of this specification also proposes a business processing method based on data blood relationship. The business processing method based on data blood relationship is for a corresponding electronic device. The electronic device and the electronic device that executes the method for determining data blood relationship can be the same electronic device or different electronic devices. As Figure 2 shown, the business processing method based on data blood relationship may include the following specific implementation steps.
[0055] S210: Determine the business field data involved in the target business.
[0056] The target business is the business that needs to be processed currently. In the embodiments of this specification, relevant associated data needs to be retrieved during the processing of the target business. For example, if the target business is a related query business, after obtaining the user identifier, it is necessary to obtain the information of the user and the information of the account related to the user. Therefore, while querying the information directly associated with the user, it is also necessary to further search for the information related to the account based on the queried account. For this type of business, the data lineage can effectively assist the processing progress and effect of the business.
[0057] Business field data refers to the business that can be directly determined in the target business. For example, when receiving the target business, the operator or target object of the target business is indicated in the target business. Corresponding data can be searched according to the correspondence between the business field data and other data to assist in business processing.
[0058] S220: Obtain the data lineage corresponding to the business field data; the data lineage is obtained through the following method: obtain the data to be analyzed; the data to be analyzed contains multiple fields; determine the hierarchical relationship between the multiple fields; the hierarchical relationship includes the unidirectional association relationship between each field; select the business field data from the multiple fields based on the hierarchical relationship; determine the irrelevant fields that have no association relationship with the business field data; delete the fields that have a unidirectional association relationship with the irrelevant fields from the data to be analyzed; select the lower-level target fields corresponding to the business field data from the data to be analyzed after deletion based on the hierarchical relationship as the new top-level target fields, and repeat the operations of determining the irrelevant fields, deleting the fields that have a unidirectional association relationship with the irrelevant fields, and selecting the new top-level target fields until all fields are traversed; determine the hierarchical relationship in the data to be analyzed after traversing all fields and deleting as the data lineage corresponding to the business field data.
[0059] After obtaining the business field data, the data lineage corresponding to the business field data can be extracted. The data lineage is the relationship determined through the Figure 1 method. Through the data lineage, the field data directly or indirectly associated with the business field data can be quickly determined.
[0060] The data lineage can be the data lineage determined by taking the business field data as the top-level target field through the Figure 1 method before processing the business. For the specific process of obtaining the data lineage, reference can be made to the introduction in the corresponding embodiments of Figure 1 and will not be elaborated here.
[0061] S230: Extract associated field data corresponding to the service field data based on the data lineage relationship.
[0062] After obtaining the data lineage relationship, the data lineage relationship can be used to extract the associated field data corresponding to the service field data. The specific extraction method of the associated field data can be determined according to requirements. For example, all the data involved in the data lineage relationship can be extracted, or only the upper-layer data or lower-layer data of the associated field data can be extracted. The specific extraction method can be operated based on the actual application situation and will not be elaborated here.
[0063] S240: Process the target service using the associated field data.
[0064] After obtaining the associated field data, the associated field data can be used to process the target service. When processing the target service, the associated field data can be directly used as the processing result of the target service. For example, when the target service is a query service, the associated field data is directly used as the search result; or the associated field data can be processed and then used as the processing result of the target service. For example, when processing an account transaction service, relevant account data is operated according to the transaction content in the target service. In actual applications, the method of using the associated field data to process the target service can be set according to specific requirements.
[0065] Based on Figure 1 the corresponding data lineage relationship determination method, an embodiment of this specification introduces a data lineage relationship determination device. The data lineage relationship determination device can be set on the corresponding electronic device. As Figure 3 shown, the data lineage relationship determination device includes the following modules.
[0066] The to-be-analyzed data acquisition module 310 is used to acquire to-be-analyzed data; the to-be-analyzed data contains multiple fields;
[0067] The hierarchical relationship determination module 320 is used to determine the hierarchical relationship between the multiple fields; the hierarchical relationship includes the one-way association relationship between each field;
[0068] The top-level target field selection module 330 is used to select at least one top-level target field from the multiple fields based on the hierarchical relationship;
[0069] The irrelevant field determination module 340 is used to determine the irrelevant fields that have no association relationship with each top-level target field;
[0070] The field deletion module 350 is used to delete the fields that have a one-way association relationship with the irrelevant fields from the to-be-analyzed data;
[0071] The traversal module 360 is configured to select, based on the hierarchical relationship, the lower-level target fields corresponding to the top-level target field from the to-be-analyzed data after deletion as the new top-level target fields, and repeatedly perform the operations of determining irrelevant fields, deleting the fields having a one-way association relationship with the irrelevant fields, and selecting new top-level target fields until all fields are traversed.
[0072] The data lineage determination module 370 is configured to determine the hierarchical relationship in the to-be-analyzed data obtained after deleting all fields as the data lineage corresponding to the top-level target field.
[0073] In some embodiments, the to-be-analyzed data includes SQL statements; the hierarchical relationship determination module includes: an abstract syntax tree conversion unit configured to convert the to-be-analyzed data into an abstract syntax tree; the abstract syntax tree is used to represent the hierarchical relationship between fields.
[0074] Based on the above embodiments, the field deletion module includes: a field deletion unit configured to delete the fields associated with the irrelevant fields based on the hierarchical relationship in the abstract syntax tree.
[0075] In some embodiments, the irrelevant field determination module includes: an irrelevant field screening unit configured to screen out the irrelevant fields that have no association relationship with the top-level target field from the lower-level target fields corresponding to each top-level target field; correspondingly, the field deletion module includes: a field pruning unit configured to prune the fields having a one-way association relationship with the irrelevant fields from the to-be-analyzed data; a traversal unit configured to use the remaining lower-level target fields in the to-be-analyzed data as the new top-level target fields, and repeatedly perform the operations of screening irrelevant fields and pruning fields until all levels of fields are traversed.
[0076] In some embodiments, the irrelevant field determination module includes: an association relationship determination unit configured to determine the association relationship between each lower-level target field and the top-level target field based on a preset inference logic; the preset inference logic is used to determine the association relationship based on the semantic conflict relationship.
[0077] In some embodiments, the data lineage determination module includes: a blood relationship information determination unit configured to determine that the fields in the to-be-analyzed data after deletion correspond to blood relationship information when the fields in the to-be-analyzed data after deletion are from a preset data table; a data lineage determination unit configured to use the hierarchical relationship between the fields corresponding to the blood relationship information as the data lineage corresponding to the top-level target field.
[0078] Based on Figure 2The corresponding business processing method based on data lineage introduces a business processing device based on data lineage according to an embodiment of this specification. The business processing device based on data lineage can be set on a corresponding electronic device. As Figure 4 shown, the business processing device based on data lineage includes the following modules.
[0079] A business field data determination module 410, configured to determine business field data involved in a target business;
[0080] A data lineage acquisition module 420, configured to acquire the data lineage corresponding to the business field data; the data lineage is acquired through the following method: acquiring data to be analyzed; the data to be analyzed contains multiple fields; determining the hierarchical relationship between the multiple fields; the hierarchical relationship includes one-way association relationships between each field; selecting the business field data from the multiple fields based on the hierarchical relationship; determining irrelevant fields that have no association relationship with the business field data; deleting fields that have a one-way association relationship with the irrelevant fields from the data to be analyzed; selecting the lower-layer target fields corresponding to the business field data from the data to be analyzed after deletion based on the hierarchical relationship as new top-layer target fields, and repeatedly performing the operations of determining irrelevant fields, deleting fields that have a one-way association relationship with the irrelevant fields, and selecting new top-layer target fields until all fields are traversed; using the hierarchical relationship in the data to be analyzed after traversing all fields and deleting as the data lineage corresponding to the business field data;
[0081] An associated field extraction module 430, configured to extract associated field data corresponding to the business field data based on the data lineage;
[0082] A target business processing module 440, configured to process the target business by using the associated field data.
[0083] Based on Figure 1 the corresponding data lineage determination method and business processing method, an embodiment of this specification provides an electronic device. The electronic device can be equivalent to the above data lineage determination method and business processing device. The electronic device can include a memory and a processor.
[0084] In this embodiment, the memory can be implemented in any suitable manner. For example, the memory can be a read-only memory, a mechanical hard disk, a solid-state drive, or a USB flash drive, etc. The memory can be used to store computer programs / instructions.
[0085] In this embodiment, the processor may be implemented in any suitable manner. For example, the processor may take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an Application Specific Integrated Circuit (ASIC), a programmable logic controller, and an embedded microcontroller, etc. The processor may execute the computer program instructions to implement as Figure 1 the corresponding data lineage determination method and / or Figure 2 the corresponding service processing method based on data lineage.
[0086] This embodiment of the specification provides a computer-readable storage medium, on which computer programs / instructions are stored. The computer-readable storage medium may be read by the processor based on the internal bus of the device, and then the program instructions in the computer-readable storage medium are implemented through the processor.
[0087] In this embodiment, the computer-readable storage medium may be implemented in any suitable manner. The computer-readable storage medium includes but is not limited to Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), Memory Card, etc. The computer storage medium stores computer program instructions. When the computer program instructions are executed, the corresponding data lineage determination method and / or Figure 1 the corresponding program instructions or modules of the service processing method based on data lineage in this specification are implemented. Figure 2
[0088] This embodiment of the specification also provides a computer program product, including computer programs / instructions. The computer program product may be a program written in a corresponding computer programming language, stored in a corresponding storage device in a program manner, and may be transmitted through a computer network. The computer program product may be executed by the processor. In this embodiment of the specification, when the computer program product is executed, the corresponding data lineage determination method and / or Figure 1 the corresponding program instructions or modules of the service processing method based on data lineage are implemented. Figure 2
[0089] It should be noted that the above methods for determining data lineage, business processing methods, devices, and equipment can be applied to the field of big data technology, and can also be applied to other technical fields other than the big data technology field, and there is no limitation in this regard.
[0090] In addition, all operations such as obtaining, processing, storing, and forwarding of all data including target object data in the above methods for determining data lineage, business processing methods, devices, and equipment comply with the relevant provisions of national laws and regulations.
[0091] Although the process flows described above include multiple operations that occur in a specific order, it should be clearly understood that these processes can include more or fewer operations, and these operations can be executed sequentially or in parallel (for example, using a parallel processor or a multi-threaded environment).
[0092] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0093] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0095] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0096] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0097] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic tape storage, magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0098] Those skilled in the art will appreciate that the embodiments of this specification may be provided as a method, system, or computer program product. Accordingly, the embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0099] The embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The embodiments of this specification may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.
[0100] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the corresponding description in the method embodiment. In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of this specification. In this specification, the schematic expression of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0101] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for determining data lineage, characterized in that Including: Obtain the data to be analyzed; The data to be analyzed contains multiple fields; Determine the hierarchical relationship between the multiple fields; The hierarchical relationship includes one-way association relationships between fields; Based on the hierarchical relationship, select at least one top-level target field from the multiple fields; Determine the irrelevant fields that have no association relationship with each top-level target field; Delete the fields in the data to be analyzed that have a one-way association relationship with the irrelevant fields; Based on the hierarchical relationship, select the lower-level target fields corresponding to the top-level target fields from the data to be analyzed after deletion as new top-level target fields, and repeat the operations of determining irrelevant fields, deleting the fields that have a one-way association relationship with the irrelevant fields, and selecting new top-level target fields until all fields are traversed; Determine the hierarchical relationship in the data to be analyzed obtained after traversing all fields and deleting as the data lineage corresponding to the top-level target field; Among them, the data to be analyzed includes SQL statements; determining the hierarchical relationship between the multiple fields includes: converting the data to be analyzed into an abstract syntax tree; the abstract syntax tree is used to represent the hierarchical relationship between fields; Determining the hierarchical relationship in the data to be analyzed obtained after traversing all fields and deleting as the data lineage corresponding to the top-level target field includes: when the fields in the data to be analyzed after deletion are from a preset data table, determining that the fields in the data to be analyzed after deletion correspond to bloodline information; taking the hierarchical relationship between the fields corresponding to the bloodline information as the data lineage corresponding to the top-level target field.
2. The method according to claim 1, wherein Deleting the fields in the data to be analyzed that have a one-way association relationship with the irrelevant fields includes: Deleting the fields associated with the irrelevant fields based on the hierarchical relationship in the abstract syntax tree.
3. The method according to claim 1, characterized in that Determining the irrelevant fields that have no association relationship with each top-level target field includes: Respectively screening out the irrelevant fields that have no association relationship with the top-level target field from the lower-level target fields corresponding to each top-level target field; Correspondingly, deleting the fields in the data to be analyzed that have a one-way association relationship with the irrelevant fields includes: Pruning the fields in the data to be analyzed that have a one-way association relationship with the irrelevant fields; among them, including: sequentially deleting the fields corresponding to the irrelevant fields at each level based on the one-way association relationship between fields.
4. The method according to claim 1, characterized in that, Determining the irrelevant fields that have no association relationship with each top-level target field includes: Determining the association relationship between each lower-level target field and the top-level target field based on a preset inference logic; the preset inference logic is used to determine the association relationship based on the semantic conflict relationship.
5. A business processing method based on data lineage, characterized in that, Including: Determine the business field data involved in the target business; Obtain the data lineage corresponding to the business field data; The data lineage is obtained in the following way: obtain the data to be analyzed; the data to be analyzed contains multiple fields; Determine the hierarchical relationship between the multiple fields; The hierarchical relationship includes one-way association relationships between fields; Based on the hierarchical relationship, select the business field data from the multiple fields; Determine the irrelevant fields that have no associated relationship with the business field data; Delete the fields in the data to be analyzed that have a one-way association with the irrelevant fields from the data to be analyzed; Based on the hierarchical relationship, select the lower-layer target fields corresponding to the business field data from the data to be analyzed after deletion as the new top-layer target fields, and repeat the operations of determining the irrelevant fields, deleting the fields that have a one-way association with the irrelevant fields, and selecting the new top-layer target fields until all fields are traversed; Determine the hierarchical relationship in the data to be analyzed obtained after traversing all fields and deleting as the data lineage corresponding to the business field data; Extract the associated field data corresponding to the business field data based on the data lineage; Process the target business using the associated field data; Wherein, the data to be analyzed includes SQL statements; determining the hierarchical relationship between the multiple fields includes: converting the data to be analyzed into an abstract syntax tree; the abstract syntax tree is used to represent the hierarchical relationship between the fields; Determining the hierarchical relationship in the data to be analyzed obtained after traversing all fields and deleting as the data lineage corresponding to the business field data includes: when the fields in the data to be analyzed after deletion are from a preset data table, determining that the fields in the data to be analyzed after deletion correspond to blood relationship information; taking the hierarchical relationship between the fields corresponding to the blood relationship information as the data lineage corresponding to the top-layer target field.
6. A data lineage determination device, characterized in that, Including: A data acquisition module to be analyzed, configured to acquire data to be analyzed; The data to be analyzed contains multiple fields; A hierarchical relationship determination module, configured to determine the hierarchical relationship between the multiple fields; The hierarchical relationship includes the one-way association relationship between each field; A top-layer target field selection module, configured to select at least one top-layer target field from the multiple fields based on the hierarchical relationship; An irrelevant field determination module, configured to determine the irrelevant fields that have no associated relationship with each top-layer target field; A field deletion module, configured to delete the fields in the data to be analyzed that have a one-way association with the irrelevant fields; A traversal module, configured to select the lower-layer target fields corresponding to the top-layer target fields from the data to be analyzed after deletion as the new top-layer target fields based on the hierarchical relationship, and repeat the operations of determining the irrelevant fields, deleting the fields that have a one-way association with the irrelevant fields, and selecting the new top-layer target fields until all fields are traversed; A data lineage determination module, configured to determine the hierarchical relationship in the data to be analyzed obtained after traversing all fields and deleting as the data lineage corresponding to the top-layer target field; Wherein, the data to be analyzed includes SQL statements; determining the hierarchical relationship between the multiple fields includes: converting the data to be analyzed into an abstract syntax tree; the abstract syntax tree is used to represent the hierarchical relationship between the fields; Determining the hierarchical relationship in the data to be analyzed obtained after traversing all fields and deleting it as the data lineage corresponding to the top-level target field includes: when the fields in the data to be analyzed after deletion are from a preset data table, determining that the fields in the data to be analyzed after deletion correspond to lineage information; and taking the hierarchical relationship between the fields corresponding to the lineage information as the data lineage corresponding to the top-level target field.
7. A business processing device based on data lineage, characterized in that, Including: A business field data determination module, configured to determine business field data involved in a target business; A data lineage acquisition module, configured to acquire the data lineage corresponding to the business field data; the data lineage is acquired by the following method: acquiring data to be analyzed; the data to be analyzed contains multiple fields; Determining the hierarchical relationship between the multiple fields; The hierarchical relationship includes one-way association relationships between each field; Selecting the business field data from the multiple fields based on the hierarchical relationship; Determining irrelevant fields that have no association relationship with the business field data; Deleting from the data to be analyzed fields that have a one-way association relationship with the irrelevant fields; Based on the hierarchical relationship, selecting a lower-level target field corresponding to the business field data from the data to be analyzed after deletion as a new top-level target field, and repeating the operations of determining irrelevant fields, deleting fields that have a one-way association relationship with the irrelevant fields, and selecting a new top-level target field until all fields are traversed; taking the hierarchical relationship in the data to be analyzed obtained after traversing all fields and deleting it as the data lineage corresponding to the business field data; An associated field extraction module, configured to extract associated field data corresponding to the business field data based on the data lineage; A target business processing module, configured to process the target business by using the associated field data; Wherein, the data to be analyzed includes SQL statements; determining the hierarchical relationship between the multiple fields includes: converting the data to be analyzed into an abstract syntax tree; the abstract syntax tree is used to represent the hierarchical relationship between fields; Determining the hierarchical relationship in the data to be analyzed obtained after traversing all fields and deleting it as the data lineage corresponding to the business field data includes: when the fields in the data to be analyzed after deletion are from a preset data table, determining that the fields in the data to be analyzed after deletion correspond to lineage information; and taking the hierarchical relationship between the fields corresponding to the lineage information as the data lineage corresponding to the top-level target field.
8. An electronic device, comprising a memory and a processor; characterized in that, The memory is used to store computer programs / instructions; the processor is configured to execute the computer programs / instructions to implement the steps of the method according to any one of claims 1-7.
9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, The computer programs / instructions, when executed by the processor, implement the steps of the method according to any one of claims 1-5.
10. A computer program product, comprising a computer program / instructions, characterized in that, The computer programs / instructions, when executed by the processor, implement the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Data management full link-based field-level blood relationship analysis method
CN114116856A
Field-level data consanguinity extraction method and device, equipment and storage medium
CN114817298A