Method and device for data blood relationship analysis, storage medium and electronic equipment
By disassembling upstream sources of structured query processing information and building a data blood relationship tree, combined with large model analysis, the index accuracy and keen perception problems caused by complex data blood relationship are solved, and the linearization and standardization of data blood relationship is achieved, and the comprehensibility of risk interpretation is improved.
Patent Information
- Application Number
- CN202510089019.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
In complex business risk control systems, the data relationship between multiple indicators is complex, making it difficult to maintain the accuracy and keen perception of indicators.
By disassembling the upstream source of the structured query processing information of the initial target field, a data blood relationship tree is constructed, and prompt information is generated through large-scale model analysis to realize data blood relationship analysis.
Linearize and standardize the complex blood relationship of data, improve the humanization level of risk interpretation, and transform the complex data processing process into concise and easy-to-understand language, lower the technical threshold, and improve the accessibility of data analysis.
Smart Images

Figure CN119988460A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer technology, and in particular to a method, device, storage medium and electronic device for data lineage analysis. Background Art
[0002] In the risk control systems of various businesses, the number of full risk monitoring indicators may be very large. For example, in some business scenarios, there may be more than 600 indicators that need to be tested for risk control, and each indicator may be processed by multiple tables. In such a large and complex data blood relationship network, how to maintain indicator accuracy and keen perception has become a challenge. Summary of the invention
[0003] The purpose of the embodiments of this specification is to provide a method, device, storage medium and electronic device for data lineage analysis.
[0004] The embodiment of this specification provides a method for data lineage analysis, which can obtain structured query processing information corresponding to each layer associated with the first target field by disassembling the upstream source of the initial first target field, thereby constructing a data lineage relationship tree corresponding to the first target field, and generating prompt information by parsing the data lineage relationship tree, thereby obtaining the data lineage analysis result corresponding to the first target field through the lineage analysis large model, which can realize the linearization and standardization of complex data lineage relationships, as well as the humanization of risk interpretation. The method includes:
[0005] For an initial first target field, obtaining structured query processing information corresponding to a first target table to which the first target field belongs, wherein the structured query processing information is used to characterize a data processing process corresponding to the first target table;
[0006] According to the structured query processing information, determine whether the termination condition is currently met; if not, perform upstream source disassembly on the structured query processing information to obtain one or more second target fields located upstream of the first target field, obtain the latest structured query processing information corresponding to the second target table to which the second target field belongs, and so on; if so, construct a data lineage relationship tree based on the obtained multiple target fields and the target table to which each target field belongs, wherein each node in the data lineage relationship tree corresponds to a target field, and the direction of the edge between two nodes is used to characterize the data flow direction between the target fields corresponding to the two nodes respectively;
[0007] The data lineage relationship tree is interpreted to generate corresponding prompt information, and the prompt information is input into a trained data lineage analysis model to obtain a data lineage analysis result output by the data lineage analysis model.
[0008] In some embodiments, judging whether a termination condition is currently satisfied according to the structured query processing information includes:
[0009] Simplifying the structured query processing information according to the first target field to obtain simplified structured query processing information;
[0010] According to the simplified structured query processing information, determining whether a termination condition is currently met;
[0011] Wherein, the upstream source disassembly of the structured query processing information includes:
[0012] The streamlined structured query processing information is disassembled from upstream sources.
[0013] In some embodiments, the simplifying the structured query processing information according to the first target field to obtain the simplified structured query processing information includes:
[0014] Determining, in the structured query processing information, non-associated text information that satisfies a preset association relationship with the first target field;
[0015] The structured query processing information is simplified by deleting the non-related text information to obtain simplified structured query processing information.
[0016] In some embodiments, the method further comprises:
[0017] Determining non-critical text information that meets a preset criticality level in the structured query processing information;
[0018] The step of simplifying the structured query processing information by deleting the non-related text information to obtain the simplified structured query processing information includes:
[0019] The structured query processing information is simplified by deleting the non-related text information and the non-key text information to obtain simplified structured query processing information.
[0020] In some embodiments, the termination condition includes at least one of the following:
[0021] The current level has reached the upper limit of the upstream level;
[0022] The first target field has no upstream dependency.
[0023] In some embodiments, the edge between the two nodes corresponds to structured processing aperture information between the target fields respectively corresponding to the two nodes, wherein the structured processing aperture information is used to characterize the data processing process between the target fields respectively corresponding to the two nodes;
[0024] The step of constructing a data lineage tree based on at least one target table involved in each acquired data level and at least one target field involved in each target table includes:
[0025] A data lineage tree is constructed based on the structured processing caliber information corresponding to at least one target table involved in each data level, at least one target field involved in each target table, and two target fields in an upper and lower hierarchical relationship.
[0026] In some embodiments, the method further comprises:
[0027] According to the structured query processing information, structured processing caliber information between the first target field and the second target field is determined.
[0028] In some embodiments, the method further comprises:
[0029] pruning nodes of the data kinship tree according to preset logic reinforcement rules;
[0030] The generating of corresponding prompt information by interpreting the data blood relationship tree includes:
[0031] By interpreting the pruned data kinship tree, corresponding prompt information is generated.
[0032] In some embodiments, the logic enforcement rule is used to indicate that structured processing aperture information of existence operator keywords is retained.
[0033] The embodiment of this specification also provides a device for data lineage analysis, including:
[0034] An obtaining module, configured to obtain, for an initial first target field, structured query processing information corresponding to a first target table to which the first target field belongs, wherein the structured query processing information is used to characterize a data processing process corresponding to the first target table;
[0035] A construction module is used to determine whether the termination condition is currently met based on the structured query processing information; if not, to perform upstream source disassembly on the structured query processing information, to obtain one or more second target fields located upstream of the first target field, to obtain the latest structured query processing information corresponding to the second target table to which the second target field belongs, and so on; if yes, to construct a data lineage relationship tree based on the obtained multiple target fields and the target table to which each target field belongs, wherein each node in the data lineage relationship tree corresponds to a target field, and the direction of the edge between two nodes is used to characterize the data flow direction between the target fields corresponding to the two nodes respectively;
[0036] The output module is used to generate corresponding prompt information by interpreting the data lineage relationship tree, input the prompt information into the trained data lineage analysis model, and obtain the data lineage analysis result output by the data lineage analysis model.
[0037] The embodiments of the present specification also provide a storage medium, wherein the storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the above method.
[0038] An embodiment of the present specification also provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the above method.
[0039] The embodiments of the present specification also provide a computer program product having at least one instruction stored thereon, wherein the at least one instruction implements the steps of the above method when executed by a processor.
[0040] In an embodiment of the present specification, by first obtaining structured query processing information corresponding to the first target table to which the initial first target field belongs, and then judging whether the termination condition is currently satisfied based on the structured query processing information, if not, the structured query processing information is disassembled from the upstream source, and if so, a data lineage tree is constructed based on the obtained multiple target fields and the target table to which each target field belongs, thereby obtaining a data lineage tree corresponding to the first target field, and then interpreting the data lineage tree to generate corresponding prompt information, so as to input the prompt information into the trained data lineage analysis model to obtain the data lineage analysis result, thereby linearizing and standardizing the complex data lineage relationship and humanizing the risk interpretation, so as to provide a wider range of users with services that are easy to access, understand and operate by translating the profound data processing process into simple and easy-to-understand language, thereby lowering the technical threshold and making data analysis within reach, based on which, the value mining of data can be deepened, and each indicator (that is, the target field) can be given a clear purpose and powerful analysis capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A flowchart of a method for data lineage analysis provided in an embodiment of this specification.
[0042] Figure 2 A flowchart of a method for data lineage analysis provided as an example in an embodiment of this specification.
[0043] Figure 3 A schematic diagram of the structure of a device for data lineage analysis provided in an embodiment of this specification.
[0044] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0046] See also Figure 1, is a flow chart of a method for data lineage analysis provided in an embodiment of this specification. In an embodiment of this specification, the method for data lineage analysis is applied to a device for data lineage analysis (hereinafter referred to as "data lineage analysis device") or an electronic device equipped with a data lineage analysis device. Figure 1 The process shown is described in detail, and the method for data lineage analysis may specifically include the following steps:
[0047] S102: For an initial first target field, obtain structured query processing information corresponding to a first target table to which the first target field belongs, wherein the structured query processing information is used to characterize a data processing process corresponding to the first target table.
[0048] In some embodiments, the first target field is used to characterize the indicator that needs to be risk controlled. The first target field can be expressed as the core feature of the indicator, such as the unique identifier of the field; in some embodiments, the first target field is expressed as "data table_field", for example, the first target field is expressed as "T1_rate", which means the field rate in data table T1. In some embodiments, the initial first target field can be determined based on the risk control requirements of the business scenario.
[0049] In some embodiments, the level where the initial first target field is located is regarded as the starting level, and the first target table to which the first target field belongs can be determined based on the representation of the first target field (such as the aforementioned "data table_field"), or it may be determined by querying the table in the starting level, or it may be determined in combination with the interaction with the user, and this specification does not limit this. In some embodiments, the structured query processing information is also the table processing sql (Structured Query Language) involved in the current query level, and the table processing sql includes multiple fields and processing information of each field. As an example, the initial first target field is represented as "T1_rate", that is, the characterization indicator is the field rate in the starting layer data table T1. For the first target field, the table processing sql corresponding to the starting data table T1 to which it belongs is obtained, and the sql can characterize the data processing process corresponding to T1.
[0050] S104, judging whether the termination condition is currently satisfied according to the structured query processing information; if not, performing upstream source disassembly on the structured query processing information to obtain one or more second target fields located upstream of the first target field, and obtaining the latest structured query processing information corresponding to the second target table to which the second target field belongs, and so on; if yes, constructing a data lineage tree based on the obtained multiple target fields and the target table to which each target field belongs, wherein each node in the data lineage tree corresponds to a target field, and the direction of the edge between two nodes is used to characterize the data flow direction between the target fields corresponding to the two nodes respectively.
[0051] In some embodiments, the termination condition includes any condition for indicating to stop the query (i.e., stop the upstream source disassembly operation). In some embodiments, the termination condition includes at least one of the following: the current level has reached the upper limit of the upstream level; the first target field has no upstream dependency. For example, the upper limit of the upstream level is 10, and if the current level has reached the upper limit of the upstream level of the first target field, the termination condition is met.
[0052] In some embodiments, the upstream target field obtained by disassembling the upstream source may be one or more. For example, the target field at the current level may be obtained by performing corresponding operations on multiple upstream target fields. It should be noted that if there are multiple upstream target fields, the operation of determining whether the termination condition is currently met will be performed only after the structured query processing information corresponding to the multiple target fields is obtained. It should be noted that in the embodiments of this specification, by disassembling the upstream source layer by layer, the flow of data in each link can be systematically identified and recorded, as well as how the data in each data flow is transferred from one link to another, and these specific processing logics can be clearly linked to the relevant data elements (that is, the target fields).
[0053] In some embodiments, if the termination condition is currently met, a data lineage tree is constructed based on the multiple target fields of multiple levels that have been obtained and the target table to which each target field belongs. By disassembling the upstream source of the first target field layer by layer to sort out the data lineage, a data structure including nodes and edges can be constructed, and with the help of the advanced capabilities of the graph database, a comprehensive data lineage tree is successfully created; the data lineage tree not only helps to track the entire process of data flowing from its source to the end point, but also records in detail the conversion, aggregation and use of the data. In some embodiments, the data lineage tree uses graphical means to intuitively display data elements and their connections with each other, helping users to understand at a glance the path of data flow, the processing steps of the application, and the final use of the data; the data lineage tree can be presented in the form of a directed graph. In some embodiments, the data lineage tree includes nodes and edges, each edge carries edge information, wherein each node represents a target field, and each level may have one or more nodes (the starting layer is a node corresponding to the first target field); the direction of the edge is used to characterize the data flow between the target fields corresponding to the two nodes connected to it, for example, the data flows from the node corresponding to the first target field in the starting layer to the node corresponding to the second target field in the upstream first layer; the edge information carried by each edge is used to indicate the processing caliber information (that is, the processing logic) corresponding to the edge, for example, the edge information between the node corresponding to the first target field and the node corresponding to the second target field is used to indicate how to process the first target field to obtain the second target field. The embodiments of this specification make the data lineage tree a powerful tool through a rigorous construction process for the data lineage tree, making the flow and transformation process of data transparent and easy to understand, greatly enhancing data governance capabilities and data utilization efficiency.
[0054] It should be noted that, for the target fields of each layer, the implementation method of obtaining the structured query processing information corresponding to the target table to which the target field belongs is the same or similar. For example, for the initial first target field, the table processing SQL corresponding to the first target table to which the first target field belongs is obtained. If it is then judged that the termination condition is not met, the upstream source is disassembled to obtain one or more second target fields of the upstream first layer, and then the table processing SQL corresponding to the target table to which each second target field belongs is obtained in the same manner as for the starting layer, and so on to obtain the table processing SQL corresponding to the target table to which the target fields in each upstream layer belong.
[0055] S106, generating corresponding prompt information by interpreting the data lineage relationship tree, inputting the prompt information into the trained data lineage analysis model, and obtaining the data lineage analysis result output by the data lineage analysis model.
[0056] In some embodiments, by parsing the data lineage tree, a prompt (i.e., prompt) with scenario description, goal orientation, and compliance guidelines can be constructed to achieve a better interactive experience and accuracy. In some embodiments, the prompt information includes an explanation of each layer of the process in the data lineage tree, such as the data tables and fields involved in the layer and the corresponding code. As an example format, the prompt information includes "context" (for example: you are a data quality assurance robot, helping humans to interpret data lineage, and the following is the upstream processing process of field A in table T1), "goal" (for example: please combine the sql of multiple processes to restore the full-link lineage restoration analysis generated by field A in table T1. It is required to sort out the processing process in each process, and the end point describes the logical processing with conditions in sql), "format" (for example: [lineage interpretation]: interpret the lineage link comprehensively and in detail to natural language; [risk assessment]: systematically and in detail explain the risk assessment, focusing on describing the business impact and data impact), "audience" (for example: business personnel who do not have a technical foundation), "starting layer process", "upstream first layer process", "upstream second layer process", etc. It should be noted that the above examples are only examples, not limitations of this application. In actual application scenarios, the format of prompt information can be defined in combination with actual needs. In the embodiments of this specification, by cleverly constructing an efficient prompt with rich scenario descriptions, clear goal orientation, and strict compliance with relevant guidelines, and with the help of the powerful capabilities of the large model and deep knowledge base support, it is possible to accurately generate interpretation texts of blood relationships, which not only greatly improves the humanization level of risk interpretation, but also enables complex information to be presented to users in a more understandable and intimate way, effectively promoting the quality and efficiency of decision-making.
[0057] In the embodiments of the present specification, a strategy of hierarchical tracking of structured query processing information is adopted to establish a data lineage link. Specifically, for the initial first target field, the structured query processing information corresponding to the first target table to which the first target field belongs is obtained. Then, according to the structured query processing information, it is determined whether the termination condition is currently satisfied. If not, the structured query processing information is disassembled from the upstream source. If so, a data lineage relationship tree is constructed based on the obtained multiple target fields and the target table to which each target field belongs. Thus, the data lineage relationship tree corresponding to the first target field can be obtained. Then, the data lineage relationship tree can be interpreted to generate corresponding prompt information, and the prompt information can be input into the trained data lineage analysis model to obtain the data lineage analysis result. Thus, the complex data lineage relationship can be linearized, standardized, and the risk interpretation can be humanized. Therefore, by translating the profound data processing process into a simple and easy-to-understand language, it is possible to provide a wider range of users with services that are easy to access, understand and operate, reduce the technical threshold, and make data analysis within reach. Based on this, the value mining of data can be deepened, and each indicator (that is, the target field) can be given a clear purpose and powerful analysis capabilities.
[0058] In some embodiments, judging whether the termination condition is currently met according to the structured query processing information includes: simplifying the structured query processing information according to the first target field to obtain simplified structured query processing information; judging whether the termination condition is currently met according to the simplified structured query processing information; wherein, performing upstream source disassembly on the structured query processing information includes: performing upstream source disassembly on the simplified structured query processing information. In some embodiments, the same simplification strategy is implemented for each layer, for example, a simplification strategy of deleting non-associated text information is adopted for each layer of target fields to obtain simplified sql. In some embodiments, the simplified structured query processing information (i.e., simplified sql) includes all processing information or key processing information associated with the target field. As an example, for the initial first target field T1_rate, after obtaining the table processing sql corresponding to the first target table T1 to which the first target field belongs, the sql is simplified to delete the part irrelevant to the rate field, and only retain the part related to the processing process of the rate field, thereby obtaining the simplified sql, and then performing upstream source disassembly on the simplified sql.
[0059] In some embodiments, by calling the capabilities of FastSQL (a project used to accelerate SQL queries), lexical and grammatical analysis of SQL is performed, an abstract syntax tree is extracted, and then a streamlining operation is performed, such as stripping off non-critical elements and retaining the logical parts closely related to the target field; then the SQL is rebuilt to completely map the context of the key data, and the rebuilt SQL is the streamlined SQL mentioned above, thereby realizing a data streamlining method from analysis to disassembly and then to re-assembly. Through the streamlining operation, the interpretation of data lineage can be more accurate, detailed and efficient.
[0060] In some embodiments, the streamlining of the structured query processing information according to the first target field to obtain the streamlined structured query processing information includes: determining in the structured query processing information non-associated text information that satisfies a preset association relationship with the first target field; and streamlining the structured query processing information by deleting the non-associated text information to obtain the streamlined structured query processing information. In some embodiments, the preset association relationship is used to indicate whether there is an association or an association condition that needs to be met. For example, if the preset association relationship is that the association degree with the target field is less than a preset threshold, then when performing the streamlining operation, the data processing process with the association degree with the target field less than the preset threshold is deleted as non-associated text information, and only the data processing process with the association degree with the target field greater than or equal to the preset threshold is retained.
[0061] In some embodiments, the method further includes: determining non-key text information that meets a preset criticality in the structured query processing information; wherein, the step of simplifying the structured query processing information by deleting the non-related text information to obtain the simplified structured query processing information includes: simplifying the structured query processing information by deleting the non-related text information and the non-key text information to obtain the simplified structured query processing information. In some embodiments, the processing of a field whose criticality is lower than a predetermined threshold may be determined as non-key text information. In some embodiments, the criticality may be preset based on a business scenario. In some embodiments, the inclusion of preset operation logic or the non-inclusion of preset operation logic may be used as a condition for determining whether the preset criticality is met. For example, if a processing process in sql does not include summation logic, it is determined that the processing process meets the preset criticality, that is, the code corresponding to the processing process is determined as non-key text information that meets the preset criticality. In some embodiments, some preset general processing processes may be regarded as non-key text information that meets the preset criticality. Therefore, in the process of performing the streamlining operation, while deleting non-related text information that satisfies the preset association relationship, non-critical text information that meets the preset criticality can also be deleted (for example, while deleting the processing process of other fields that are not related to the target field, deleting the processing process that is related to the target field but does not contain the summation logic), so that the processing logic that the business scenario is concerned about can be obtained more accurately.
[0062] In some embodiments, the edge between the two nodes corresponds to the structured processing aperture information between the target fields corresponding to the two nodes, wherein the structured processing aperture information is used to characterize the data processing process between the target fields corresponding to the two nodes; wherein the data lineage tree is constructed based on the at least one target table involved in each data level and the at least one target field involved in each target table, including: constructing a data lineage tree based on the structured processing aperture information corresponding to the at least one target table involved in each data level, the at least one target field involved in each target table, and the two target fields with upper and lower hierarchical relationships. In some embodiments, the constructed data lineage tree includes multiple levels, each level includes at least one node, each node is represented in the "data table_field" format, the edge between the nodes is a directed edge, used to indicate the direction of data flow, and the edge information carried by the edge is the structured processing aperture information, used to characterize how to process from one target field to obtain another target field.
[0063] In some embodiments, the method further includes: determining the structured processing aperture information between the first target field and the second target field according to the structured query processing information. In some embodiments, the processing process related to the first target field is queried from the structured query structure information, and the structured processing aperture information between the first target field and the second target field is determined based on the query result. In some embodiments, the structured query processing information is first simplified to obtain a simplified sql, and then the structured processing aperture information between the first target field and the second target field is determined from the simplified sql.
[0064] In some embodiments, the method further includes: pruning nodes of the data lineage tree according to preset logic reinforcement rules; wherein, generating corresponding prompt information by interpreting the data lineage tree includes: generating corresponding prompt information by interpreting the pruned data lineage tree. In some embodiments, the preset logic reinforcement rules are used to indicate which processing logics to retain or which processing logics to delete, or indicate which nodes to retain or which nodes to delete. In some embodiments, for each business line, logic reinforcement rules can be preset according to the business scenarios corresponding to each business line, that is, different business lines can customize different logic reinforcement rules, thereby supporting personalized logic reinforcement rules for different business scenarios, so as to obtain more accurate data lineage interpretation results oriented to business needs. In some embodiments, only leaf nodes in the data lineage tree need to be pruned based on logic reinforcement rules, and non-leaf nodes do not need to be pruned. In some embodiments, if all structured processing caliber information pointing to the node is pruned, the node is deleted.
[0065] In some embodiments, the logic reinforcement rule is used to indicate that the structured processing aperture information with the operator keyword is retained. As an example, the preset logic reinforcement rule indicates that the structured processing aperture information with the operator keyword "sum" is retained. Then, for the constructed data lineage tree, pruning can be performed based on the logic reinforcement rule to delete the structured processing aperture information that does not contain "sum". When deleting the structured processing aperture information, the leaf node pointed to by the data flow direction can be deleted (in actual business scenarios, whether to delete the leaf node pointed to by the data flow direction while deleting the structured processing aperture information can be set based on the logic reinforcement rule. For example, after the deletion operation is performed on the data flow direction, if there is no other operation logic, the leaf node pointed to by the data flow direction can be deleted. If there is still sum logic, the leaf node pointed to by the data flow direction is retained).
[0066] It should be noted that through the above-mentioned logic reinforcement rules, the data lineage tree can be further optimized and pruned, and "minor" processing caliber information or leaf nodes can be deleted. Such a strategy ensures the integrity of the key logic in the lineage structure and ensures that the information provided to the data lineage analysis model is not only effective but also accurate, thereby improving the quality and efficiency of the entire data processing process.
[0067] Figure 2 This is a flow chart of a method for data lineage analysis provided as an example in the embodiments of this specification. The method can be divided into the following four stages:
[0068] 1) Sort out the lineage. This stage mainly includes the following steps: 1.1) Query the table processing sql involved in the hierarchy. For example, after determining the initial first target field, query the table processing sql involved in the starting layer for the first target field (that is, the structured query processing information corresponding to the first target table to which the first target field belongs). 1.2) Streamline the sql according to the field. For example, based on the first target field of the current query, streamline the sql queried in step 1.1) and delete the code that is not related to the first target field. 1.3) Temporarily store the streamlined sql. 1.4) Determine whether the termination condition is reached. If the termination condition is not reached, execute the following step 1.5). If the termination condition is reached, execute the following step 2.1. 1.5) Disassemble the upstream source. For example, after executing the above steps 1.1) to 1.4) for the initial first target field, if it is determined that the termination condition is not reached, the operation of disassembling the upstream source is performed based on the streamlined SQL to obtain one or more target fields of the first upstream layer, and then repeat the above steps 1.1) to 1.4) for the one or more target fields, and so on, until the termination condition is reached.
[0069] 2) Logical reinforcement. This stage mainly includes the following steps: 2.1) Construct a lineage tree (also called "data lineage tree" in this context) hierarchically. The nodes in the lineage tree are used to express lineage elements (that is, target fields), and the lines in the lineage tree are used to express the processing between lineage elements (including data flow and data processing caliber information). 2.2) Modify the "logicless" leaves according to the logical reinforcement rules, for example, retain the leaf nodes pointed to by the processing process of the operator keyword "sum". 2.3) Visualize the lineage tree, using graphical means to intuitively display the lineage elements and their connections, to help users understand at a glance the data flow path, the application processing steps, and the ultimate use of the data.
[0070] 3) Bloodline interpretation. This stage mainly includes the following steps: 3.1) By interpreting the pruned bloodline tree obtained after executing step 2.2), a prompt is constructed, and the prompt is input into the big model to obtain the business language version of the bloodline interpretation (data bloodline analysis result) output by the big model.
[0071] 4) Prepare a report. Specifically, prepare a blood relationship interpretation result report based on the output of the large model. The blood relationship analysis result of the data output by the large model can be directly used as the blood relationship interpretation result report, or the blood relationship interpretation result report can be prepared by further processing the blood relationship analysis result of the data output by the large model.
[0072] Figure 3 The present invention provides a schematic diagram of a device for data lineage analysis according to an embodiment of the present invention. The device for data lineage analysis (hereinafter referred to as "data lineage analysis device 1") can be implemented as all or part of an electronic device through software, hardware or a combination of both. According to some embodiments, the data lineage analysis device 1 includes an acquisition module 11, a construction module 12, and an output module 13.
[0073] An obtaining module, configured to obtain, for an initial first target field, structured query processing information corresponding to a first target table to which the first target field belongs, wherein the structured query processing information is used to characterize a data processing process corresponding to the first target table;
[0074] A construction module is used to determine whether the termination condition is currently met based on the structured query processing information; if not, to perform upstream source disassembly on the structured query processing information, to obtain one or more second target fields located upstream of the first target field, to obtain the latest structured query processing information corresponding to the second target table to which the second target field belongs, and so on; if yes, to construct a data lineage relationship tree based on the obtained multiple target fields and the target table to which each target field belongs, wherein each node in the data lineage relationship tree corresponds to a target field, and the direction of the edge between two nodes is used to characterize the data flow direction between the target fields corresponding to the two nodes respectively;
[0075] The output module is used to generate corresponding prompt information by interpreting the data lineage relationship tree, input the prompt information into the trained data lineage analysis model, and obtain the data lineage analysis result output by the data lineage analysis model.
[0076] In some embodiments, judging whether a termination condition is currently satisfied according to the structured query processing information includes:
[0077] Simplifying the structured query processing information according to the first target field to obtain simplified structured query processing information;
[0078] According to the simplified structured query processing information, determining whether a termination condition is currently met;
[0079] Wherein, the upstream source disassembly of the structured query processing information includes:
[0080] The streamlined structured query processing information is disassembled from upstream sources.
[0081] In some embodiments, the simplifying the structured query processing information according to the first target field to obtain the simplified structured query processing information includes:
[0082] Determining, in the structured query processing information, non-associated text information that satisfies a preset association relationship with the first target field;
[0083] The structured query processing information is simplified by deleting the non-related text information to obtain simplified structured query processing information.
[0084] In some embodiments, the data lineage analysis device 1 is further used for:
[0085] Determining non-critical text information that meets a preset criticality level in the structured query processing information;
[0086] The step of simplifying the structured query processing information by deleting the non-related text information to obtain the simplified structured query processing information includes:
[0087] The structured query processing information is simplified by deleting the non-related text information and the non-key text information to obtain simplified structured query processing information.
[0088] In some embodiments, the termination condition includes at least one of the following:
[0089] The current level has reached the upper limit of the upstream level;
[0090] The first target field has no upstream dependency.
[0091] In some embodiments, the edge between the two nodes corresponds to structured processing aperture information between the target fields respectively corresponding to the two nodes, wherein the structured processing aperture information is used to characterize the data processing process between the target fields respectively corresponding to the two nodes;
[0092] The step of constructing a data lineage tree based on at least one target table involved in each acquired data level and at least one target field involved in each target table includes:
[0093] A data lineage tree is constructed based on the structured processing caliber information corresponding to at least one target table involved in each data level, at least one target field involved in each target table, and two target fields in an upper and lower hierarchical relationship.
[0094] In some embodiments, the data lineage analysis device 1 is further used for:
[0095] According to the structured query processing information, structured processing caliber information between the first target field and the second target field is determined.
[0096] In some embodiments, the data lineage analysis device 1 is further used for:
[0097] pruning nodes of the data kinship tree according to preset logic reinforcement rules;
[0098] The generating of corresponding prompt information by interpreting the data blood relationship tree includes:
[0099] By interpreting the pruned data kinship tree, corresponding prompt information is generated.
[0100] In some embodiments, the logic enforcement rule is used to indicate that structured processing aperture information of existence operator keywords is retained.
[0101] The above device embodiments correspond to the method embodiments. For specific descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, please refer to the corresponding method embodiments.
[0102] The embodiment of the present specification also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded by a processor to execute the method of the embodiment of the present specification.
[0103] The embodiments of the present specification also provide a computer program product, which stores at least one instruction, and the at least one instruction is loaded by the processor to execute the method of the embodiments of the present specification.
[0104] An embodiment of the present specification also provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the above method.
[0105] The embodiments of this specification also provide Figure 4 The structural diagram of the electronic device shown in FIG. Figure 4 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above method.
[0106] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0107] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0108] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0109] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0111] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0112] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0113] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0114] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.
Claims
1. A method for data lineage analysis, comprising: For an initial first target field, obtaining structured query processing information corresponding to a first target table to which the first target field belongs, wherein the structured query processing information is used to characterize a data processing process corresponding to the first target table; According to the structured query processing information, determine whether the termination condition is currently met; if not, perform upstream source disassembly on the structured query processing information to obtain one or more second target fields located upstream of the first target field, obtain the latest structured query processing information corresponding to the second target table to which the second target field belongs, and so on; if so, construct a data lineage relationship tree based on the obtained multiple target fields and the target table to which each target field belongs, wherein each node in the data lineage relationship tree corresponds to a target field, and the direction of the edge between two nodes is used to characterize the data flow direction between the target fields corresponding to the two nodes respectively; The data lineage relationship tree is interpreted to generate corresponding prompt information, and the prompt information is input into a trained data lineage analysis model to obtain a data lineage analysis result output by the data lineage analysis model.
2. According to the method of claim 1, judging whether a termination condition is currently satisfied based on the structured query processing information comprises: Simplifying the structured query processing information according to the first target field to obtain simplified structured query processing information; According to the simplified structured query processing information, determining whether a termination condition is currently met; Wherein, the upstream source disassembly of the structured query processing information includes: The streamlined structured query processing information is disassembled from upstream sources.
3. The method according to claim 2, wherein the step of simplifying the structured query processing information according to the first target field to obtain the simplified structured query processing information comprises: Determining, in the structured query processing information, non-associated text information that satisfies a preset association relationship with the first target field; The structured query processing information is simplified by deleting the non-related text information to obtain simplified structured query processing information.
4. The method according to claim 3, further comprising: Determining non-critical text information that meets a preset criticality level in the structured query processing information; The step of simplifying the structured query processing information by deleting the non-related text information to obtain the simplified structured query processing information includes: The structured query processing information is simplified by deleting the non-related text information and the non-key text information to obtain simplified structured query processing information.
5. The method according to claim 2, wherein the termination condition comprises at least one of the following: The current level has reached the upper limit of the upstream level; The first target field has no upstream dependency.
6. The method according to claim 1, wherein the edge between the two nodes corresponds to the structured machining aperture information between the target fields respectively corresponding to the two nodes, wherein: The structured processing aperture information is used to characterize the data processing process between the target fields respectively corresponding to the two nodes; The step of constructing a data lineage tree based on at least one target table involved in each acquired data level and at least one target field involved in each target table includes: A data lineage tree is constructed based on the structured processing caliber information corresponding to at least one target table involved in each data level, at least one target field involved in each target table, and two target fields in an upper and lower hierarchical relationship.
7. The method according to claim 6, further comprising: According to the structured query processing information, structured processing caliber information between the first target field and the second target field is determined.
8. The method according to claim 1, further comprising: pruning nodes of the data kinship tree according to preset logic reinforcement rules; The generating of corresponding prompt information by interpreting the data blood relationship tree includes: By interpreting the pruned data kinship tree, corresponding prompt information is generated.
9. The method according to claim 8, wherein the logic reinforcement rule is used to indicate that structured processing aperture information of existence operator keywords is retained.
10. A device for data lineage analysis, comprising: An obtaining module, configured to obtain, for an initial first target field, structured query processing information corresponding to a first target table to which the first target field belongs, wherein the structured query processing information is used to characterize a data processing process corresponding to the first target table; A construction module is used to determine whether the termination condition is currently met based on the structured query processing information; if not, to perform upstream source disassembly on the structured query processing information, to obtain one or more second target fields located upstream of the first target field, to obtain the latest structured query processing information corresponding to the second target table to which the second target field belongs, and so on; if yes, to construct a data lineage relationship tree based on the obtained multiple target fields and the target table to which each target field belongs, wherein each node in the data lineage relationship tree corresponds to a target field, and the direction of the edge between two nodes is used to characterize the data flow direction between the target fields corresponding to the two nodes respectively; The output module is used to generate corresponding prompt information by interpreting the data lineage relationship tree, input the prompt information into the trained data lineage analysis model, and obtain the data lineage analysis result output by the data lineage analysis model.
11. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
12. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the method as claimed in any one of claims 1 to 9.
13. A computer program product having at least one instruction stored thereon, characterized in that: When the at least one instruction is executed by the processor, the steps of the method described in any one of claims 1 to 9 are implemented.
Citation Information
Cited By
Data consanguinity analysis method of e-commerce data warehouse, product, equipment and medium
CN120725717A
Data lineage analysis method, product, device and medium of e-commerce data warehouse
CN120725717B
Method and device for constructing field consanguinity tree, storage medium and terminal
CN120780710A
Method, device, storage medium and terminal for constructing a field bloodline tree
CN120780710B