A data pedigree map display method, an electronic device, and a storage medium

By adding virtual nodes and removing back loops in the data lineage graph, the display complexity caused by loops is resolved, improving the visualization effect and user experience of the data lineage graph.

CN117131248BActive Publication Date: 2025-12-12ZHEJIANG MEIRI HUDONG NETWORK TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311084905.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-25
Publication Date
2025-12-12
Estimated Expiration
2043-08-25

AI Technical Summary

Technical Problem

When loops exist in the data lineage graph, a large number of nodes or lineage levels increase the difficulty of display and result in poor visualization, affecting the user experience.

Method used

When back connections exist in the data lineage diagram, virtual nodes are added and back connections are deleted based on specific conditions met by the number of nodes and the number of levels. The data flow path is displayed through virtual nodes, simplifying the display.

Benefits of technology

This effectively avoids the problems of chaotic data lineage diagrams and poor visualization caused by a large number of nodes or levels, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117131248B_ABST
    Figure CN117131248B_ABST
Patent Text Reader

Abstract

The application provides a data bloodline graph display method, an electronic device and a storage medium. The method comprises the following steps: S100, generating a corresponding basic data bloodline graph based on a center node selected by a user on a preset data bloodline graph; updating a current data bloodline graph based on a click operation of the user on the current data bloodline graph; if there is a back connection line in the current data bloodline graph, and the number of nodes and the number of levels meet a set condition, adding a virtual data node corresponding to a corresponding back connection data node behind a back connection processing node in the current data bloodline graph, connecting the virtual data node with the back connection processing node, and deleting the back connection line. The application can improve the visual display effect of the data bloodline graph and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer technology application, and in particular to a data bloodline graph display method, an electronic device and a storage medium. BACKGROUND

[0002] With the explosive growth of data, the relationship between data becomes more and more complex. In this background, data bloodline with characteristics such as plasticity and attribution will play an increasingly important role in data governance process. The bloodline of data is of great significance for analyzing data, tracking the dynamic evolution of data, measuring the credibility of data and ensuring the quality of data. At present, a data bloodline graph is generated based on the data bloodline analysis result, so that users can intuitively know the flow path of data. The data bloodline graph is generally composed of data nodes, processing nodes and connection lines. The processing node is used to mark the processing mode and processing rule required for the upstream data node connected by the processing node to flow into the downstream data node. However, in some application scenarios, there may be data loop conditions, for example, a parent node data may be obtained after the child node data passes through the processing node, in this case, two same node data are generally connected back through the connection line. However, this connection back through the connection line increases the display difficulty and leads to poor visualization effect in the case of a large number of nodes or a large number of bloodline levels, which makes it difficult to accurately know the two data nodes with connection back relationship, and affects the user experience. SUMMARY

[0003] In view of the above technical problems, the technical scheme adopted by the present application is as follows:

[0004] The present application provides a data bloodline graph display method, which comprises the following steps:

[0005] S100, generating a corresponding basic data bloodline graph based on a center node selected by a user on a preset data bloodline graph; the preset data bloodline graph comprises data nodes, processing nodes and connection lines connecting the data nodes and the processing nodes, and the connection lines have directions;

[0006] S200, updating the current data bloodline graph based on a click operation of the user on the current data bloodline graph;

[0007] S300, if there is a back connection line in the current data bloodline graph, and if m>D2, or D1

[0008] The embodiment of the present application further provides an electronic device, comprising a processor and the non-transitory computer readable storage medium.

[0009] The embodiment of the present application further provides an electronic device, comprising a processor and the non-transitory computer readable storage medium.

[0010] The embodiment of the present application has at least the following beneficial effects:

[0011] The data bloodline graph display method provided by the embodiment of the present application can delete the back connection line and display the corresponding virtual node when there is a loop in the displayed data bloodline graph, if the node number and the level number of the data bloodline graph meet the preset condition, thus avoiding the problem of data bloodline graph confusion and poor visualization effect caused by too many nodes or too many bloodline levels. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without any creative effort based on these drawings.

[0013] Figure 1 The flow chart of the data bloodline graph display method provided by the embodiment of the present application. DETAILED DESCRIPTION

[0014] With reference to the drawings and the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments of the present application, all the other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0015] It should be noted that the terms "first", "second" and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate specified orders or sequences. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in other than the order illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units need not be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0016] The embodiments of the present application provide a data bloodline graph display method, and the technical thought is that when a data bloodline graph exists a loop, the current displayed data bloodline graph is adjusted based on the number of nodes and the number of levels of the current displayed data bloodline graph, so as to avoid the problems of data bloodline graph confusion and poor visualization effect caused by too many nodes or too many bloodline levels.

[0017] Further, as shown in Figure 1 The data bloodline graph display method provided by the embodiments of the present application can include the following steps:

[0018] S100, generating a corresponding basic data bloodline graph based on a center node selected by a user on a preset data bloodline graph; the preset data bloodline graph includes data nodes and processing nodes and a connection line connecting the data nodes and the processing nodes, and the connection line has a direction. The connection line can be a straight line or an arc line.

[0019] In the embodiments of the present application, the preset data bloodline graph can be generated based on data in a preset space. The preset space can be a self-defined space, for example, a data processing platform.

[0020] In the embodiments of the present application, the data nodes can include structured data nodes and unstructured data nodes. The structured data nodes can include data table nodes (MySQL, Oracle, Hive, PostgreSQL, HBase, SQLServer, Impala, ClickHouse, Iceberg, Jdbc, MongoDB, etc.), data table field nodes, API data nodes, etc., and the unstructured data nodes can include file nodes (HDFS, FTP, SFTP, Codis, S3, ODPS, HETU, LocalFile, JuiceFS, OwnCloud, etc.), message queue nodes (Kafka, ActiveMQ, RocketMQ, etc.), etc.

[0021] In the embodiments of the present application, the processing node is used to mark the processing mode and processing rule required for the upstream data node connected to the processing node to be transferred to the downstream data node, which is determined based on the task implemented by the corresponding workflow.

[0022] Those skilled in the art know that any method for generating a corresponding data bloodline graph based on data in space is within the protection scope of the present application.

[0023] In the embodiments of the present application, each node can be drawn using a solid line circle or a box. Each node of the preset data bloodline graph is provided with corresponding node attribute information. The attribute information of the processing node can at least include a node name, a responsible person, a creation time, an execution time, an execution engine, a workflow name, a workflow ID, a source system, a task main function path, a project, an execution tenant, an execution queue, a running cluster, etc. For the data table node, the corresponding node attribute information can include a node name, a node type (the above structured data type, such as MySQL, etc.), a data source name, a table name, an asset name, an asset path, a subject domain, data layering, a source system, an update frequency, a responsible person, a creation time, an update time, a storage cluster, etc. For the file node, the corresponding node attribute information can include a node name, a node type (the above unstructured data type, such as HDFS, etc.), a data source name, a file path, a subject domain, data layering, a source system, an update frequency, a responsible person, a creation time, an update time, a storage cluster, etc. For the API node, the corresponding node attribute information can include a node name, a URL, a request method, a request header, a request input parameter, a request output parameter, a responsible person, a creation time, etc. For the message queue node, the corresponding node attribute information can include a node name, a node type (Kafka, etc.), a data source name, an asset name, a subject, a data format, a separator, field information, a responsible person, a creation time, etc.

[0024] In the embodiments of the present application, the center node can be a data node or a processing node. If the center node is a data node, the corresponding basic data bloodline graph includes data nodes and processing nodes upstream of the center node and data nodes and processing nodes downstream of the center node, i.e., the corresponding basic data bloodline graph is a 5-layer structure. If the center node is a processing node, the corresponding basic data bloodline graph includes data nodes upstream of the center node and data nodes downstream of the center node, i.e., the corresponding basic data bloodline graph is a 3-layer structure.

[0025] S200, updating the current data bloodline graph based on the clicking operation of the user on the current data bloodline graph.

[0026] Further, S200 can specifically include:

[0027] S201, adding corresponding nodes in the current data bloodline graph based on the clicking operation of the user on the current data bloodline graph.

[0028] In the embodiments of the present application, an expansion mark, such as "+", can be set on the nodes of the last layer of the data bloodline graph. When it is detected that the user clicks the expansion mark of a node, the downstream nodes corresponding to the node are expanded, i.e., corresponding nodes are added in the current data bloodline graph.

[0029] S202, if the number of nodes in the current data bloodline graph is greater than a preset threshold, taking the node corresponding to the clicking operation as a new center node, and displaying other nodes in the current data bloodline graph in a hidden state except the basic data bloodline graph corresponding to the new center node.

[0030] In the embodiments of the present application, if the number of nodes in the current data bloodline graph is greater than a preset threshold, it means that the number of nodes is too large, which can cause the visual effect to be poor. Since the clicking operation of the user can reflect the idea of the user who wants to see the subsequent nodes, the node corresponding to the clicking operation can be taken as a new center node, and other nodes in the current data bloodline graph can be displayed in a hidden state except the basic data bloodline graph corresponding to the new center node, i.e., only the basic data bloodline graph corresponding to the new center node is reserved. In this way, the display interface can be simple and the visual effect can be good. It is known to those skilled in the art that the hidden state display can be achieved by using existing methods.

[0031] In the embodiments of the present application, the preset value can be determined based on actual needs, and can be specifically determined based on the shape and size of the nodes and the shape and size of the connection lines. In one non-limiting illustrative embodiment, the preset threshold can be 300.

[0032] S300, if there is a back connection line in the current data pedigree, and if m>D2, or D1

[0033] In the embodiment of the present application, the direction of the back connection line is opposite to the data flow direction in the current data pedigree, for example, the data flow direction of the current data pedigree is forward, and the direction of the back connection line is backward.

[0034] In the embodiment of the present application, only data nodes have virtual nodes, and the virtual nodes are the abstraction of the data nodes to which the back connection processing nodes are connected, and are only used for display, and are actually the same nodes. The virtual nodes and the corresponding real data nodes have the same attributes and operations, and the shape of the virtual data nodes can be the same as that of the corresponding original nodes, and the difference is that the virtual data nodes are drawn using a dashed line.

[0035] The inventor of the present application has found through practice that the number of nodes and the level of pedigree are positively correlated with the complexity of a workflow. In the embodiment of the present application, if m>D2, it indicates that the number of nodes in the current data pedigree is large but the level is low, so the number of longitudinal nodes is large, and the back connection line needs to bypass a large number of longitudinal nodes to connect to the original node, which can result in poor page display effect. If a user wants to know to which node the back connection line is connected, the user needs to drag the current data pedigree, which can result in poor visual effect and poor user experience. If D1

[0036] In the embodiments of the present application, D1 to D3 can be empirical values, and can be obtained through user experience of N test users. Specifically, D1 to D3 are obtained through the perceptual experience of N test users on the data bloodline graph with different numbers of nodes and different levels of backlink lines. In an illustrative embodiment, D1 = 10, D2 = 50, and D3 = 7, that is, when the number of nodes is greater than 50 or the number of nodes is between 10 and 50 but the number of levels is greater than 7, the user experience is the worst, and the backlink line needs to be set as a virtual node.

[0037] As known by those skilled in the art, in the case of n≤D3, or, m≤D1, or, D1≤m<D2 and n≤7, the current data bloodline graph can be maintained, that is, the backlink line does not need to be set as a virtual node.

[0038] In the embodiments of the present application, for the center node, a center node identifier is also set to represent that the node is a center node. The center node identifier can be set based on actual needs, and the present application does not make special limitations.

[0039] Further, in the embodiments of the present application, the method further comprises:

[0040] If it is detected that a user clicks any one of the virtual data node and the corresponding backlink data node, the virtual data node and the corresponding backlink data node are displayed in linkage, that is, when one of the virtual node and the corresponding backlink data node is selected, the other also presents a selected state, for example, in the form of highlighting or highlighting plus flashing, so as to enable the user to quickly know the data flow path, and the visual effect is good.

[0041] Further, the method provided in the embodiments of the present application further comprises:

[0042] S400, if external data imported from other spaces into the preset space is received, the current preset data bloodline graph is updated.

[0043] Specifically, S400 can specifically comprise:

[0044] S401, the ID of any node in the external data is used to query in the preset database, if the query is found, the node is marked, and if the query is not found, the node is created in the preset database and added in the preset data bloodline graph. The imported external data at least includes attribute information of each node.

[0045] S402, based on the upstream and downstream node information of the node, the preset data bloodline graph is updated; wherein the upstream and downstream node information is obtained based on the task engine and the corresponding analysis rule included in the space corresponding to the external data.

[0046] Specifically, when the task engine is Hive SQL, the field lineage can be parsed by trying the Hive hook function such as org.apache.hadoop.hive.ql.hooks.LineageLogger. When the task engine is Spark SQL, the field lineage can be parsed by obtaining the Output of the logical plan through the onSuccess method of QueryExecutionListener. When the task engine is Flink SQL, the logical plan tree (AST) of the SQL can be obtained through Calcite, and the Input\Output can be obtained by traversing the AST, so as to parse the field lineage. For the task engine being Spark\Mr, the field lineage can be parsed by obtaining the Input\Output through the framework specification input and output parameters, and then analyzing the yarn log.

[0047] Those skilled in the art know that the field lineage parsed based on the corresponding parsing rule can be prior art.

[0048] In the embodiment of the application, the parsing of the lineage of the external data can be performed in the corresponding space or in the preset space, and preferably, in the preset space, so that the external space only needs to directly import data.

[0049] The embodiment of the application further provides a non-transitory computer readable storage medium, which can be arranged in an electronic device to save at least one instruction or at least one program related to a method in the method embodiment, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided by the above-mentioned embodiment.

[0050] The embodiment of the application further provides an electronic device, which comprises a processor and the aforementioned non-transitory computer readable storage medium.

[0051] The embodiment of the application further provides a computer program product, which comprises program code, and when the program product runs on an electronic device, the program code is used to make the electronic device execute the steps in the method according to various exemplary embodiments of the application described in the specification.

[0052] Although some specific embodiments of the application have been described in detail by examples, those skilled in the art should understand that the above examples are only for illustration, but not for limiting the scope of the application. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the application. The scope disclosed by the application is defined by the appended claims.

Claims

1. A data pedigree map display method, characterized by, The method comprises the following steps: S100, generating a corresponding basic data bloodline graph based on a center node selected by a user on a preset data bloodline graph; the preset data bloodline graph comprises data nodes and processing nodes and connecting lines connecting the data nodes and the processing nodes, and the connecting lines have directions; S200, updating a current data bloodline graph based on a click operation of the user on the current data bloodline graph; S300, if there is a back connection line in the current data bloodline graph, and if m>D2 or D1 2. The method of claim 1, wherein, S200 specifically comprises: S201, based on a click operation of the user on the current data bloodline graph, adding a corresponding node in the current data bloodline graph; S202, if the number of nodes in the current data bloodline graph is greater than a preset threshold, taking the node corresponding to the click operation as a new center node and hiding other nodes in the current data bloodline graph except the basic data bloodline graph corresponding to the new center node.

3. The method of claim 1, wherein, If the center node is a data node, the corresponding basic data bloodline graph comprises data nodes and processing nodes located upstream of the center node and data nodes and processing nodes located downstream of the center node; If the center node is a processing node, the corresponding basic data bloodline graph comprises data nodes located upstream of the center node and data nodes located downstream of the center node.

4. The method of claim 1, wherein, Further comprising: If it is detected that the user clicks any one of the virtual data node and the corresponding back connection data node, the virtual data node and the corresponding back connection data node are displayed in linkage.

5. The method of claim 1, wherein, The preset data bloodline graph is generated based on data in a preset space; Further comprising: S400, if external data imported from other spaces into the preset space is received, updating a current preset data bloodline graph.

6. The method of claim 5, wherein, S400 specifically comprises: S401, querying a preset database by using the ID of any node in the external data, if the query is successful, marking, if the query is not successful, creating the node in the preset database and adding the node in the preset data bloodline graph; S402, updating the preset data bloodline graph based on upstream and downstream node information of the node; wherein the upstream and downstream node information is obtained based on a task engine and corresponding analysis rules included in a space corresponding to the external data.

7. The method of claim 1, wherein, D1=10, D2=50, D3=7.

8. The method of claim 1, wherein, Each node of the preset data bloodline graph is provided with corresponding node attribute information, and the center node is further provided with a center node identifier. 9.A non-transitory computer-readable storage medium having stored therein at least one instruction or at least one piece of program, characterized in that, The at least one instruction or the at least one program is loaded and executed by the processor to implement the method of any one of claims 1-8.

10. An electronic device, comprising: A non-transitory computer readable storage medium as claimed in claim 9 and a processor.

Citation Information

Patent Citations

  • Metadata blood relationship processing method and device, storage medium and electronic device

    CN110399423A

  • Data blood relationship analysis system and method based on metadata management

    CN115757655A