Data lineage processing method, apparatus, device, and storage medium
By generating a data lineage diagram that matches role permissions, the problem of complex data relationships in the data warehouse that rely on manual analysis is solved. This enables automated collection, monitoring, and backtracking of data lineage, improving the usability, accuracy, and readability of the data warehouse.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2022-01-25
- Publication Date
- 2026-07-24
AI Technical Summary
Data warehouses contain complex data relationships, making manual analysis difficult. This results in poor usability, low accuracy, poor readability, and difficulty in repairing data anomalies, which is time-consuming and labor-intensive.
By receiving upstream and downstream node information of data lineage that matches role permissions, a data lineage relationship diagram matching role permissions is generated, supporting personalized display and automated monitoring, providing a data lineage visualization system, and realizing automated data lineage collection, monitoring and backtracking.
It improves the stability, usability, accuracy and readability of the data warehouse, reduces operation and maintenance costs, and increases operational efficiency and the ease of data anomaly repair.
Smart Images

Figure CN114491450B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to the fields of big data, data warehousing and data lineage. Background Technology
[0002] A data warehouse is a collection of data. Data warehouses play a crucial guiding role in business operations, and the direction of business development often relies on data-driven processes. The relationships between data in a data warehouse are complex, making manual analysis difficult. Summary of the Invention
[0003] This disclosure provides a data lineage processing method, apparatus, device, and storage medium.
[0004] According to one aspect of this disclosure, a data lineage processing method is provided, comprising:
[0005] Receive information about upstream and downstream nodes of data lineage that matches role permissions, wherein the information about upstream and downstream nodes is obtained by performing data lineage analysis on the collected data;
[0006] Generate a data lineage diagram that matches the role permissions based on the information of the upstream and downstream nodes.
[0007] According to another aspect of this disclosure, a data lineage processing method is provided, comprising:
[0008] Information about upstream and downstream nodes of the data lineage that matches the role's permissions is sent to the front end. This information is used to generate a data lineage graph that matches the role's permissions.
[0009] According to another aspect of this disclosure, a data lineage processing apparatus is provided, comprising:
[0010] The first receiving module is used to receive information about upstream and downstream nodes of data lineage that matches role permissions. The information about upstream and downstream nodes is obtained by performing data lineage analysis on the collected data.
[0011] The generation module is used to generate a data lineage diagram that matches the role permissions based on the information of the upstream and downstream nodes.
[0012] According to another aspect of this disclosure, a data lineage processing apparatus is provided, comprising:
[0013] The first sending module is used to send information about upstream and downstream nodes of the data lineage that matches the role's permissions to the front end. The information about the upstream and downstream nodes is used to generate a data lineage relationship diagram that matches the role's permissions.
[0014] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0015] At least one processor; and
[0016] The memory is communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods of any embodiment of the present disclosure.
[0018] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform a method according to any embodiment of this disclosure.
[0019] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements a method according to any embodiment of this disclosure.
[0020] Using this disclosure can significantly improve the stability, usability, accuracy, and readability of data warehouses.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0022] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0023] Figure 1A This is a schematic flowchart of a data lineage processing method according to an embodiment of the present disclosure;
[0024] Figure 1B This is an exemplary data lineage diagram according to an embodiment of the present disclosure;
[0025] Figure 2 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure;
[0026] Figure 3 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure;
[0027] Figure 4 This is a schematic flowchart of a data lineage processing method according to an embodiment of the present disclosure;
[0028] Figure 5 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure;
[0029] Figure 6 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure;
[0030] Figure 7 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure;
[0031] Figure 8 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure;
[0032] Figure 9 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure;
[0033] Figure 10 This is a schematic diagram of the structure of a data lineage processing apparatus according to an embodiment of the present disclosure;
[0034] Figure 11 This is a schematic diagram of the structure of a data lineage processing apparatus according to another embodiment of the present disclosure.
[0035] Figure 12 This is a schematic diagram of the structure of a data lineage processing apparatus according to an embodiment of the present disclosure;
[0036] Figure 13 This is a schematic diagram of the structure of a data lineage processing apparatus according to another embodiment of the present disclosure;
[0037] Figure 14 This is a schematic diagram of the structure of a data lineage processing apparatus according to another embodiment of the present disclosure;
[0038] Figure 15 This is an architecture diagram of a data lineage analysis system according to an embodiment of the present disclosure;
[0039] Figure 16 This is a schematic diagram of a data lineage analysis process according to an embodiment of the present disclosure;
[0040] Figure 17 This is a schematic diagram showing the display effect of a data kinship diagram according to an embodiment of the present disclosure;
[0041] Figure 18 This is a schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure. Detailed Implementation
[0042] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0043] Figure 1A This is a schematic flowchart of a data lineage processing method according to an embodiment of the present disclosure. The method includes:
[0044] S101. Receive information about upstream and downstream nodes of data lineage that matches role permissions, wherein the information about upstream and downstream nodes is obtained by performing data lineage analysis on the collected data.
[0045] S102. Generate a data lineage diagram that matches the role permissions based on the information of the upstream and downstream nodes.
[0046] In this embodiment, the data lineage processing method can be used in systems such as data warehouses. For example, a data warehouse may include a front-end and a back-end. The front-end may include various types of clients. The back-end may include various types of servers. In the back-end, data from multiple data sources can be collected, and data lineage analysis can be performed on the collected data. Then, the back-end can send information about upstream and downstream nodes of the data lineage matching the role and permissions to the front-end based on role and permissions. After receiving the information about the upstream and downstream nodes, the front-end can generate a data lineage relationship diagram matching the role and permissions based on the information about the upstream and downstream nodes. Furthermore, the front-end can display the data lineage relationship diagram in a UI (User Interface).
[0047] In this embodiment of the disclosure, there may be multiple different roles. For example, in a data warehouse, there may be multiple roles such as administrator, analyst, and user. Different roles may have different operation permissions on the data tables in the data warehouse. For example, the administrator has permissions to view, modify, and delete data tables A1, A2, A3, B1, B2, C1, and C2; analyst 1 has permissions to view and modify A1, A2, and A3; analyst 2 has permissions to view and modify B1, B2, C1, and C2; user 1 has permissions to view B1 and B2; user 2 has permissions to view C1 and permissions to view and modify C2.
[0048] In this embodiment, data lineage represents the origin and flow of data, mainly including the data source, data processing method, mapping relationship, and data exit point. For example, multiple data tables A1, A2, and A3 may have a certain lineage relationship. A1 may be the source of A2, and the data exit point of A2 may be A3. The backend performs data lineage analysis on the collected data to obtain the upstream and downstream relationships of the data lineage, and thus obtain the information of the upstream and downstream nodes of the data lineage matching the role's permissions. The information of the upstream and downstream nodes may include the information of each node and the relationship between each node. Different nodes may correspond to different data tables. The frontend can generate a data lineage relationship diagram matching the role's permissions based on the received upstream and downstream node information. For example, as shown in the example... Figure 1B As shown, the data lineage graph can include nodes N1, N2, N3, N4, N5, and N6. N1 is the root node, N2 and N3 are child nodes of N1, N4 is a child node of N2, and N5 and N6 are child nodes of N3. Therefore, in the data lineage graph, N1 can be the top-level node, N2 and N3 can be second-level nodes with edges connecting to N1, and N4, N5, and N6 are third-level nodes, with N4 connected to N2 and N5 and N6 connected to N3. The hierarchical and connection relationships of the nodes in this data lineage graph are for illustrative purposes only and are not limitations. In practical applications, corresponding data lineage graphs can be generated based on specific role permissions.
[0049] After displaying the data lineage diagram on the front end, users can view relevant information about the data tables corresponding to each node in the diagram. For example, node N1 corresponds to data table A1; clicking node N1 can navigate to the page containing data table A1, or display a notification window asking whether to open data table A1.
[0050] In this embodiment, a data lineage diagram matching role permissions can be generated based on the information of upstream and downstream nodes in the data lineage. This allows for the provision of different data lineage diagrams for different roles and permissions. Therefore, personalized data lineage diagrams can be displayed, providing a unique visualization effect for each user.
[0051] Figure 2 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure. The method of this embodiment includes one or more features of the above-described data lineage processing method embodiments. In one possible implementation, the method further includes:
[0052] S201. Display the corresponding color in the data lineage diagram according to the state of the node; nodes in different states display different colors. Distinguishing different node states using different colors in the data lineage diagram helps highlight nodes requiring attention, facilitates rapid location, and improves operational efficiency.
[0053] In one possible implementation, the node's state includes at least one of the following: successful execution, failed execution, running, and waiting to run. These states are merely examples and not limitations, and the node's state can be modified, deleted, or added according to the needs of the actual application.
[0054] In this embodiment, successful execution corresponds to a first color, failed execution corresponds to a second color, currently running corresponds to a third color, and awaiting execution corresponds to a fourth color. The first, second, third, and fourth colors can be different colors. Furthermore, these colors can have significant differences to facilitate operator identification. For example, the first color could be green, the second red, the third white, and the fourth gray. Using different colors to distinguish the states of nodes (successful execution, failed execution, currently running, awaiting execution, etc.) in the data lineage diagram helps highlight nodes in a certain state, such as failed execution, facilitating quick location by operators and improving operational efficiency.
[0055] In one possible implementation, the data lineage graph includes a directed acyclic graph (DAG). For example, a DAG can include a directed graph without loops. A DAG can describe priority relationships or dependencies. Using a DAG in the data lineage graph can represent the upstream and downstream relationships between nodes.
[0056] In one possible implementation, the data lineage graph further includes task information, which includes at least one of the following:
[0057] Task time, node identifier, node path, Service-Level Agreement (SLA) description, whether it meets the requirements, and details.
[0058] In this embodiment of the disclosure, each role may utilize data within its authorized scope to perform various types of tasks, such as text recognition, speech recognition, facial recognition, and cloud computing. Each task may require different data. For example, task T1 needs to operate on data tables D1 and D2, while task T2 needs to operate on data table D3. Different tasks may correspond to different nodes in the data lineage graph. For example, task T1 corresponds to nodes N7 and N8, and task T2 corresponds to node N9.
[0059] The data lineage diagram can also display various task information. Task time can include at least one of the task start time, end time, and duration; for example, task T1 starts at xx hour xx minute xx second and lasts for 1 hour. Node identifiers can include the name, number, index, etc., of the node corresponding to the data table that a task needs to operate on. For example, the node identifiers for task T1 include N7 and N8. Node paths can include the storage path of the data table or table entries corresponding to the node. For example, data table D1 for task T1 is stored in URL1. The SLA description can include the goals that the task needs to achieve. Whether the goals are met can include whether the task has achieved the goals in the SLA description. Details can include various situations that may occur with the task, such as what type of exception occurred, and which nodes have upstream and downstream relationships.
[0060] In this embodiment of the disclosure, displaying task information can quickly show the relationship between tasks and nodes, thereby facilitating the rapid execution of required operations on certain nodes.
[0061] In one possible implementation, the method further includes:
[0062] S202. In response to the selection operation of the target node that needs to be repaired, a data repair request is sent, wherein the data repair request includes the identifier of the target node that needs to be repaired.
[0063] In this embodiment, some nodes may malfunction, such as failing to run, requiring repair. The abnormal nodes can be quickly located using colors displayed in the data lineage diagram. If an abnormal node needs to be operated on, the operator can select that node for repair. Specific selection operations can include various methods. For example, directly selecting the abnormal node with the mouse. Alternatively, after selecting the abnormal node with the mouse, a pop-up window prompts whether repair is needed, allowing the operator to select repair again within the prompt window. After the operator selects the abnormal node in the UI, the front-end can send a data repair request to the back-end, carrying the identifier of the abnormal node requiring repair. Using the identifier of the target node in the data repair request, the back-end can be requested to repair the target node. Furthermore, the back-end can also find data from other nodes related to the target node based on its identifier. Then, if any of these nodes are abnormal, they can also be repaired. For example, downstream nodes of the target node can be found based on its identifier, and similar repairs can be performed on the downstream nodes.
[0064] In one possible implementation, the method further includes:
[0065] S203. If the repair is successful, receive the updated status of the target node. The updated status of the target node allows for updating the data lineage diagram, displaying a more accurate node status. Furthermore, it can also receive updated statuses of other nodes related to the target node, changing the color of these nodes displayed in the data lineage diagram. For example, if target node N2 changes from an abnormal state to a normal state, its color changes to green in the data lineage diagram. Further, if N4, downstream of target node N2, also changes from an abnormal state to a normal state, N4 also changes to green.
[0066] Figure 3 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure. The method of this embodiment includes one or more features of the above-described data lineage processing method embodiments. In one possible implementation, the method further includes:
[0067] S301. In response to a query operation, a query request is sent, the query request including the time to be queried and / or the identifier of the node to be queried;
[0068] S302. Receive query results, wherein the query results are obtained by querying the information of upstream and downstream nodes in the data lineage based on the time to be queried and / or the identifier of the node to be queried.
[0069] S303. Generate a data lineage diagram based on the query results.
[0070] In this embodiment of the disclosure, controls supporting query functionality, such as input boxes and query buttons, can be displayed on the data lineage graph of the UI. Input boxes may include date input boxes and / or keyword input boxes. In the date input box, the user can enter the time to be queried, such as the date or time; in the keyword input box, the user can enter keywords to be queried, such as node identifiers. The query operation may include entering the information to be queried in the input box and / or clicking the query button. After the user performs a query operation on the UI, the front end may send a query request to the back end in response to the query operation. This query request may include the time to be queried and / or the identifier of the node to be queried entered by the user.
[0071] In this embodiment, the backend can save the data lineage diagram by date and / or time. For example, one data lineage diagram can be saved daily. Alternatively, N data lineage diagrams can be saved daily. Upon receiving a query request, if the query request includes a time (e.g., date), the backend can search for information on upstream and downstream nodes in the data lineage that match the operator's permissions for that date. If the query request also includes a node identifier, the backend can further search for information on upstream and downstream nodes related to that node within the upstream and downstream node information for that date. When searching based on node identifiers, a fuzzy matching method can be used to find more nodes related to that node identifier.
[0072] The backend can return information about the upstream and downstream nodes found, i.e., the query results, to the frontend. The frontend can then generate a corresponding data lineage diagram based on the received information. The data lineage diagram generated by the frontend based on the query results may be the same as or different from the default data lineage diagram displayed.
[0073] In this embodiment of the disclosure, it is possible to support the query of data lineage, thereby obtaining the required data lineage relationship more quickly.
[0074] Figure 4 This is a schematic flowchart of a data lineage processing method according to an embodiment of the present disclosure. The method includes:
[0075] S401. Send the information of upstream and downstream nodes of the data lineage that matches the role permissions to the front end. The information of the upstream and downstream nodes is used to generate a data lineage relationship diagram that matches the role permissions.
[0076] In this embodiment, the data lineage processing method can be used in systems such as data warehouses. For example, a data warehouse may include a front-end and a back-end. The front-end may include various types of clients. The back-end may include various types of servers. In the back-end, data from multiple data sources can be collected, and data lineage analysis can be performed on the collected data. Then, the back-end can send information about upstream and downstream nodes of the data lineage matching the role's permissions to the front-end based on role permissions. After receiving the upstream and downstream node information, the front-end can generate a data lineage relationship diagram matching the role's permissions based on the upstream and downstream node information. Furthermore, the front-end can display this data lineage relationship diagram in the UI.
[0077] In this disclosed embodiment, there may be multiple different roles. For example, in a data warehouse, there may be multiple roles such as administrator, analyst, and user. Different roles may have different operation permissions on the data tables in the data warehouse.
[0078] In this embodiment of the disclosure, data lineage can represent the origin and development of data, which can mainly include the source of data, the processing method of data, the mapping relationship, and the data output.
[0079] After the data lineage diagram is displayed on the front end, relevant information about the data tables corresponding to each node in the data lineage diagram can be viewed.
[0080] In this disclosure, specific examples of roles, data lineage, etc., can be found in the relevant descriptions of the above method embodiments, and will not be repeated here.
[0081] In this embodiment, data lineage analysis can be performed on the collected data, and then information on upstream and downstream nodes of the data lineage matching the role and permissions can be sent to the front end. This allows the front end to generate a data lineage diagram matching the role and permissions, enabling the provision of different data lineage diagrams based on different roles and permissions. Therefore, personalized display of the data lineage diagram can be achieved, presenting a unique visualization effect for each user.
[0082] In one possible implementation, the data lineage graph includes a directed acyclic graph (DAG). For example, a DAG can include a directed graph without loops. A DAG can describe priority relationships or dependencies. Using a DAG in the data lineage graph can represent the upstream and downstream relationships between nodes.
[0083] Figure 5 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure. The method of this embodiment includes one or more features of the above-described data lineage processing method embodiments. In one possible implementation, the method further includes:
[0084] S501. Perform data lineage analysis on the collected data to obtain the upstream and downstream relationships of the data lineage;
[0085] S502. Based on the upstream and downstream relationships of the data lineage, generate information on the upstream and downstream nodes of the data lineage.
[0086] In this embodiment, the backend can perform data lineage analysis on all collected data to obtain the upstream and downstream relationships of the data lineage. For example, analyzing data tables A1, A2, A3, B1, B2, C1, and C2 reveals that the downstream of A1 includes A2, B1, and C1; the downstream of A2 includes A3; the downstream of B1 includes B2; and the downstream of C1 includes C2. Based on the upstream and downstream relationships of the data lineage, information on the upstream and downstream nodes of the data lineage matching the role's permissions can be obtained. This information can include information about each node and the relationships between them. Different nodes may correspond to different data tables. For example, A1 corresponds to N1, A2 to N2, B1 to N3, C1 to N4, A2 to N5, B2 to N6, and C2 to N7. Then, the backend can send the information on the upstream and downstream nodes of the data lineage matching the role's permissions to the frontend. For example, if a user U1's role permissions include operation permissions on data tables A1, A2, and A3, the information of upstream and downstream nodes sent to this user may include: A1 corresponds to N1, A2 corresponds to N2, C1 corresponds to N4, and may also include that the downstream node of N1 is N2, the downstream node of N2 is N4, etc.
[0087] The front-end can generate a data lineage diagram matching the permissions of the role based on the received information from upstream and downstream nodes. This data lineage diagram can include the hierarchical relationship and connection relationship of the nodes. For example, the data lineage diagram of U1 can include nodes N1, N2 and N4, and there is an edge between N1 and N2 that points to N2; there is also an edge between N2 and N4 that points to N4.
[0088] In this embodiment of the disclosure, the information of upstream and downstream nodes is obtained based on the upstream and downstream relationship of the data lineage that matches the role permissions, which can support providing different data lineage diagrams for different roles and different permissions.
[0089] In one possible implementation, the collected data includes input information and / or output information of the task. S501 performs data lineage analysis on the collected data to obtain the upstream and downstream relationships of the data lineage, including: performing data lineage analysis on the input information and / or output information of each task to obtain the upstream and downstream relationships of each task.
[0090] In the embodiments of this disclosure, various types of tasks, such as text recognition, speech recognition, face recognition, and cloud computing, may require different data. The input and / or output information of a task may include the input-output logic between the data required by the task. Based on the input-output logic between the data, the upstream and downstream relationships of each task can be obtained, thereby generating information about the upstream and downstream nodes of the data lineage of each task.
[0091] Figure 6 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure. The method of this embodiment includes one or more features of the above-described data lineage processing method embodiments. In one possible implementation, the method further includes:
[0092] S601: Collect data from multiple data sources.
[0093] In this embodiment, data can be collected from big data platforms, etc. For example, data can be collected from big data platforms using MR (MapTask & ReduceTask) services, Hive services, QE (Query Engine) services, Turing clients, Turing pages, Spark clients, LS (longscheduler) page settings, etc. Here, MR is the task processing stage of Hadoop. Hive is a data warehouse tool based on Hadoop, capable of storing, querying, and analyzing large-scale data stored in Hadoop. QE is a query engine that can use SQL (Structured Query Language) for queries. Spark is a fast and general-purpose computing engine for large-scale data processing. LS is a platform that can be scheduled on a timer.
[0094] In this embodiment of the disclosure, the collected data may include the relationships between data dependencies during the processing of AFS files, Hive tables, MySQL datasets, PALO datasets, etc., using Spark and Hadoop big data processing frameworks. AFS is an open-source HDFS (Hadoop Distributed File System).
[0095] You can first acquire a large amount of data from a big data platform, perform data lineage analysis on the collected data, and then select the analysis results that match the role's permissions to send to the front end for display. Collecting data from multiple data sources can improve data completeness and provide a rich data foundation for subsequent data lineage analysis.
[0096] In one possible implementation, the information of the upstream and downstream nodes includes the node's state, which includes at least one of the following: successful execution, failed execution, running, and waiting to run. These states are merely examples and not limitations; the node's state can be modified, deleted, or added according to the needs of the actual application. Nodes with different states can be displayed in different colors on the front end; please refer to the relevant description of the front end execution method embodiment for details.
[0097] By distinguishing between the states of nodes such as successful execution, failed execution, running, and waiting to run, it is helpful to highlight nodes in a certain state, such as failed execution, so that operators can quickly locate them and improve operational efficiency.
[0098] Figure 7 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure. The method of this embodiment includes one or more features of the above-described data lineage processing method embodiments. In one possible implementation, the method further includes:
[0099] S701. Send the task information to be displayed to the front end. The task information includes at least one of the following: task time, node identifier, node path, SLA description, whether the target is met, and details.
[0100] In this embodiment of the disclosure, the task time may include at least one of the task start time, end time, and duration. For example, the start time of task T1 is xx hour xx minute xx second, and the duration is 1 hour. The node identifier may include the name, number, index, etc., of the node corresponding to the data table that a task needs to operate on. For example, the node identifiers for task T1 include N7 and N8. The node path may include the storage path of the data table corresponding to the node or the table entry in the data table. For example, the data table D1 of task T1 is stored in URL1. The SLA description may include the goals that the task needs to achieve. Whether the goals are met may include whether the task has achieved the goals in the SLA description. The details may include various situations that may occur with the task, such as what type of exception occurred, and which nodes have upstream and downstream relationships.
[0101] In this embodiment of the disclosure, sending the task information to be displayed to the front end is beneficial to show the relationship between the task and the node quickly, thereby facilitating the quick execution of the required operations on certain nodes.
[0102] Figure 8 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure. The method of this embodiment includes one or more features of the above-described data lineage processing method embodiments. In one possible implementation, the method further includes:
[0103] S801. Receive a data repair request, wherein the data repair request includes the identifier of the target node that needs to be repaired;
[0104] S802. Repair the data corresponding to the target node according to the identifier of the target node;
[0105] S803. Obtain the identifiers of upstream and downstream nodes that are related to the identifier of the target node;
[0106] S804. Based on the identifiers of the upstream and downstream nodes, repair the data corresponding to the upstream and downstream nodes that are related to the target node.
[0107] In this embodiment, some nodes may malfunction and require repair. The abnormal nodes can be quickly located using the colors displayed in the data lineage diagram. If an abnormal node needs to be repaired, the operator can select that node. This selection can be done in various ways. For example, the operator can directly select the abnormal node with the mouse. Alternatively, after selecting the abnormal node, a prompt window will appear asking if repair is needed, and the operator can select "repair needed" again in the prompt window. After the operator selects the abnormal node in the UI, the front-end can send a data repair request to the back-end, including the identifier of the abnormal node requiring repair. Using the identifier of the target node in the data repair request, the back-end can be requested to repair the target node. Furthermore, the back-end can also find data from other nodes related to the target node based on its identifier. Then, if any of these nodes are abnormal, they can also be repaired. For example, the downstream nodes of the target node can be found based on its identifier, and similar repairs can be performed on the downstream nodes.
[0108] In the above method, the timing of S802 and S803 is not restricted; they can be sequential or executed in parallel. For example, first find the identifiers of upstream and downstream nodes related to the target node's identifier, and then repair the target node. Alternatively, repair the target node first, and then find the identifiers of upstream and downstream nodes related to the target node's identifier. Yet another example is finding the identifiers of upstream and downstream nodes related to the target node's identifier first, and then simultaneously repairing both the target node and its related upstream and downstream nodes.
[0109] In one possible implementation, the method further includes:
[0110] S805. If the repair is successful, update the current state of the target node to obtain the updated state of the target node.
[0111] S806. Send the updated status of the target node to the front end.
[0112] In this embodiment, after the backend successfully repairs the data corresponding to a node, it can modify the node's state. For example, it can update an abnormal state to a normal state. Then, the backend can send the updated state of the node to the frontend. The updated state of the target node can update the data lineage diagram displayed on the frontend to show a more accurate node state. Furthermore, the backend can also send the updated states of other nodes that are related to the target node, thereby changing the color of these nodes displayed in the data lineage diagram on the frontend. For example, if the target node N2 changes from an abnormal state to a normal state, the color of the target node N2 in the data lineage diagram changes to green. Furthermore, if the downstream node N3 of the target node N2 also changes from an abnormal state to a normal state, N3 also changes to green.
[0113] Figure 9 This is a schematic flowchart of a data lineage processing method according to another embodiment of the present disclosure. The method of this embodiment includes one or more features of the above-described data lineage processing method embodiments. In one possible implementation, the method further includes:
[0114] S901. Receive a query request, wherein the query request includes the time to be queried and / or the identifier of the node to be queried;
[0115] S902. Based on the time to be queried and / or the identifier of the node to be queried, query the information of the upstream and downstream nodes of the data lineage to obtain the query result;
[0116] S903. Send the query result to the front end.
[0117] In this embodiment of the disclosure, controls supporting query functionality, such as input boxes and query buttons, can be displayed on the data lineage graph of the UI. Input boxes may include date input boxes and / or keyword input boxes. In the date input box, the user can enter the time to be queried, such as the date or time; in the keyword input box, the user can enter keywords to be queried, such as node identifiers. The query operation may include entering the information to be queried in the input box and / or clicking the query button. After the user performs a query operation on the UI, the front end may send a query request to the back end in response to the query operation. This query request may include the time to be queried and / or the identifier of the node to be queried entered by the user.
[0118] In this embodiment, the backend can save the data lineage diagram by date and / or time. For example, one data lineage diagram can be saved daily. Alternatively, N data lineage diagrams can be saved daily. Upon receiving a query request, if the query request includes a time (e.g., date), the backend can search for information on upstream and downstream nodes in the data lineage that match the operator's permissions for that date. If the query request also includes a node identifier, the backend can further search for information on upstream and downstream nodes related to that node within the upstream and downstream node information for that date. When searching based on node identifiers, a fuzzy matching method can be used to find more nodes related to that node's identifier.
[0119] The backend can return information about the upstream and downstream nodes found, i.e., the query results, to the frontend. The frontend can then generate a corresponding data lineage diagram based on the received information. The data lineage diagram generated by the frontend based on the query results may be the same as or different from the default data lineage diagram displayed.
[0120] In this embodiment of the disclosure, it is possible to support the query of data lineage, thereby obtaining the required data lineage relationship more quickly.
[0121] Figure 10 This is a schematic diagram of a data lineage processing apparatus according to an embodiment of the present disclosure. The apparatus includes:
[0122] The first receiving module 1001 is used to receive information of upstream and downstream nodes of data lineage that matches role permissions. The information of upstream and downstream nodes is obtained by performing data lineage analysis on the collected data.
[0123] The generation module 1002 is used to generate a data lineage diagram that matches the role permissions based on the information of the upstream and downstream nodes.
[0124] Figure 11 This is a schematic diagram of a data lineage processing apparatus according to another embodiment of the present disclosure. The method of this embodiment includes one or more features of the above-described data lineage processing apparatus embodiment. In one possible implementation, the apparatus further includes:
[0125] Display module 1101 is used to display the corresponding color in the data lineage diagram according to the state of the node, with different colors displayed for nodes in different states.
[0126] In one possible implementation, the state of the node includes at least one of the following: successful operation, failed operation, running, or waiting to run.
[0127] In one possible implementation, the data kinship graph comprises a directed acyclic graph.
[0128] In one possible implementation, the data lineage graph further includes task information, which includes at least one of the following:
[0129] Task time, node identifier, node path, SLA description, whether the target is met, and details.
[0130] In one possible implementation, the device further includes:
[0131] The first sending module 1102 is used to send a data repair request in response to a selection operation of a target node that needs to be repaired, wherein the data repair request includes the identifier of the target node that needs to be repaired.
[0132] In one possible implementation, the device further includes:
[0133] The second receiving module 1103 is used to receive the updated status of the target node if the repair is successful.
[0134] In one possible implementation, the device further includes:
[0135] The second sending module 1104 is used to send a query request in response to a query operation, wherein the query request includes the time to be queried and / or the identifier of the node to be queried;
[0136] The third receiving module 1105 is used to receive query results, which are obtained by querying the information of upstream and downstream nodes in the data lineage based on the time to be queried and / or the identifier of the node to be queried.
[0137] The generation module 1002 is also used to generate a data lineage diagram based on the query results.
[0138] In one possible implementation, the display module 1101 is also used to display a data lineage diagram generated based on the query results.
[0139] Figure 12 This is a schematic diagram of a data lineage processing apparatus according to an embodiment of the present disclosure. The apparatus includes:
[0140] The first sending module 1201 is used to send information about upstream and downstream nodes of data lineage that match the role permissions to the front end. The information about upstream and downstream nodes is used to generate a data lineage relationship diagram that matches the role permissions.
[0141] Figure 13This is a schematic diagram of a data lineage processing apparatus according to another embodiment of the present disclosure. The method of this embodiment includes one or more features of the above-described data lineage processing apparatus embodiment. In one possible implementation, the apparatus further includes a processing module 1301, which is used to perform data lineage analysis on the collected data to obtain the upstream and downstream relationships of the data lineage; and to generate information on the upstream and downstream nodes of the data lineage based on the upstream and downstream relationships of the data lineage.
[0142] In one possible implementation, the collected data includes task input information and / or output information, and the processing module 1301 is used to perform data lineage analysis on the input information and / or output information of each task to obtain the upstream and downstream relationships of each task.
[0143] In one possible implementation, the device further includes:
[0144] The acquisition module 1302 is used to acquire data from multiple data sources.
[0145] In one possible implementation, the information of the upstream and downstream nodes includes the status of the node, which includes at least one of the following: successful operation, failed operation, running, and waiting to run.
[0146] In one possible implementation, the device further includes:
[0147] The second sending module 1303 is used to send task information to be displayed to the front end, the task information including at least one of the following:
[0148] Task time, node identifier, node path, SLA description, whether the target is met, and details.
[0149] Figure 14 This is a schematic diagram of a data lineage processing apparatus according to another embodiment of the present disclosure. The method of this embodiment includes one or more features of the above-described data lineage processing apparatus embodiment. In one possible implementation, the apparatus further includes:
[0150] The first receiving module 1401 is used to receive a data repair request, wherein the data repair request includes the identifier of the target node that needs to be repaired;
[0151] The first repair module 1402 is used to repair the data corresponding to the target node according to the identifier of the target node;
[0152] The acquisition module 1403 is used to acquire the identifiers of upstream and downstream nodes that are related to the identifier of the target node.
[0153] The second repair module 1404 is used to repair the data corresponding to upstream and downstream nodes that are related to the target node based on the identifiers of the upstream and downstream nodes.
[0154] In one possible implementation, the device further includes:
[0155] The update module 1405 is used to update the current state of the target node when the repair is successful, so as to obtain the updated state of the target node.
[0156] The third sending module 1406 is used to send the updated status of the target node to the front end.
[0157] In one possible implementation, the device further includes:
[0158] The second receiving module 1407 is used to receive a query request, wherein the query request includes the time to be queried and / or the identifier of the node to be queried;
[0159] The query module 1408 is used to query the information of upstream and downstream nodes in the data lineage based on the time to be queried and / or the identifier of the node to be queried, and obtain the query result.
[0160] The fourth sending module 1409 is used to send the query result to the front end.
[0161] For a description of the specific functions and examples of each module and submodule of the data lineage processing apparatus in this disclosure, please refer to the relevant descriptions of the corresponding steps in the embodiments of the above-described data lineage processing method, which will not be repeated here.
[0162] In related technologies, the relationships between data in a data warehouse rely on manual communication. Data output requires manual verification, and backtracking involves lengthy investigations and feedback to downstream systems. Repairing data anomalies is extremely difficult, time-consuming, and labor-intensive. Data accuracy is compromised after output, with users discovering anomalies, reporting issues, and then investigating and locating the cause – a time-consuming and reactive process. This results in poor usability, low accuracy, poor readability, high resource consumption, and low work efficiency. The performance of a data warehouse can be assessed by key metrics such as usability, accuracy, and readability.
[0163] Based on the embodiments disclosed herein, an automated and visualized system for data lineage collection, monitoring, and backtracking can be provided. This system enables personalized viewing, clearly showing data lineage relationships, and automatically monitoring whether data has been generated and whether the generated data is correct. If backtracking is needed, clicking a trigger button can automatically backtrack downstream data. This can significantly reduce operation and maintenance costs and improve efficiency.
[0164] like Figure 15 As shown, the system architecture of this embodiment mainly includes: multiple data sources, a lineage analysis module, a data lineage database, and a data lineage platform. For example, the big data platform may include multiple data sources. Data for data lineage analysis can be collected using the big data platform. For example, data can be collected from the big data platform using MR services, Hive services, QE services, Turing clients, Turing pages, Spark clients, LS page settings, etc.
[0165] For example, data can be collected from MR using the MR service. Data can be collected from Hive using the Hive service. Data can be collected from QE using the QE service. Data can be collected from the Turing DB (database) using the Turing client or the Turing SQL capabilities within the Turing page. The Spark client can run an SQL collection service. This service can include SQL processed through Spark session extensions, the parser interface, and the BjhSqlparser to obtain Spark SQL and the Result API, which then updates a database such as the Baina database. The SQL collection service can also include RDDs, which are analyzed and then updated in the database. Operations can also be performed by updating the UDB (Universal Database) through heterogeneous services. The LS page settings can collect data through the API (Application Programming Interface) service.
[0166] The collected data can be analyzed using lineage analysis methods, such as the data lineage processing method described in any of the above embodiments. The obtained data lineage can be saved to a data lineage database (DB). The data lineage platform can be used to perform operations such as problem investigation, data analysis, data monitoring, authentication, and custom tree creation. For example, a custom tree can reflect information about the upstream and downstream nodes of the data lineage.
[0167] Based on this architecture, such as Figure 16 As shown, the data lineage analysis process of this disclosure embodiment may include:
[0168] S1601: Data Input Acquisition: This system can collect the inputs and outputs of all relevant tasks that users, such as operators in various roles, are concerned with, and provide the results to the lineage analysis system. This significantly improves the data acquisition rate.
[0169] S1602: Lineage Analysis: The data lineage analysis system can determine upstream and downstream relationships based on the input and output logic provided by each acquisition system, and generate upstream and downstream lineage nodes. This embodiment of the disclosure can realize the automatic analysis function of data lineage.
[0170] S1603: Front-end Display: Display a personalized DAG (Directed Acyclic Graph) on the front-end based on upstream and downstream nodes. For example, a personalized DAG can be implemented based on the data table permissions of different roles. Permissions for each data table can be set, and users with different roles will find or display different data lineage diagrams by default, highlighting their respective key points.
[0171] In addition, such as Figure 17 As shown, the XXX data lineage diagram, generated according to data lineage relationships, includes nodes corresponding to multiple data tables. Each data table has certain permissions, which can be displayed as "owner: **". Role permissions can also be hidden. A data table may have multiple roles with permissions, for example, the second-level derived table 13. The next level of the root node table points to the first-level derived tables 1, 2, 3, and 4. The first-level derived table 1 points to the second-level derived tables 5 and 6. The first-level derived table 2 points to the second-level derived tables 7 and 8. The first-level derived table 3 points to the second-level derived tables 9 and 10. The first-level derived table 4 points to the second-level derived tables 11, 12, and 13.
[0172] The data lineage diagram can display different colors for each node based on its business status. For example, nodes that have successfully run, failed, are currently running, and are waiting to run can be displayed in different colors. Furthermore, for nodes that have failed, information related to the cause of the failure, such as the delay reason and delay duration, can be displayed. Alternatively, information related to the cause of the failure can be hidden and revealed only after the user clicks or selects the node. Once an anomaly occurs and a user triggers an anomaly point, automatic repair can be performed. A one-click backtracking function can be set up, allowing the repair of an anomaly node to simultaneously modify other child nodes of that node, achieving automatic data monitoring and backtracking, and automatically completing the repair of data nodes. For example, selecting one-click backtracking on the first-level derived table 2 can repair both the first-level derived table 2 and the second-level derived tables 7 and 8. If the repair is successful, the colors of the first-level derived table 2, second-level derived table 7, and second-level derived table 8 can be changed, for example, to the color of the waiting-to-run node.
[0173] The data lineage diagram in this embodiment uses different colors to represent different node states, making it easier for users to quickly view and maintain the data.
[0174] This disclosure also provides date search and keyword search functions, which can more accurately locate relevant nodes.
[0175] Further, see Figure 17 The display interface also shows information about different nodes through an information table, including: task time, node name, node path, SLA description, whether the target is met, and details.
[0176] In this embodiment, the automatically generated data lineage diagram enables quick and convenient automatic backtracking, reducing the likelihood of data omissions or gaps and improving system data accuracy. Furthermore, one-click backtracking facilitates problem troubleshooting, enhancing system usability. Additionally, different colors highlighting node states improve system readability. Moreover, different colors emphasize nodes of interest, such as abnormal nodes, making monitoring and backtracking more convenient, minimizing data delays, and enabling rapid correction of abnormal data, thus improving system stability.
[0177] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0178] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0179] Figure 18 A schematic block diagram of an example electronic device 1800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0180] like Figure 18As shown, device 1800 includes a computing unit 1801, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1802 or a computer program loaded from storage unit 1808 into random access memory (RAM) 1803. The RAM 1803 may also store various programs and data required for the operation of device 1800. The computing unit 1801, ROM 1802, and RAM 1803 are interconnected via bus 1804. Input / output (I / O) interface 1805 is also connected to bus 1804.
[0181] Multiple components in device 1800 are connected to I / O interface 1805, including: input unit 1806, such as keyboard, mouse, etc.; output unit 1807, such as various types of monitors, speakers, etc.; storage unit 1808, such as disk, optical disk, etc.; and communication unit 1809, such as network card, modem, wireless transceiver, etc. Communication unit 1809 allows device 1800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0182] The computing unit 1801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1801 performs the various methods and processes described above, such as data lineage processing methods. For example, in some embodiments, the data lineage processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1800 via ROM 1802 and / or communication unit 1809. When the computer program is loaded into RAM 1803 and executed by the computing unit 1801, one or more steps of the data lineage processing method described above may be performed. Alternatively, in other embodiments, the computing unit 1801 may be configured to perform a data lineage processing method by any other suitable means (e.g., by means of firmware).
[0183] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0184] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0185] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0186] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0187] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0188] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0189] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0190] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A data lineage processing method for the front end of a data warehouse, the method comprising: The front end receives information about upstream and downstream nodes of the data lineage that matches the role and permissions sent by the back end of the data warehouse. The information about upstream and downstream nodes is generated based on the upstream and downstream relationships obtained by the back end through data lineage analysis of the collected data. The front end generates a data lineage diagram that matches the role permissions based on the information of the upstream and downstream nodes; the data lineage diagram is used to display the hierarchical relationship and connection relationship of the upstream and downstream nodes on the user interface of the front end, wherein nodes with upstream and downstream relationships have connecting edges; The user interface on the front end displays the corresponding color in the data lineage diagram according to the state of the node, with different colors displayed for nodes in different states; In response to the user interface selection operation of the target node that needs to be repaired displayed in the data lineage diagram on the front end, a data repair request is sent to the back end, the data repair request including the identifier of the target node; the data repair request is used to request the back end to repair the target node and the upstream and downstream nodes that are related to the target node. If the backend repair is successful, the frontend receives the updated status after the backend repairs the data corresponding to the target node and its upstream and downstream nodes that are related to the target node. The updated status is obtained by the backend obtaining the identifiers of the upstream and downstream nodes that are related to the target node based on the identifier of the target node, and updating the current status of the target node and its upstream and downstream nodes after successfully repairing the data corresponding to the target node and its upstream and downstream nodes. Based on the updated status of the target node and other nodes related to the target node, the user interface on the front end changes the color displayed by the updated node in the data lineage diagram.
2. The method according to claim 1, wherein, The node's state includes at least one of the following: successful execution, failed execution, running, or waiting to run.
3. The method according to claim 1, wherein, The data kinship diagram includes a directed acyclic graph.
4. The method according to any one of claims 1 to 3, wherein, The data lineage diagram also includes task information, which includes at least one of the following: Task time, node identifier, node path, Service Level Agreement (SLA) description, whether it meets the standard, and details.
5. The method according to any one of claims 1 to 3, further comprising: The front end responds to the query operation by sending a query request to the back end, the query request including the time to be queried and / or the identifier of the node to be queried; The front end receives query results from the back end. The query results are obtained by querying the information of upstream and downstream nodes in the data lineage based on the time to be queried and / or the identifier of the node to be queried. The front end generates a data lineage diagram based on the query results.
6. A data lineage processing method for the backend of a data warehouse, the method comprising: The backend performs data lineage analysis on the collected data to obtain upstream and downstream relationships, and generates information on upstream and downstream nodes. The backend sends information about upstream and downstream nodes of the data lineage that matches the role permissions to the frontend of the data warehouse. The information about upstream and downstream nodes is used to generate a data lineage relationship graph that matches the role permissions. The data lineage relationship graph is used to display the hierarchical relationship and connection relationship of the upstream and downstream nodes on the user interface of the frontend, wherein nodes with upstream and downstream relationships have connecting edges. The backend receives a data repair request sent by the frontend, and the data repair request includes the identifier of the target node that needs to be repaired as shown in the data lineage diagram; The backend repairs the data corresponding to the target node based on the target node's identifier; The backend obtains the identifiers of upstream and downstream nodes that are related to the identifier of the target node. The backend repairs the data corresponding to upstream and downstream nodes that are related to the target node based on the identifiers of the upstream and downstream nodes. If the target node is successfully repaired, the backend updates the current state of the target node to obtain the updated state of the target node. The updated state is obtained by the backend obtaining the identifiers of upstream and downstream nodes that are related to the target node based on the identifier of the target node, and updating the current state of the target node and the upstream and downstream nodes after successfully repairing the data corresponding to the target node and the upstream and downstream nodes. If the upstream and downstream nodes of the target node are successfully repaired, the backend updates the upstream and downstream nodes to obtain the update status of the upstream and downstream nodes. The backend sends the updated status of the target node and its upstream and downstream nodes to the frontend so that the updated node can be displayed in the data lineage graph in the user interface of the frontend.
7. The method according to claim 6, further comprising: Data lineage analysis is performed on the collected data to obtain the upstream and downstream relationships of the data lineage; Based on the upstream and downstream relationships of the data lineage, information about the upstream and downstream nodes of the data lineage is generated.
8. The method according to claim 7, wherein, The collected data includes task input information and / or output information. Data lineage analysis is performed on the data to obtain upstream and downstream relationships within the data lineage, including: Data lineage analysis is performed on the input and / or output information of each task to obtain the upstream and downstream relationships of each task.
9. The method according to any one of claims 6 to 8, further comprising: Collect data from multiple data sources.
10. The method according to any one of claims 6 to 8, wherein, The information of the upstream and downstream nodes includes the status of the nodes, and the status of the nodes includes at least one of the following: successful operation, failed operation, running, and waiting to run.
11. The method according to any one of claims 6 to 8, further comprising: Send task information to be displayed to the front end, the task information including at least one of the following: Task time, node identifier, node path, SLA description, whether the target is met, and details.
12. The method according to any one of claims 6 to 8, further comprising: Receive a query request from the front end, the query request including the time to be queried and / or the identifier of the node to be queried; Based on the time to be queried and / or the identifier of the node to be queried, the query is performed on the information of the upstream and downstream nodes of the data lineage to obtain the query results; The query results are sent to the front end.
13. A data lineage processing apparatus for the front end of a data warehouse, the apparatus comprising: The first receiving module is used to receive information on upstream and downstream nodes of data lineage that matches role permissions from the backend of the data warehouse at the frontend. The information on upstream and downstream nodes is generated based on the upstream and downstream relationships obtained by the backend through data lineage analysis of the collected data. A generation module is used to generate a data lineage diagram matching the role permissions on the front end based on the information of the upstream and downstream nodes; the data lineage diagram is used to display the hierarchical relationship and connection relationship of the upstream and downstream nodes on the user interface of the front end, including the hierarchical relationship and connection relationship of the nodes, wherein nodes with upstream and downstream relationships have connecting edges. The display module is used to display the corresponding color in the data lineage diagram on the front-end user interface according to the state of the node, with different colors displayed for nodes in different states; The first sending module is further configured to send a data repair request to the backend in response to a selection operation on the user interface of the frontend for a target node that needs to be repaired displayed in the data lineage diagram, so as to request the backend to repair the target node. The data repair request includes an identifier of the target node that needs to be repaired. The data repair request is used to request the backend to repair the target node and the upstream and downstream nodes that have a lineage relationship with the target node. The second receiving module is configured to receive, at the front end, an updated status after the back end repairs the data corresponding to the target node and its upstream and downstream nodes that are related to the target node, when the back end repairs the data successfully. The updated status is obtained by the back end obtaining the identifiers of the upstream and downstream nodes that are related to the target node based on the identifier of the target node, and updating the current status of the target node and its upstream and downstream nodes after the back end successfully repairs the data corresponding to the target node and its upstream and downstream nodes. The display module is also used to change the color of the updated node displayed in the data lineage diagram on the front-end user interface according to the update status of the target node and other nodes that are related to the target node.
14. The apparatus according to claim 13, wherein, The node's state includes at least one of the following: successful execution, failed execution, running, or waiting to run.
15. The apparatus according to claim 13, wherein, The data kinship diagram includes a directed acyclic graph.
16. The apparatus according to any one of claims 13 to 15, wherein, The data lineage diagram also includes task information, which includes at least one of the following: Task time, node identifier, node path, SLA description, whether the target is met, and details.
17. The apparatus according to any one of claims 13 to 15, further comprising: The second sending module is used to send a query request in response to a query operation, wherein the query request includes the time to be queried and / or the identifier of the node to be queried; The third receiving module is used to receive query results, which are obtained by querying the information of upstream and downstream nodes in the data lineage based on the time to be queried and / or the identifier of the node to be queried. The generation module is also used to generate a data lineage diagram based on the query results.
18. A data lineage processing apparatus for the backend of a data warehouse, the apparatus comprising: The lineage analysis module is used to perform lineage analysis on the collected data at the back end to obtain the upstream and downstream relationships and generate information on upstream and downstream nodes. The first sending module is used to send information about upstream and downstream nodes of data lineage matching role permissions from the backend to the frontend of the data warehouse. The information about upstream and downstream nodes is used to generate a data lineage relationship diagram matching the role permissions. The data lineage relationship diagram is used to display the hierarchical relationship and connection relationship of the upstream and downstream nodes on the user interface of the frontend, wherein nodes with upstream and downstream relationships have connecting edges. The first receiving module is configured to receive a data repair request sent by the front end at the back end, wherein the data repair request includes the identifier of the target node that needs to be repaired as shown in the data lineage diagram; The first repair module is used to repair the data corresponding to the target node in the backend according to the identifier of the target node; The acquisition module is used to acquire the identifiers of upstream and downstream nodes that are related to the identifier of the target node in the backend. The second repair module is used to repair the data corresponding to upstream and downstream nodes that are related to the target node in the backend according to the identifiers of the upstream and downstream nodes. An update module is used to update the current state of the target node in the backend when the target node is successfully repaired, thereby obtaining an updated state of the target node; and to update the upstream and downstream nodes in the backend when the upstream and downstream nodes of the target node are successfully repaired, thereby obtaining an updated state of the upstream and downstream nodes; the updated state is obtained by the backend obtaining the identifiers of the upstream and downstream nodes that are related to the target node based on the identifier of the target node, and updating the current state of the target node and the upstream and downstream nodes after successfully repairing the data corresponding to the target node and the upstream and downstream nodes. The third sending module is used to send the updated status of the target node and its upstream and downstream nodes from the backend to the frontend, so as to change the color of the updated node displayed in the data lineage diagram on the user interface of the frontend. The second receiving module is used to receive query requests, which include the time to be queried and / or the identifier of the node to be queried.
19. The apparatus according to claim 18 further includes a processing module, the processing module being used to perform data lineage analysis on the collected data to obtain the upstream and downstream relationships of the data lineage; and to generate information on the upstream and downstream nodes of the data lineage based on the upstream and downstream relationships of the data lineage.
20. The apparatus according to claim 19, wherein, The collected data includes task input information and / or output information. The processing module is used to perform data lineage analysis on the input information and / or output information of each task to obtain the upstream and downstream relationships of each task.
21. The apparatus according to any one of claims 18 to 20, further comprising: The data acquisition module is used to acquire data from multiple data sources.
22. The apparatus according to any one of claims 18 to 20, wherein, The information of the upstream and downstream nodes includes the status of the nodes, and the status of the nodes includes at least one of the following: successful operation, failed operation, running, and waiting to run.
23. The apparatus according to any one of claims 18 to 20, further comprising: The second sending module is used to send task information to be displayed to the front end, the task information including at least one of the following: Task time, node identifier, node path, SLA description, whether the target is met, and details.
24. The apparatus according to any one of claims 18 to 20, further comprising: The second receiving module is used to receive a query request, which includes the time to be queried and / or the identifier of the node to be queried. The query module is used to query the information of upstream and downstream nodes in the data lineage based on the time to be queried and / or the identifier of the node to be queried, and obtain the query results. The fourth sending module is used to send the query results to the front end.
25. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.
26. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-12.
27. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-12.