Method and apparatus for managing data lineage, and device and medium
By creating a unified data lineage model, the challenges of managing different types of data throughout their lifecycle are solved, achieving connectivity and visibility of data throughout its entire lifecycle, improving the efficiency and security of data management, and ensuring data quality and consistency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-11-01
- Publication Date
- 2026-05-07
AI Technical Summary
Existing data lineage management technologies are ineffective and struggle to manage the flow and transformation of different types of data throughout their lifecycles, leading to problems such as invisible data flow, uncontrollable risks, and unknowable impacts of data updates.
By creating a data lineage model, multiple service nodes and edge relationships can be managed in a unified manner. A single model is used to describe different types of lineages, including client, online and offline lineages. Relational, graph and key-value databases are used to store node and edge information, realizing the connectivity and interoperability of multiple types of lineages.
It enables connectivity and visibility of data throughout its entire lifecycle, improves the efficiency and security of data management, ensures data quality and consistency, and allows for better tracking of data sources and changes.
Smart Images

Figure CN2024129450_07052026_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices, and media for managing data lineage Technical Field
[0001] Exemplary implementations of this disclosure generally relate to data management, and more particularly to methods, apparatus, devices, and computer-readable storage media for managing data lineage. Background Technology
[0002] In a computer environment, various types of data can flow between different locations and / or processing processes to achieve the intended task. Data lineage refers to the relationships between data throughout its lifecycle, including its source, flow, and transformation. Data lineage records the data's propagation path and processing procedures from acquisition, transmission, storage, and analysis. Through data lineage, we can better understand the source, history, and reliability of data, and process data more securely and efficiently. However, the performance of existing data lineage management technologies is unsatisfactory, thus there is a desire to improve the efficiency of data lineage management.
[0003] Summary of the Invention
[0004] In a first aspect of this disclosure, a method for managing data lineage is provided. In this method, in response to receiving a data lineage, multiple nodes are determined based on multiple services associated with the data lineage. A first node among the multiple nodes is created based on a first service among the multiple services. A first type of the first node includes a set of dimensions, each dimension corresponding to a set of entities in a set of hierarchical levels of the service. Based on the relationships between the multiple services, an edge between the first node and a second node among the multiple nodes is determined. Based on the multiple nodes and the edge, a data lineage model describing the data lineage is created.
[0005] In a second aspect of this disclosure, an apparatus for managing data lineage is provided. The apparatus includes: a node determination module configured to, in response to receiving a data lineage, determine multiple nodes based on multiple services associated with the data lineage, wherein a first node among the multiple nodes is created based on a first service among the multiple services, and a first type of the first node includes a set of dimensions, each set of dimensions corresponding to a set of entities in a set of hierarchical levels of the services; an edge determination module configured to determine an edge between the first node and a second node among the multiple nodes based on the association relationships between the multiple services; and a creation module configured to create a data lineage model describing the data lineage based on the multiple nodes and the edge.
[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.
[0008] In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] In the following detailed description, the above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent, taken in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1A shows a block diagram of data flow according to some implementations of this disclosure;
[0012] Figure 1B shows a block diagram of data lineage according to some implementations of this disclosure;
[0013] Figure 2 shows a block diagram for managing data lineage according to some implementations of this disclosure;
[0014] Figure 3 shows a block diagram of a data lineage model according to some implementations of this disclosure;
[0015] Figure 4 shows a block diagram of a data lineage model for storing data according to some implementations of this disclosure;
[0016] Figure 5 shows a block diagram of a data lineage model according to some implementations of this disclosure;
[0017] Figure 6 shows a block diagram of multiple sides in a data lineage model according to some implementations of this disclosure;
[0018] Figure 7 shows a block diagram of a data lineage model according to some implementations of this disclosure;
[0019] Figure 8 shows a block diagram of a data lineage model according to some implementations of this disclosure;
[0020] Figure 9 shows a flowchart of a method for managing data lineage according to some implementations of this disclosure;
[0021] Figure 10 shows a block diagram of an apparatus for managing data lineage according to some implementations of this disclosure; and
[0022] Figure 11 shows a block diagram of a device capable of implementing various implementations of the present disclosure. Detailed Implementation
[0023] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0024] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.
[0025] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0026] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0027] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0028] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0029] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0030] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.
[0031] Example Environment
[0032] Data lineage refers to the relationships of data throughout its lifecycle, including its source, flow, and transformation. Data lineage records the data propagation path and processing procedures from acquisition, transmission, storage, and analysis. Through data lineage, we can better understand the source, history, and reliability of data, and process data more securely and efficiently. See Figure 1A for a description of the data lifecycle; Figure 1A shows a block diagram 100A of data flow according to some implementations of this disclosure. Legend 110, 112, and 114 illustrate different types of lineage nodes, and arrow 120 indicates the lineage relationship, i.e., the data flow relationship.
[0033] Figure 1A (left) illustrates the client-side phase, which includes the data collection, storage, and transmission process on the device. This phase may involve security information, such as sensitive data like phone numbers, and is typically a single line of data. The lineage of this phase can be referred to as the client-side lineage or the mobile-side lineage. Figure 1A (middle) illustrates the online phase, encompassing the entire computation and transmission process from the gateway to the online database (MySQL / Redis / ...). The security information involved is typically a single line of data. The lineage of this phase can be referred to as the online-side lineage. Figure 1A (right) illustrates the offline phase, including the process of data moving from the online database to the offline database and performing various computational tasks between the offline databases. The security information involved is typically batch data. The lineage of this phase is called the offline-side lineage.
[0034] As shown in Figure 1A, the complete lifecycle of data can be observed: from end nodes (such as apps, web pages, etc.) through gateway nodes to the backend online service, then it is stored in a database (such as MySQL, possibly with intermediate caching such as Redis), and the data stored in the database can be queried by the backend online service, and so on. Furthermore, the data stored in the database can be synchronized to offline databases (such as Hive) through certain synchronization nodes for data analysis (or used as training data, etc.), and finally presented to users or fed back to the online service to improve processing capabilities.
[0035] Data lineage has multiple applications. First, it plays a crucial role in data management (e.g., security management), helping organizations understand the source and destination of data, trace its history, ensure quality and reliability, and improve efficiency and effectiveness. Data management may face challenges such as invisible data flow, uncontrollable risks, and unknown impacts of data updates. Data lineage management can address these issues. Through data lineage, users can identify the source and evolution of data to ensure quality and consistency. Furthermore, regarding data security requirements, data lineage can assist in identifying the dissemination of secure data and tracking data usage and processing.
[0036] Technical solutions for representing data lineage have been proposed. Different types of lineages focus on different node types and dimensions, and thus describe different lineage relationships. First, they focus on different node types. Online lineages focus on online services and online storage, such as RPC services and MySQL storage; offline lineages focus on offline storage, such as Hive. Second, they focus on different dimensions. Client-side lineage can be viewed as the interaction between various service modules within an app, while online lineage can be viewed as the call relationships between various services and the read / write relationships between services and storage. Online services / storage themselves include multiple modules, and these modules interact with each other. Mobile lineage can be considered a special service node. Mobile lineage focuses on the lineage flow within a service, while online lineage focuses on the lineage flow between services; therefore, the dimensions they focus on are different. Third, they represent different relationships. Online lineages represent the call relationships between services and the read / write relationships between services and storage. Offline lineages represent the transformation relationships between tables / fields.
[0037] In addition, online lineage differs from other lineages in its specific presentation. Typically, data lineage is described in the form of a graph, which can be called a lineage graph. The points in the graph are the nodes that the lineage focuses on, and the edges represent the relationships between two nodes, i.e., the lineage relationships. Online lineage focuses on the actual call relationships of traffic and has the concept of a master node. One node corresponds to one lineage graph, and this node is the master node in that lineage graph. See Figure 1B for more details, which shows a block diagram 100B of data lineage according to some implementations of this disclosure. As shown in Figure 1B, node A can represent a service, with multiple services connected upstream and downstream. There can be four call paths: ① to ④. However, some paths do not actually exist.
[0038] In Figure 1B, legend 140 represents a regular node, and legend 142 represents a master node. The path BAD at the bottom of Figure 1B shows the lineage with node A as the master node. All four edges (①, ②, ③, ④) in the graph have actual traffic (path: ①→③, ②→④), meaning that for a master node in an online lineage graph, all other nodes have actual traffic relationships with the master node. Path CAE shows the lineage with node C as the master node. For master node C, there is a link ②→④, and its lineage graph will not contain B and D. Similarly, for master node B, there is a link ①→③, and its lineage graph will not contain C and E.
[0039] For online lineages, the number of nodes corresponds one-to-one with the number of lineage graphs; that is, the number of lineage graphs equals the number of nodes. Searching for any node X in an online lineage might not result in a node Y existing in X's lineage graph; this type of lineage can be classified as asymmetric. Conversely, other lineages have no traffic flow relationships and therefore no concept of a master node; a single lineage graph can encompass the lineages of all nodes. This type of lineage is symmetric. Although lineage graphs can represent data flow, it is necessary to store multiple lineage graphs with each node as a master node.
[0040] Given the aforementioned shortcomings, the aim is to integrate different types of lineages to achieve data connectivity throughout its entire lifecycle. Specifically, the goal is to define a unified data model that includes the node types, dimensions, and relationships represented by different lineages. It is also to provide interoperability between various data lineages to connect them and achieve connectivity across multiple lineage types. Furthermore, the aim is to enable the coexistence of symmetrical and asymmetrical lineages, thereby representing the lineage relationships of both characteristics and how to store and query them.
[0041] Overview of managing data lineage
[0042] To at least partially address the shortcomings of the prior art, a method for managing data lineage is proposed according to an exemplary implementation of this disclosure. In summary, a data lineage model can be provided in which nodes can have multiple hierarchical types. Thus, without maintaining multiple lineage graphs, a single model can be used to describe richer information, thereby describing data lineage in a more accurate manner.
[0043] Referring to Figure 2, which describes an outline of an exemplary implementation of the present disclosure, Figure 2 shows a block diagram 200 for managing data lineage according to some implementations of the present disclosure. As shown in Figure 2, in response to receiving a data lineage 210, a data lineage model 240 can be determined. This model may include multiple nodes and edges between the multiple nodes. For ease of description, only a first node (e.g., node 210) and a second node (e.g., node 220) and edge 230 among the multiple nodes are described as examples. The multiple nodes can be determined separately based on multiple services associated with the data lineage, wherein the first node among the multiple nodes is created based on a first service among the multiple services.
[0044] As shown in Figure 2, the first type of the first node (e.g., type 212) includes a set of dimensions, each corresponding to a set of entities in a set of levels of the service. Specifically, the type here can be represented by entities in multiple levels of the service. For example, for an online service, this node can include multiple levels (i.e., multiple dimensions): service, service method (service.method), and method field (service.method.field). Nodes can be determined in a similar manner; for example, node 210 corresponding to service 1 and node 220 corresponding to service 2 can be identified.
[0045] Furthermore, based on the relationships between multiple services, the edge between the first and second nodes among multiple nodes can be determined. As shown in Figure 2, assuming service 1 calls service 2, a directed edge (edge 230) can be added between node 210 and node 220. Furthermore, based on multiple nodes and edges, a data lineage model 240 describing data lineage can be created.
[0046] Using the exemplary implementation of this disclosure, different types of lineage can be represented in a unified manner, achieving data connectivity throughout its entire lifecycle. In this way, the node types, dimensions, and relationships involved in different lineages can be defined.
[0047] Detailed process of managing data lineage
[0048] Having outlined some implementations of this disclosure, further details regarding methods for managing data lineage will be described below. According to some implementations of this disclosure, to integrate different types of lineages, node types can cover different types within all lineages. Furthermore, nodes of the same type can include information at different levels (dimensions). For example, an online service node can include three dimensions: "service," "service method (or interface)," and "(method) fields." For offline storage Hive, it includes three dimensions: "the service storing itself," "table," and "fields."
[0049] According to some implementations of this disclosure, not all types of nodes have three dimensions. For example, message middleware such as Kafka / RocketMQ only have two layers: "its own service" (corresponding to the topic) and "fields". Alternatively and / or additionally, there are nodes with only one dimension (e.g., a cron job script service) and nodes with more than three dimensions, which can be determined based on the entities to be focused on when constructing the data lineage. Specifically, the entities to be focused on in a particular business can be determined, and thus the corresponding dimensions can be determined. For example, for a certain business, if it only focuses on the service method dimension, only the "method" level can be defined. If it needs to focus on both services and methods, two levels can be defined: "service" and "service.method". If it needs to focus on services, methods, and fields, three levels can be defined: "service", "service.method", and "service.method.field". Table 1 below shows the definitions of node types according to some implementations of this disclosure:
[0050] Table 1 Node Types
[0051] It should be understood that for an HTTP service, a specific HTTP method (e.g., GET, POST) and path are typically required to map to the specific service's handler method. Table 1 only illustratively illustrates examples of node types; alternatively and / or additionally, there may be more, fewer, or different dimensions, and the specific content of each dimension may vary depending on the specific business requirements. Before constructing the data lineage model, it is necessary to clearly define all types of nodes to be focused on and the necessary hierarchical dimensions within each type of node. It is possible to define some types of nodes first, and then add new types of nodes as needed during the subsequent construction process.
[0052] According to some implementations of this disclosure, the type of a node can be described using "type" (string). Corresponding to the node definitions in Table 1 above, for an RPC service, the type can include: rpc.service, rpc.method, and rpc.field.
[0053] According to some implementations of this disclosure, the first node includes: node attributes, which include at least one attribute corresponding to the first type; and an identifier, which is determined by the first type and the node attributes. In this way, richer attributes can be defined for the node, thus facilitating the unified representation of multiple lineage types within a single data lineage model.
[0054] Specifically, attributes can be represented using "base_props" (mapped [string]string). Here, attributes can include one or more base attributes (i.e., a set of base attributes) related to the node. Specifically, if type = rpc.service, the key in the mapping can be represented as service; if type = rpc.method, the keys can be represented as service and method; if type = rpc.field, the keys can be represented as service, method, and field. The attribute identification method is similar for other node types.
[0055] According to some implementations of this disclosure, if it is necessary to focus on the regional information involved in each service, Internet Data Centers (IDCs) or regions can be used to distinguish service nodes in different regions. In this case, the attribute can further include keywords represented by IDC or region. Taking region as an example, if type = rpc.service, the keywords in the mapping can include service and region; if type = rpc.method, the keywords in the mapping can include service, method, and region; if type = rpc.field, the keywords in the mapping can include service, method, field, and region. Table 2 shows an example of the attribute data structure:
[0056] Example of data structure for attributes in Table 2
[0057] According to some implementations of this disclosure, a node's identifier can be uniquely determined based on its type and attributes. Specifically, the concatenated string can be determined using type + base_props, and this string can be used as the node's identifier. For example, if type = rpc.method, then the node's identifier is unique_key = $type + $service + $method + ($region).
[0058] It should be understood that during data flow, multiple services may have the same name. For example, `rpc.service` and `mysql.service` can have the same name, making it impossible to distinguish them solely by `base_props`. In this case, a unique identifier for the node can be determined by a combination of `type` and `base_props`. Alternatively and / or additionally, if different types of nodes have different names, the node identifier can be determined based on `base_props`.
[0059] Alternative and / or additional locations require matching multiple fields in the database when determining node equality. To improve processing efficiency, a digest identifier field can be further provided, for example, based on various digest algorithms (e.g., MD5, etc.) to determine the corresponding digest identifier. It should be understood that the concatenation order of each member during identifier determination should be consistent. In this way, it can be ensured that the determined identifier and the process of further determining the digest identifier are consistent.
[0060] According to some implementations of this disclosure, the first node further includes: a lineage type, which represents the stage of the first service in the data lineage. The lineage type includes at least one of the following: client stage, online stage, or offline stage. Using some implementations of this disclosure, data lineages spanning different lifecycles can be represented in a unified way. Specifically, the lineage type can be represented using the `lineage_type(int32, 32-bit integer)` field. One or more bits can be used to represent a type; for example, the first bit can represent the client lineage type, the second bit can represent the online lineage type, the third bit can represent the offline lineage type, and so on.
[0061] In one example, if the current node corresponds only to the client lineage, lineage_type = 1 (001) can be defined. If the current node corresponds to both online and offline lineages, lineage_type = 6 (110) can be defined. If the current node corresponds to all three lineages, lineage_type = 7 (111) can be defined. Alternatively and / or additionally, this representation can be extended; for example, a fourth bit could be used to represent the lineage type related to machine learning, etc. The remaining bits can be reserved for future expansion.
[0062] According to some implementations of this disclosure, the first node further includes at least one of the following: a node extension field, which represents an extension attribute associated with the first node, the extension attribute being labeled using a lineage type; a node name, which represents the name of the first node being presented; an update time, which represents the time when the first node was updated; and a creation time, which represents the time when the first node was created. Using some implementations of this disclosure, the data lineage model can include richer information, thereby facilitating interoperability between multiple data lineages to connect various lineages and achieve connectivity between multiple types of lineages.
[0063] According to some implementation methods of this disclosure, a node can belong to multiple data lineages. Each lineage may require extended attributes for the node. In this case, additional information can be stored using node extended fields. For example, in addition to basic attributes, this field can store basic information such as the node's owner, the business to which the node belongs, etc. To prevent keyword conflicts between different lineages, each lineage's keywords can be specified with a prefix, that is, the extended attributes can be marked using the lineage type. Specifically, the extended attributes related to the client lineage can be represented as: cl_xxx; the extended attributes related to the online lineage can be represented as: ol_xxx; and the extended attributes related to the offline lineage can be represented as: of_xxx.
[0064] According to some implementations of this disclosure, if data lineage needs to be displayed to users at the platform layer, using summary identifiers or identifiers that are not user-friendly is problematic. In this case, a name field can be added to facilitate application display at the platform layer. Users can specify the name, font, shape, etc., of the nodes to be presented. This allows users to easily understand the services represented by each node, thus providing a clearer understanding of the data flow process. Alternatively and / or additionally, a unified clock can be used to represent update and creation times. Having described the detailed information of nodes, according to some implementations of this disclosure, nodes can be represented using the data structure shown in Table 3 below.
[0065] Table 3 shows an example of the data structure for the nodes.
[0066] According to some implementations of this disclosure, an edge includes: a start node, which represents the starting point of the edge; an end node, which represents the ending point of the edge; a master node identifier, which represents the master node to which the edge belongs; and an edge type, which represents the association relationship between the first node and the second node. Using some implementations of this disclosure, various types of lineage-related information can be stored in the edge, thereby storing richer information in a single data lineage model.
[0067] According to some implementations of this disclosure, the start node (e.g., using a "from" (string) field) can store the identifier of the start node of the edge, and the end node (e.g., using a "to" (string) field) can store the identifier of the end node of the edge. According to some implementations of this disclosure, the master node identifier can represent the master node to which the edge belongs, for example, stored using a "master_node_id" (string) field. Assuming the master node is A, then each edge in the lineage graph with A as the master node can be labeled as A.
[0068] According to some implementations of this disclosure, the edge type can be determined based on the association between two nodes, and the edge type is stored using a "type" (string) field. For example, "invoke" can be used to represent a call relationship. When RPC node A calls an interface of RPC node B, this relationship is "invoke". As another example, "read / write" can be used to represent a read / write relationship. When RPC node A reads / writes a table on storage node B, this relationship is "read / write". According to some implementations of this disclosure, the read / write relationship can be further subdivided. For MySQL nodes, it can be subdivided into insert / update / delete / select. For Redis nodes, based on the definition of Redis data storage, more relationships can be subdivided, such as get / set, etc. For another example, "contain" can be used to represent a containment relationship. Specifically, a containment relationship exists between node type = rpc.service and node type = rpc.method.
[0069] According to some implementations of this disclosure, an edge can be uniquely identified based on the from, to, and type defined above. For example, the edge identifier can be determined based on from + to + type. Alternatively and / or additionally, the edge's digest identifier can be determined using a digest algorithm in a manner similar to that used for processing nodes, and a display name can be assigned to the edge, etc.
[0070] According to some implementations of this disclosure, edge determination includes: in response to determining that the master node identifier is empty, setting the edge type based on the lineage type. For example, a `custom_type` field can be added to the edge attribute to store the original edge type. When storing the edge type, the following rule can be used: if `master_node_id` is not empty, then the edge type `type` is set to `master_node_id`; if `master_node_id` is empty, then the edge type is set to the specific lineage type. Specifically, the edge type `type` of an online lineage can be set to `master_node_id` or `online`, the edge type `type` of a mobile lineage can be set to `mobile`, and the edge type `type` of an offline lineage can be set to `offline`. Using some implementations of this disclosure, richer information can be stored through combinations of master node identifiers and types, thereby achieving the goal of enabling a single data lineage model to cover multiple lineage types.
[0071] According to some implementations of this disclosure, an edge further includes at least one of the following: an edge extension field, representing an extended attribute associated with the edge; an update time, representing the time when the edge was updated; and a creation time, representing the time when the edge was created. Here, the edge extension field can store any desired information associated with edges of different types. For example, key-value pairs can be used to store the edge's extended information, including but not limited to: the edge's discovery method, the edge's confidence level, etc. Alternatively and / or additionally, the update time can represent the time when the edge was updated; and the creation time can represent the time when the first node was created. Using some implementations of this disclosure, the edges of the data lineage model can include richer information, thereby facilitating interoperability between multiple data lineages to connect various lineages and achieve connectivity between multiple types of lineages.
[0072] The details of the edges have been described. According to some implementations of this disclosure, the edges can be represented using the data structure shown in Table 4 below.
[0073] Example of data structure for Table 4
[0074] According to some implementations of this disclosure, data lineages represented in a traditional data lineage graph and / or data flow manner (or other manner) can be received, and the nodes and edges between them can be generated according to the process described above. Taking online lineages as an example, a data lineage model as shown in Figure 3 can be generated, which shows a block diagram 300 of a data lineage model according to some implementations of this disclosure.
[0075] As shown in Figure 3, the data lineage model includes multiple nodes and multiple edges, where different shapes can represent different types of nodes. For example, nodes 310, 311, 312, 313, and 314 can represent service-related nodes, nodes 320, 321, 322, and 323 can represent service-related nodes, and nodes 330, 331, 332, 333, 334, and 335 can represent nodes associated with fields of the service's method.
[0076] In constructing a data lineage model, there are typically retrieval relationships between dimensions at the same level, and containment relationships between dimensions at different levels. However, since each type of node has a different number of levels and focuses on different things, there is no need to forcibly restrict the lineage connections between dimensions at the same level. For example, the cronjob service represented by node 311 in Figure 3 has only one level of dimension; this node can be connected to both RPC services and RPC methods. In this way, the data lineage model can support the construction of multiple lineage relationships.
[0077] According to some implementations of this disclosure, multiple nodes can be stored in a relational database, edges can be stored in a graph database, and node extension fields and edge extension fields can be stored in a key-value database. In this way, the created data lineage model can be stored, thereby enabling interoperability between multiple lineages. Specifically, after determining the data lineage model, an interface can be provided via a merged lineage storage service to support the writing of data lineages. See Figure 4 for further details, which shows a block diagram 400 for storing a data lineage model according to some implementations of this disclosure.
[0078] As shown in Figure 4, each bloodline is responsible for generating its own bloodline and transforming the data model of the bloodline it produces through its own "bloodline synchronization service" (if the data model used during construction is the same, no further transformation is needed), and then connecting the data with the "fusion bloodline storage service" so that the data can be stored in the storage.
[0079] Based on some implementations of this disclosure, three types of storage can be provided: relational databases, graph databases, and key-value databases (optional). Specifically, relational databases can store node-related information. When performing lineage queries, one can first input the node's attributes to perform the query, and then query the corresponding edges based on that node. Relational databases are suitable for this application scenario and facilitate quickly finding the corresponding nodes. Typically, for large-scale data lineages, the number of edges may be in the tens of billions or even more, making it difficult for general relational databases to support such large-scale data storage. Furthermore, relational databases are not very friendly to graph retrieval (requiring multi-level joins). Therefore, graph databases can be used to store edges. Graph databases circumvent these problems and can also effectively support common graph query syntax.
[0080] It should be understood that extended fields for nodes and edges can involve additional information defined using key-value pairs. In this case, a key-value database can be used to store this additional information. This additional information can further store various aspects of the data lineage, thus providing a richer data foundation for downstream processing. Such additional information may consume significant storage space; storing it directly in a relational or graph database could degrade its processing performance. Therefore, storing this information in an additional key-value database facilitates the use of the key-value database's own functionalities to support data storage and retrieval.
[0081] As shown in Figure 4, the lineage processing service 410 can provide various services: client-side lineage service 420, online lineage service 422, offline lineage service 424, and other lineage services 426. Specifically, various types of lineage services can generate data lineage models conforming to the format described above, which can include multiple nodes and multiple edges. Furthermore, each lineage model can be provided to the fused lineage storage service 450 to store different parts of the data lineage model in different databases.
[0082] Specifically, the client-side lineage generation service 430 can generate a data lineage model whose lifecycle is in the client-side stage according to the above format. Then, the client-side lineage synchronization service 440 can synchronize the generated data lineage model to the converged storage service 450. The online lineage generation service 432 can generate a data lineage model whose lifecycle is in the online stage according to the above format. Then, the online lineage synchronization service 442 can synchronize the generated data lineage model to the converged storage service 450. The offline lineage generation service 434 can generate a data lineage model whose lifecycle is in the offline stage according to the above format. Then, the offline lineage synchronization service 444 can synchronize the generated data lineage model to the converged storage service 450. Other lineage generation services 436 can generate data lineage models whose lifecycle is in other stages according to the above format. Then, other lineage synchronization services 446 can synchronize the generated data lineage models to the converged storage service 450.
[0083] Furthermore, the fusion lineage storage service 450 can store the fused multiple nodes in the relational database 460, the fused multiple edges in the graph database 462, and the corresponding attribute data in the key-value database 464. Using some implementation methods of this disclosure, different parts of the data lineage model can be stored in different databases in a manner more suitable for data storage and retrieval.
[0084] According to some implementations of this disclosure, another data lineage model can be created in response to receiving another data lineage, as described above. Furthermore, multiple data lineage models can be merged into a single data lineage model. Specifically, matching nodes from two data lineage models can be merged. In response to determining that a third node in another data lineage matches a first node, the data lineage model is updated based on the edges associated with the third node; and in response to determining that a fourth node in another data lineage does not match any of the multiple nodes, the data lineage model is updated based on the fourth node and the edges associated with the fourth node.
[0085] Here, we can determine whether two nodes in the two models match. If the corresponding attributes of two nodes match, they are considered identical and can be represented by a single node in the merged model. If the corresponding attributes of two nodes do not match perfectly (e.g., different types or lineages, etc.), they are considered different and require different nodes to be represented in the merged model. Then, following the principles of graph merging, the nodes and edges in the merged model can be updated. In this way, a single model can be used to store information from different data lineages.
[0086] According to some implementations of this disclosure, storing the first node in the relational database includes at least one of the following: throwing an exception in response to determining that writing the first node to the relational database has failed; or rewriting the edge associated with the first node to the graph database in response to determining that writing the first node to the relational database has succeeded and in response to determining that writing the edge associated with the first node to the graph database has failed. Using some implementations of this disclosure, data consistency across databases can be ensured, thereby ensuring the integrity of the data lineage model.
[0087] It should be understood that introducing multiple storage types may lead to data consistency issues. For example, a write to database 1 might succeed while a write to database 2 fails. Suppose edges are written first, followed by nodes. If the edge write succeeds but the node write fails, consumers of the data lineage can retrieve the newly written lineage information. However, some nodes in this lineage are missing due to the write failure, resulting in data inconsistency.
[0088] In this disclosure, nodes can be written before edges. The "integrated lineage storage service" supports update and insert semantics, and data consistency can be achieved through retries on failures. Specifically, if writing to a node fails, an exception can be thrown directly so that the upstream can retry. In this case, there is no impact on consumers of the data lineage; consumers will not notice the failed lineage information. If writing to a node succeeds but writing to an edge fails, an exception can also be thrown directly so that the upstream can retry. This also has no impact on consumers of the data lineage; a failed edge write means that the failed lineage information cannot be queried, and consumers will not notice the failed lineage information.
[0089] The above has described the specific process for solving the data layer interoperability problem. However, there is still the issue of connection between different types of lineages. See Figure 5 for more details, which shows a block diagram 500 of a data lineage model according to some implementations of this disclosure. The key lies in the lineage type on the nodes; each type of lineage writes to the nodes of its respective lineage. If node A appears in both the "client lineage" and the "online lineage," then the lineage type corresponding to synchronizing node A to the merged lineage storage service in the "client lineage" is 1, and the lineage type corresponding to synchronizing node A in the "online lineage" is 2.
[0090] When writing to a node using the fusion lineage storage service, this field needs to be aggregated: It checks if the node exists in the database; if not, it's written directly; if it exists, the data is read and the lineage type is merged, ultimately updating the lineage type to 3, indicating that the node exists in both the "client lineage" and the "online lineage." Such a node can be called a connection point. When constructing lineages, each type of lineage must establish a lineage relationship (edge) with the connection point, connecting the various lineages to achieve interaction between them. According to some implementation methods disclosed herein, a node can belong to multiple types of data lineages. However, an edge belongs to only one type of data lineage.
[0091] A data lineage model can be used to achieve the coexistence of asymmetric and symmetric lineages. In the context of this disclosure, an edge is uniquely identified by its start node, end node, and edge type. For the lineage graphs in Figure 1B with A as the master node and B as the master node, both include a "call" edge from B to A. In this case, the proposed model is expected to represent that the edge belongs to both A and B. According to some implementations of this disclosure, different types of lineages can be connected by nodes, each node being fixed and unique, and the edge type can be adjusted. For a lineage graph containing a master node, the edge type is set to the ID of this master node (i.e., master_node_id in the node model), while the original edge type is stored in the edge's extended attribute custom_type.
[0092] For lineage graphs that do not contain master nodes, the edge type is set to the lineage type. For example, the "contains" edge type of the "online lineage" is set to "online," and the types of all edges in the "offline lineage" and "client lineage" are set to "offline" and "client," respectively. This rule can be used to easily distinguish edges from different lineages. Each lineage may be used differently, so it's necessary to distinguish which specific lineage an edge belongs to. For example, if both the "client lineage" and "offline lineage" have edges with type=depend, then this "depend" cannot be used to distinguish which lineage the edge belongs to. If other lineages besides "online" also need to introduce the concept of master nodes, since nodes may belong to multiple lineage types, it's difficult to distinguish which lineage an edge belongs to. Therefore, an extended attribute "lineage_type (lineage type)" can be added to the edge to facilitate this distinction.
[0093] See Figure 6 for further details. Figure 6 illustrates a block diagram 600 of multiple edges in a data lineage model according to some implementations of this disclosure. Based on Figure 1B as input, the data lineage model shown in Figure 6 can be obtained, and this single data lineage model can include information from multiple data lineage graphs in Figure 1B. Figure 6 includes five nodes, each corresponding to a lineage graph. In actual storage, not five lineage graphs are stored, but only one is stored; this is achieved by setting the edge type to the ID of the master node. For a lineage graph containing a master node, traversal is performed during lineage retrieval based on the ID of the input node as the edge type.
[0094] Specifically, edges 514, 520, 530, and 540 labeled with node A correspond to edges in the lineage graph with node A as the master node; edges 512 and 532 labeled with node B correspond to edges in the lineage graph with node B as the master node; edges 522 and 542 labeled with node C correspond to edges in the lineage graph with node D as the master node; and edges 510 and 534 labeled with node D correspond to edges in the lineage graph with node D as the master node. In this way, a single data lineage model can be used to store information from multiple lineage graphs.
[0095] For the case of multiple lineages, see Figure 7 for further details, which shows a block diagram 700 of a data lineage model according to some implementations of this disclosure. As shown in Figure 7, other non-"online lineages" are set to the type of this lineage and an extended custom_type field is added.
[0096] When performing a lineage lookup, if we start with M in the "client lineage," we can first find the lineage relationship M->B in the "client lineage." We then discover that B also belongs to the "online lineage," so we use node B as the master node to find the lineage relationship B->A->D in the "online lineage." Next, we find that D also belongs to the "offline lineage," so we can find the lineage relationship D->X->Y in the "offline lineage." In this way, we can find M->B->A->D->X->Y, spanning the entire lifecycle of the data.
[0097] If you enter B in the query, it belongs to both the "client lineage" and the "offline lineage." You can search in both lineages separately and find the lineage relationships M->B and B->A->D respectively. Similar to the description above, you can eventually find M->B->A->D->X->Y, that is, throughout the entire lifecycle of the data.
[0098] Based on some implementations of this disclosure, edges can be defined in different ways. Specifically, if the edge is an edge of a master node in an "online lineage," the edge type can be set to the master node's ID, i.e., `master_node_id`. If the edge is an "online lineage" but does not belong to a master node, for example, if the type is "containment," then the edge type can be set to "online." Simultaneously, the `custom_type` field can be set to "containment." For "client lineage" and "offline lineage," this field can be set to "client" or "online," respectively. Furthermore, if it is desired to introduce the concept of a master node in "client lineage" and "offline lineage," then a similar approach to "online lineage" can be used for modification.
[0099] See Figure 8 for further details, which illustrates a block diagram 800 of a data lineage model according to some implementations of this disclosure. When performing a lineage query, a starting point A is first set, and data rows with `from = A` are searched in the database. Assuming two rows are found: `from = A` and `to = B / C`, the first hop of the lineage relationship is found. Setting `from` to B / C continues the search in the database; for example, finding the relationship between `from = B` and `to = D`, the second hop of the lineage relationship is found. Continuing to set `from` to D allows for the discovery of more lineage information. Similarly, if a key-value database is used to store edges, the connection relationship between `from` and `to` can be predefined.
[0100] Based on some implementations of this disclosure, the methods described above can be used to establish a data lineage model, thereby improving data transparency. The proposed data lineage model can illustrate in detail the various transformations and movements of data from its origin to systems and processes by drawing a comprehensive data map. This supports managers in understanding the sources, operations, and utilization of data within the infrastructure. The proposed data lineage model can improve the efficiency of data propagation management. By identifying the propagation links of data of interest, such as collection, processing, and storage, and combining data coloring capabilities, it can support managers in understanding the assets and propagation paths involved in the data of interest. The proposed data lineage model can improve data integrity verification, allowing the verification of the accuracy, integrity, and consistency of data throughout its entire lifecycle, thereby ensuring data integrity. Furthermore, it can detect and correct discrepancies or inconsistencies, ensuring compliance with various data security specifications.
[0101] According to some implementations of this disclosure, data lineage models can serve as a basis for data management, thereby meeting various data security requirements. Furthermore, data lineage can be widely applied in the fields of data assets and application development. For example, the frequency of data usage can be represented by citation popularity calculations. Frequent consumption and widespread citation of data serve as strong evidence of its authority. Similar to page ranking values in web page citations, citation popularity values can be defined based on the downstream lineage of data. Data with high popularity is more trustworthy. Data lineage models can be used to understand data context. When searching for data, lineage can be used to determine whether the current data is needed by the user or whether it is trustworthy. Alternatively and / or additionally, data lineage models can be used to implement tag propagation. Security tags for some data can be automatically identified (or manually) according to rules, and based on lineage, the tags can be automatically propagated to a wider range of downstream data.
[0102] In application development, data lineage models can be used to perform impact analysis. Before upstream developers modify data, downstream developers are notified to make the corresponding changes. Alternatively and / or additionally, in attribution analysis, when a service experiences a problem, the upstream services in the lineage are examined to identify the cause of the problem.
[0103] Example process
[0104] Figure 9 illustrates a flowchart of a method 900 for managing data lineage according to some implementations of this disclosure. At block 910, in response to receiving a data lineage, multiple nodes are determined based on multiple services associated with the data lineage. A first node among the multiple nodes is created based on a first service among the multiple services. A first type of the first node includes a set of dimensions, each dimension corresponding to a set of entities in a set of service hierarchies. At block 920, an edge between the first node and a second node among the multiple nodes is determined based on the association relationships between the multiple services. At block 930, a data lineage model describing the data lineage is created based on the multiple nodes and the edge.
[0105] According to some implementations of this disclosure, the first node includes: node attributes, the node attributes including at least one attribute corresponding to the first type; and an identifier, the identifier being determined by the first type and the node attributes.
[0106] According to some implementations of this disclosure, the first node further includes: lineage type, which indicates the stage of the first service in the data lineage, and the lineage type includes at least one of the following: client stage, online stage, or offline stage.
[0107] According to some implementations of this disclosure, the first node further includes at least one of the following: a node extension field, which represents an extension attribute associated with the first node, the extension attribute being marked using a lineage type; a node name, which represents the name of the first node being presented; an update time, which represents the time when the first node was updated; and a creation time, which represents the time when the first node was created.
[0108] According to some implementations of this disclosure, an edge includes: a start node, which represents the starting point of the edge; an end node, which represents the ending point of the edge; a master node identifier, which represents the master node to which the edge belongs; and an edge type, which represents the association relationship between the first node and the second node.
[0109] According to some implementations of this disclosure, determining the edge includes: in response to determining that the master node identifier is empty, setting the edge type based on the lineage type.
[0110] According to some implementations of this disclosure, an edge further includes at least one of the following: an edge extension field, which represents an extension attribute associated with the edge; an update time, which represents the time when the edge is updated; and a creation time, which represents the time when the edge is created.
[0111] According to some implementations of this disclosure, multiple nodes are stored in a relational database, edges are stored in a graph database, and node extension fields and edge extension fields are stored in a key-value database.
[0112] According to some implementations of this disclosure, the first node is stored in the relational database in the following manner: in response to determining that writing the first node to the relational database has failed, an exception is thrown; or in response to determining that writing the first node to the relational database has succeeded and in response to determining that writing the edge associated with the first node to the graph database has failed, the edge associated with the first node is rewritten to the graph database.
[0113] According to some implementations of this disclosure, the method further includes: in response to receiving another data lineage, creating another data lineage model; in response to determining that a third node in the other data lineage matches a first node, updating the data lineage model based on the edge associated with the third node; and in response to determining that a fourth node in the other data lineage does not match any of the multiple nodes, updating the data lineage model based on the fourth node and the edge associated with the fourth node.
[0114] Example devices and equipment
[0115] Figure 10 shows a block diagram of an apparatus 1000 for managing data lineage according to some implementations of the present disclosure. The apparatus includes: a node determination module 1010 configured to, in response to receiving a data lineage, determine multiple nodes based on multiple services associated with the data lineage, wherein a first node among the multiple nodes is created based on a first service among the multiple services, and a first type of the first node includes a set of dimensions, each set of dimensions corresponding to a set of entities in a set of hierarchical levels of the services; an edge determination module 1020 configured to determine an edge between the first node and a second node among the multiple nodes based on the association relationships between the multiple services; and a creation module 1030 configured to create a data lineage model describing the data lineage based on the multiple nodes and the edges.
[0116] According to some implementations of this disclosure, the first node includes: node attributes, the node attributes including at least one attribute corresponding to the first type; and an identifier, the identifier being determined by the first type and the node attributes.
[0117] According to some implementations of this disclosure, the first node further includes: lineage type, which indicates the stage of the first service in the data lineage, and the lineage type includes at least one of the following: client stage, online stage, or offline stage.
[0118] According to some implementations of this disclosure, the first node further includes at least one of the following: a node extension field, which represents an extension attribute associated with the first node, the extension attribute being marked using a lineage type; a node name, which represents the name of the first node being presented; an update time, which represents the time when the first node was updated; and a creation time, which represents the time when the first node was created.
[0119] According to some implementations of this disclosure, an edge includes: a start node, which represents the starting point of the edge; an end node, which represents the ending point of the edge; a master node identifier, which represents the master node to which the edge belongs; and an edge type, which represents the association relationship between the first node and the second node.
[0120] According to some implementations of this disclosure, the edge determination module is further configured to: set the edge type based on lineage type in response to determining that the master node identifier is empty.
[0121] According to some implementations of this disclosure, an edge further includes at least one of the following: an edge extension field, which represents an extension attribute associated with the edge; an update time, which represents the time when the edge is updated; and a creation time, which represents the time when the edge is created.
[0122] According to some implementations of this disclosure, multiple nodes are stored in a relational database, edges are stored in a graph database, and node extension fields and edge extension fields are stored in a key-value database.
[0123] According to some implementations of this disclosure, the first node is stored in the relational database in the following manner: in response to determining that writing the first node to the relational database has failed, an exception is thrown; or in response to determining that writing the first node to the relational database has succeeded and in response to determining that writing the edge associated with the first node to the graph database has failed, the edge associated with the first node is rewritten to the graph database.
[0124] According to some implementations of this disclosure, the apparatus further includes a merging module configured to: create another data lineage model in response to receiving another data lineage; update the data lineage model based on the edge associated with the third node in response to determining that a third node in the other data lineage matches a first node; and update the data lineage model based on the fourth node and the edge associated with the fourth node in response to determining that a fourth node in the other data lineage does not match any of the plurality of nodes.
[0125] Figure 11 shows a block diagram of a device 1100 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1100 shown in Figure 11 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 1100 shown in Figure 11 can be used to implement the methods described above.
[0126] As shown in Figure 11, computing device 1100 is in the form of a general-purpose computing device. Components of computing device 1100 may include, but are not limited to, one or more processors or processing units 1110, memory 1120, storage devices 1130, one or more communication units 1140, one or more input devices 1150, and one or more output devices 1160. Processing unit 1110 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1120. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 1100.
[0127] Computing device 1100 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 1100, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1120 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1130 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 1100.
[0128] The computing device 1100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG11, disk drives for reading or writing from removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading or writing from removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 1120 may include a computer program product 1125 having one or more program modules configured to perform various methods or actions of various implementations of this disclosure.
[0129] The communication unit 1140 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 1100 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 1100 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.
[0130] Input device 1150 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1160 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 1100 can also communicate as needed with one or more external devices (not shown) via communication unit 1140. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 1100, or with any device that enables computing device 1100 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interfaces (not shown).
[0131] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
[0132] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0133] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0134] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0136] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for managing data lineage, comprising: In response to receiving a data lineage, multiple nodes are determined based on multiple services associated with the data lineage. A first node among the multiple nodes is created based on a first service among the multiple services. A first type of the first node includes a set of dimensions, which respectively correspond to a set of entities in a set of levels of the service. Based on the association relationships between the multiple services, determine the edge between the first node and the second node among the multiple nodes; as well as Based on the multiple nodes and the edges, a data lineage model describing the data lineage is created.
2. The method according to claim 1, wherein the first node comprises: Node attributes, wherein the node attributes include at least one attribute corresponding to the first type; as well as The identifier is determined by the first type and the node attribute.
3. The method according to claim 2, wherein the first node further comprises: Lineage type, which indicates the stage in which the first service is located in the data lineage, includes at least one of the following: client stage, online stage, or offline stage.
4. The method of claim 3, wherein the first node further comprises at least one of the following: Node extension fields, which represent extended attributes associated with the first node, and the extended attributes are marked using the lineage type; Node name, where the node name represents the name of the first node being presented; Update time, where the update time represents the time when the first node was updated; as well as Creation time, which refers to the time when the first node was created.
5. The method of claim 2, wherein the edge comprises: The start node represents the starting point of the edge; End node, where the end node represents the endpoint of the edge; Master node identifier, wherein the master node identifier indicates the master node to which the edge belongs; and Edge type, which represents the association between the first node and the second node.
6. The method of claim 5, wherein determining the edge comprises: In response to determining that the master node identifier is empty, the edge type is set based on the lineage type.
7. The method of claim 5, wherein the edge further comprises at least one of the following: An edge extension field, wherein the edge extension field represents an extension attribute associated with the edge; Update time, where the update time represents the time when the edge was updated; and Creation time, which refers to the time when the edge was created.
8. The method of claim 1, wherein the plurality of nodes are stored in a relational database, the edges are stored in a graph database, and the node extension field and the edge extension field are stored in a key-value database.
9. The method of claim 8, wherein the first node is stored in the relational database in the following manner: In response to determining that writing to the first node to the relational database has failed, an exception is thrown; or In response to determining that writing the first node to the relational database was successful, and in response to determining that writing the edge associated with the first node to the graph database failed, the edge associated with the first node is rewritten to the graph database.
10. The method of claim 1, further comprising: In response to receiving another data lineage, create another data lineage model; In response to determining that a third node in the other data lineage matches the first node, the data lineage model is updated based on the edges associated with the third node; as well as In response to determining that a fourth node in the other data lineage does not match any of the plurality of nodes, the data lineage model is updated based on the fourth node and the edges associated with the fourth node.
11. An apparatus for managing data lineage, comprising: A node determination module is configured to determine multiple nodes based on multiple services associated with the data lineage in response to receiving a data lineage. A first node among the multiple nodes is created based on a first service among the multiple services. A first type of the first node includes a set of dimensions, which respectively correspond to a set of entities in a set of levels of the service. An edge determination module is configured to determine the edge between a first node and a second node among the plurality of nodes based on the association relationships between the plurality of services; as well as. A creation module is configured to create a data lineage model describing the data lineage based on the plurality of nodes and the edges.
12. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.
13. A computer-readable storage medium having a computer program stored thereon, the computer program causing the processor to implement the method according to any one of claims 1 to 10 when executed by a processor.
14. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 10.