Generation, Access, and Display of System Metadata

A dedicated lineage server using a specialized data structure in random access memory enhances data lineage tracking efficiency, addressing slow query responses and resource inefficiencies in existing systems by enabling rapid lineage metadata access and visualization.

JP7712998B2Active Publication Date: 2025-07-24AB INITIO TECHNOLOGY LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023216403
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-12-01
Filing Date
2023-12-22
Publication Date
2025-07-24
Estimated Expiration
2037-12-01

AI Technical Summary

Technical Problem

Existing data processing systems face inefficiencies in managing and accessing lineage metadata, leading to slow query responses and high resource usage when tracking data lineage due to the large amount of processing time required to generate and display lineage diagrams.

Method used

Implementing a dedicated lineage server that stores lineage metadata in a specialized data structure optimized for speed and efficiency, utilizing random access memory to minimize storage space and processing time, allowing for faster query responses and reduced resource consumption.

Benefits of technology

Enables lineage metadata queries to be processed up to 500 times faster than traditional methods, improving the efficiency and speed of data lineage tracking and visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007712998000001
    Figure 0007712998000001
  • Figure 0007712998000002
    Figure 0007712998000002
  • Figure 0007712998000003
    Figure 0007712998000003
Patent Text Reader

Abstract

To provide data structures and methods for generating, accessing, and displaying lineage metadata.SOLUTION: A method includes: receiving a portion of metadata from a data source, the portion of metadata describing nodes and edges; generating instances of a data structure representing the portion of metadata; storing the instances of the data structure in a random access memory; receiving a query that includes identification of at least one particular element of data; and using at least one instance of the data structure to cause a display of a computer system to display representation of lineage of the particular element of data. At least one instance of the data structure includes an identification value that identifies a corresponding node, one or more property values representing respective properties of the corresponding node, and one or more pointers corresponding to respective identification values. Each pointer represents a side associated with a node identified by the corresponding respective identification value.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Technical Field This application relates to data structures and methods for generating, accessing, and displaying lineage metadata, such as the lineage of data elements stored within a data storage system.

Background Art

[0002] Background Enterprises use data processing systems such as data warehousing, customer relationship management, and data mining to manage data. In many data processing systems, data is pulled into a central repository from many different data sources such as database files, operating systems, flat files, the Internet, and other information sources. Often, the data is transformed before being loaded into the data system. The transformation includes cleansing, integration, and extraction. Metadata can be used to track the data, the source of the data, and the transformations occurring on the data stored within the data system. Metadata (sometimes called "data about data") is data that describes the attributes, format, origin, history, interrelationships, etc. of other data. The management of metadata can play a central role in complex data processing systems.

[0003] Users may want to investigate how specific data is derived from various data sources. For example, a user may want to know how a dataset or data object was generated, or from which source a dataset or data object was imported. Tracking a dataset back to its source of derivation is called data lineage tracking (or "upstream data lineage tracking"). A user may also want to investigate how a specific dataset is being used, for example, which applications have read a given dataset ("downstream data lineage tracking" or "impact analysis"). A user may also be interested in knowing how one dataset relates to other datasets. For example, a user may want to know if a dataset has been modified and which tables are affected.

[0004] A lineage, which is a type of metadata, enables users to obtain answers to questions about data lineage (such as "where did a given value come from?", "how was the output value calculated?", and "which applications created and depend on this data?"). Users can understand the consequences of proposed changes (such as "if this part is changed, what else will be affected?" and "if this source format is changed, which applications will be affected?"). Users can also obtain answers to questions that involve both technical and business metadata (such as "which group is responsible for creating and using this data?", "who last modified this application?", and "what changes did they make?").

[0005] The answers to these questions can analyze a complex data processing system and help address the problem. For example, if a certain data element has an outlier, any number of past inputs or data processing steps may be involved in that outlier. Therefore, the system may be presented to the user in the form of a diagram that includes visual elements representing the data element of interest, as well as visual elements representing other data elements that affect or are affected by the data element of interest. The user can look at the diagram and visually identify other data elements and / or transformations that affect the data element of interest. As an example, by using this information, the user can know whether any of the data elements and / or transformations could be the cause of the outlier, and if a problem is discovered, correct (or flag for correction) any of the underlying data processing steps. As another example, by using this information, the user can identify any data elements or transformations that may be essential to a part of the system (e.g., removing them from the system would affect the data element of interest) and / or data elements or transformations that may not be essential to a part of the system (e.g., removing them from the system would not affect the data element of interest). Summary of the Invention Means for Solving the Problems

[0006] Overview In particular, the present inventors describe a method executed by a data processing device, the method comprising receiving a portion of metadata from a data source, wherein the portion of metadata describes nodes and edges, at least some of the edges representing the influence of one node on another node respectively, each edge having a single direction; generating an instance of a data structure representing the portion of metadata, wherein at least one instance of the data structure includes an identification value identifying a corresponding node, one or more characteristic values representing respective characteristics of the corresponding node, and one or more pointers for each identification value, each pointer representing an edge associated with the node identified by the corresponding identification value; storing the instance of the data structure in random access memory; receiving a query including an identification of at least one specific data element; and using at least one instance of the data structure to display a representation of the lineage of the specific data element on a display of a computer system.

[0007] These techniques can be implemented in several ways, as a method, as a system, and / or as a computer program product stored on a computer-readable storage device.

[0008] Aspects of these techniques may include one or more of the following advantages. Lineage metadata can be stored using a dedicated data structure designed to obtain speed and efficiency when responding to queries against the lineage metadata. The lineage metadata can be stored in memory, so that a computer system storing the lineage metadata can respond to queries against the lineage metadata more quickly than if the lineage metadata were not stored in memory (e.g., if the lineage metadata were stored on a hard disk or in another type of storage technique and accessed therefrom). Using the techniques described herein, lineage data can be obtained much faster, e.g., 500 times faster, than with other techniques. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Description of the Drawings

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 2E

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6

Best Mode for Carrying Out the Invention

[0010] Like reference numerals in the various drawings indicate like elements.

[0011] Description A system for managing access to metadata can receive a query from a user requesting a lineage of specific data elements and, in response, send a diagram representing the lineage of the data elements. If the data elements belong to a data storage system that stores a relatively large amount of data, the system for managing access to metadata may need to spend a large amount of processing time processing the lineage of the data elements and generating the corresponding diagram. However, by introducing a system dedicated to processing lineage metadata and optimized for that type of processing, the processing can be accelerated and made more efficient. Accordingly, this specification describes a technique for using a dedicated system to process and store lineage metadata in a generally faster and more efficient manner than without using a dedicated system.

[0012] FIG. 1A shows a metadata processing environment 100 that includes a lineage server 102 that stores lineage metadata and provides it to other systems within the environment. The metadata processing environment 100 also includes a metadata server 104 that generally responds to requests for obtaining metadata. The metadata server 104 can access metadata 106 stored within a metadata database 108. The metadata 106 comes from data sources 110A-C that continuously supply metadata 112A-C to the metadata database 108. For example, the data sources 110A-C can be any combination of relational databases, flat files, network sources, etc.

[0013] In use, the metadata server 104 responds to a query 114 received from a user terminal 116 operated by a user 118. For example, the user terminal 116 can be a computing device such as a personal computer, a laptop computer, a tablet device, a smartphone, etc. In some examples, for instance, when the metadata server 104 is configured to provide access to data over a network and includes or communicates with a web server that can interface with a web browser, the user terminal 116 operates a network-based user application such as a web browser. Generally, much of the interaction between the computer systems described herein can be performed using a similar communication network that uses communication protocols commonly used on the Internet or a network of that kind.

[0014] The metadata server 104 is configured to respond to queries for multiple types of metadata. Since one type of metadata is lineage metadata, the metadata server 104 can process a query 114 that requests the lineage of a particular data element, such as a data element 120 stored by one of data sources 110A - C, by accessing lineage metadata 106 that describes the lineage of the particular data element stored in the metadata database 108. The metadata server 104 can then provide the user terminal 116 with lineage metadata 122, such as lineage metadata in the form of a lineage diagram (described below with respect to FIGS. 2A - 2E).

[0015] In some examples, processing query 114 regarding lineage metadata is a task that requires a relatively large amount of processing time and / or uses a relatively large amount of processing resources on metadata server 104. For example, to process query 114, metadata server 104 may need to access metadata 106 stored in metadata database 108. In this example, to access all of the required metadata, metadata server 104 may expend processing resources to generate a query to metadata database 108. Further, the process of transmitting a query to metadata database 108 and waiting for a response incurs latency, such as communication network latency. Additionally, in some examples, metadata server 104 must process metadata received from metadata database 108 to extract the metadata required to create a lineage diagram. For example, since metadata database 108 stores various types of metadata in addition to lineage metadata, the metadata received from metadata database 108 may include information not directly related to the lineage of the data elements of interest, and thus additional processing time is used to identify and remove information not related to the lineage.

[0016] In some implementations, for example, to improve the performance of metadata processing environment 100, lineage server 102 is used to provide lineage metadata to metadata server 104. Lineage server 102 is a dedicated system that stores lineage metadata 124 in a form that is generally more quickly and efficiently accessible than techniques that do not use a lineage server. Specifically, lineage server 102 stores lineage metadata 124 using a dedicated data structure that is designed to obtain speed and efficiency when responding to queries for lineage metadata. The data structure defines the configuration of the data such that all data stored using a particular data structure is configured in the same way. The data structure techniques will be described in more detail below with respect to FIG. 3.

[0017] In use, the system server 102 transmits a query 126 to the metadata database 108 to obtain system metadata 128. The system server 102 preferably stores system metadata regarding most or all of a wide number of data elements 120, such as those stored by data sources 110A - C. In this way, the system server 102 can respond to system queries for most of the data elements 120 for which a query can be made. When the system server 102 receives system metadata 128 from the metadata database 108, the system server 102 updates its data structure that includes stored system metadata 124. In some examples, the system server 102 periodically transmits a new query 126 to the metadata database 108, such as hourly, daily, or at another interval, to store a relatively large number of relatively up - to - date system metadata. For example, this interval can be a scheduled interval corresponding to, for example, schedule data maintained by the system server 102.

[0018] As shown in more detail in FIG. 1B, the metadata server 104 can receive a query 114 for system metadata (e.g., a query for system metadata that can be used to display a system diagram on the user terminal 116) and provide that query 114 to the system server 102. The system server can then return system metadata 122 in response to query 114. The metadata server 104 does not need to spend much processing time or use many processing resources to prepare the received system metadata 122 for providing to the user terminal 116 by comparing it with system metadata obtained using other techniques such as obtaining system metadata from the metadata database 108.

[0019] Figures 1C and 1D show the elements of the lineage server 102 and the metadata server 104 and how they interact. As described above, the metadata server 104 receives a query 114 that requests a lineage of specific data elements. The query 114 identifies the specific data element for which the lineage is requested (e.g., one of the data elements 120 in FIG. 1A). The metadata server 104 uses the identification of the data element to select a walk plan 130 from a set of walk plans 132 that can be used to gather the lineage metadata related to that data element. The walk plan 130 is a data structure (e.g., a structured document including tagged parts such as an XML document) that describes how to traverse (i.e., "walk") through a set of lineage metadata in a specific way. In some examples, the walk plan can be selected based on the data type of the data element. For example, a specific data type may be associated with a specific one of the walk plans 132 (e.g., an association stored in a related index accessible to the metadata server 104). The walk plans 132 will be described in detail below with respect to FIG. 4.

[0020] Once the walk plan 130 is selected, the metadata server 104 transmits the query 114 and the walk plan 130 to the lineage server 102. In response, the lineage server 102 identifies the lineage metadata related to the query 114 from within its data structure 134 of lineage metadata. The data structure 134 is a representation of the lineage metadata configured in a way that minimizes the amount of storage space required to include the lineage metadata without omitting any data necessary to respond to queries against the lineage metadata. Thus, the lineage server 102 can typically use the data stored within its data structure 134 to provide to the metadata server 104 all of the lineage metadata that the metadata server 104 needs to respond to the query 114. The data structure 134 will be described in detail below with respect to FIG. 3.

[0021] In some implementations, data structure 134 is loaded into memory 135 of the orchestration server 102 for fast access (e.g., fast data reads and writes). An example of a memory is a random access memory. A random access memory stores data items in such a way that each data item can be accessed in approximately the same amount of time as any other item of the same size (e.g., bytes or words). In contrast, other types of data storage areas, such as magnetic disks, have physical constraints that cause some data items to take longer to access than others, depending on the current physical state of the disk (e.g., the position of the magnetic read / write head). Data items stored in random access memory are typically stored at an address unique to that data item or at an address shared among a small number of data items. Random access memory is typically volatile, so that when the random access memory is disconnected from active power (e.g., when the computer system loses power), the data stored in the random access memory is lost. In contrast, magnetic disks and some other types of data storage areas are non-volatile and retain data without active power.

[0022] Since the system server 102 stores the data structure 134 in the memory 135, the system server 102 can read and write system metadata faster than techniques that do not store the data structure in the memory 135. Specifically, the data structure 134 is configured in a way that minimizes the data usage. For example, the data structure 134 can omit data such as text strings in the original system metadata obtained from the metadata server 104. Therefore, during the use of the system server 102, all of the data structure 134, for example, all of the data representing the system metadata, can be stored in the memory 135. A computer system generally has constraints regarding the amount of random access memory available at a given time (e.g., due to address limitations). Further, random access memory tends to be more expensive per byte than other types of data storage areas (e.g., magnetic disks). Therefore, when using random access memory, the data structure 134 may have an upper limit regarding its total size on a particular computer system. Thus, the techniques described herein (e.g., the techniques described below with respect to FIG. 3) minimize the size of the data structure while maintaining the information regarding the system that can be requested by the query 114.

[0023] In some implementations, the metadata server 104 also stores the system metadata 137 (e.g., the system metadata received from the metadata database 108 shown in FIG. 1A). However, since the metadata server 104 does not use, for example, the data structure of the system server 102, it does not store most of its stored system metadata 137 in random access memory. Therefore, even if the metadata server 104 stores some system metadata 137, the metadata server 104 can access the system metadata 124 of the system server 102 to obtain any metadata that is not stored locally in the metadata server 104. When the system server 102 is not in use, the metadata server 104 that stores some system metadata 137 generally accesses the system metadata stored in the metadata database 108 (FIG. 1A), which may result in a performance disadvantage compared to using the system server 102 as described above.

[0024] Here, a random access memory has been used as the main example, but other types of memory can also be used with the system server 102. For example, another type of memory is flash memory. Unlike random access memory, flash memory is non-volatile. However, flash memory typically has restrictions on accessing data items. Some types of flash memory are configured in such a way that a set of data items (e.g., a block of data items) is the smallest data unit that can be accessed at one time, as opposed to individually accessible data items. For example, to delete a data item on some types of flash memory, an entire block has to be deleted. To keep the remaining data items, the remaining data items can be rewritten to the flash memory.

[0025] The system server 102 traverses the data structure 134 using the walk plan 130 and collects system metadata stored within the data structure in response to the query 114. As shown in FIG. 1D, the system server 102 then sends back a response 138 that includes the system metadata 139 to the metadata server 104. The metadata server 104 can use the system metadata 139 to generate its response 140 to the query 114. The response 140 can take one of several forms. In some examples, the response 140 includes the same system metadata 139 received from the system server 102 in a form that involves, for example, minimal post-processing. In some examples, the metadata server 104 performs post-processing on the system metadata 139. For example, if the system metadata 139 is received in a non-human-readable encoding format, the metadata server 104 can, for example, change the format of the system metadata 139 to a human-readable format. In some examples, the metadata server 104 generates a system diagram based on the system metadata 139 and incorporates data representing the system diagram into the response 140. (As will be described in detail below with respect to FIGS. 2A-2E) For example, if the response 140 is a system diagram, in some examples the response 140 is transmitted to the user terminal 116 (FIG. 1A). In some examples, the response 140 is transmitted to an intermediate system and then transmitted to the user terminal and / or processed into a form suitable for transmission to the user terminal.

[0026] FIG. 2A shows an example of information displayed within a metadata viewing environment. In some examples, the metadata viewing environment is an interface that runs on a user terminal, such as the user terminal 116 shown in FIG. 1A. In the example of FIG. 2A, the metadata viewing environment displays information regarding the data system diagram 200A. An example of a metadata viewing environment is a web-based application that enables a user (such as the user 118 shown in FIG. 1A) to visualize and edit metadata. Using the metadata viewing environment, a user can search, analyze, and manage metadata from anywhere within an enterprise using a standard web browser. Each type of metadata object has one or more views or visual representations. The metadata viewing environment of FIG. 2A shows a system diagram regarding the target element 206A.

[0027] For example, this pedigree diagram represents the data representing the metadata objects stored in the metadata server 104 (FIG. 1A) and / or the end-to-end pedigree for the processing nodes, i.e., the objects (its sources) on which a given starting object depends and the objects (its targets) that a given starting object affects. In this example, the connection between two examples of metadata objects, the data element 202A and the transformation 204A, is shown. These metadata objects are represented by the nodes in the figure. The data element 202A can represent, for example, a dataset, a table within a dataset, a column within a table, and a field within a file, message, or report. An example of the transformation 204A is an element of an executable file that describes how a single data element output is created. The connections between the nodes are based on the relationships between the metadata objects.

[0028] FIG. 2B shows a corresponding pedigree diagram 200B for the same target element 206A shown in FIG. 2A, but each element 202B is grouped and displayed within the group based on context. For example, the data element 202B is grouped within a dataset 208B (e.g., a table, file, message, report), an application 210B (including executable files such as graphs, plans, programs, etc. and the datasets on which they operate), and a system 212B. The system 212B is a functional grouping of data and the applications that process the data, and the system is composed of application and data groups (e.g., a database, a file group, a messaging system, a group of datasets). The transformation 204B is grouped within the executable file 214B, the application 210B, and the system 212B. Executable files such as graphs, plans, programs, etc. read and write datasets. Parameters can be set to determine which groups are expanded by default and which groups are collapsed. By removing unnecessary levels of detail, setting such parameters allows the user to view only the details of the groups that are important to them.

[0029] Using a metadata viewing environment to perform data lineage calculations is useful for several reasons. For example, calculating and showing the relationship between data elements and transformations can help users determine how the reported values were calculated for a given field report. Users can also see which datasets store a particular type of data and which executable files read and write to that dataset. In the case of business terms, a data lineage diagram can show which data elements (e.g., columns or fields) are related to a particular business term (e.g., as defined within an enterprise).

[0030] The data lineage diagrams presented within the metadata viewing environment can also assist users in impact analysis. In particular, users may want to know which downstream executable files are affected when a column or field is added to a dataset and who needs to be notified. Impact analysis can reveal where a given data element is used and can also reveal the ramifications of changing that data element. Similarly, users can see which datasets are affected by changes within an executable file or whether it is okay to remove a particular database table from production.

[0031] Using a metadata viewing environment to perform data lineage calculations to generate a data lineage diagram is useful for managing business terms. For example, for employees within an enterprise, it is often desirable to see consistency in the meaning of business terms across the enterprise, the relationships between those terms, and the data that those terms refer to. Using business terms consistently can enhance the transparency of enterprise data and help convey business requirements. Therefore, it is important to know where to find the physical data that underlies business terms and which business logic is used in calculations.

[0032] Viewing the relationships between data nodes can also be useful when managing and maintaining metadata. For example, a user may want to know who changed a piece of metadata, what the source of the piece of metadata (or "source of the record") is, or what changes were made when loading or reloading metadata from an external source. When maintaining metadata, it may be desirable to allow a designated user to create metadata objects (such as business terms), edit the characteristics of the metadata objects (such as the description and relationships of the object to other objects), or delete metadata objects that are no longer in use.

[0033] The metadata viewing environment provides several graphical views of the objects, enabling users to explore and analyze the metadata. For example, a user can view the content of the system and applications, explore the details of any object, and use a data lineage view that enables the user to easily perform various types of dependency analysis such as the above-mentioned data lineage analysis and impact analysis to view the relationships between objects. The hierarchical structure of the objects can also be viewed, and the hierarchical structure can be searched for a specific object. When an object is found, a bookmark can be created for the object to enable the user to easily return to that object.

[0034] With appropriate permissions, a user can edit metadata within the metadata viewing environment. For example, a user can update the description of an object, create business terms, define the relationships between objects (such as linking business terms to fields in a report or columns in a table), move an object (such as moving a dataset from one application to another), or delete an object.

[0035] In FIG. 2C, a corresponding system diagram 200C for the target element 206A is shown, and the level of resolution is set according to the applications involved in the calculation of the target data element 206A. In particular, applications 202C, 204C, 206C, 208C, and 210C are shown, because only those applications are directly involved in the calculation regarding the target data element 206A. If the user wants to view any part of the system diagram at a different level of resolution (e.g., to show more or less details in the figure), the user can activate the corresponding expand / collapse button 212C.

[0036] FIG. 2D shows the corresponding system diagram 200D at different levels of resolution. In this example, the expand / collapse button 212C has been activated by the user, and the metadata viewing environment is again showing the same system diagram, but application 202C has been expanded to display the data set 214D and the executable file 216D within application 202C.

[0037] FIG. 2E shows the corresponding system diagram 200E at different levels of resolution. In this example, the user has decided to display everything that is expanded by the custom expansion. Any field or column that is the final data source (e.g., having no upstream system) is expanded. In addition, fields with a specific flag set are also expanded. In this example, specific flags are set on the data sets and fields at important intermediate points within the system, and one column is the column in which the system is being displayed.

[0038] Other examples of systems are described in U.S. Patent Application No. 12 / 629,466, entitled "VISUALIZING RELATIONSHIPS BETWEEN DATA ELEMENTS AND GRAPHICAL REPRESENTATIONS OF DATA ELEMENT ATTRIBUTES", which is hereby incorporated by reference in its entirety.

[0039] Viewing elements and relationships within the metadata viewing environment can be made more useful by adding information associated with each of the nodes that represent them. One exemplary way to add related information to a node is to graphically overlay the information on a particular node. These graphics can indicate some value or characteristic of the data represented by the node, and can be any characteristic within the metadata database. This approach has the advantage of typically combining two or more pieces of information that would not otherwise have a common point (the relationship between the data node and the characteristics of the data represented by the node) to put useful information "in context". For example, along with a visual representation of the relationships between data nodes, characteristics such as the quality of the metadata, the freshness of the metadata, the source of the record information, etc. can be displayed. Some of this information can be made accessible in tabular form, but viewing the characteristics of the data along with the relationships between the various nodes of the data may be more useful to the user. The user can choose which characteristics of the data to display on the data elements and / or transformation nodes within the metadata viewing environment. Which characteristics to display can also be set according to default system settings.

[0040] As described above with respect to FIG. 1A, the lineage server 102 stores lineage metadata in memory (e.g., random access memory) using the data structure 134. FIG. 3 shows an example of a data structure 300. In use, the lineage server 102 includes many instances of the data structure 300. An instance of a data structure is a collection of data (e.g., a collection of bits) that is formatted in the manner defined by the data structure. Instances of the data structure 300 described herein may be referred to as "nodes".

[0041] Each instance of the data structure 300 represents a metadata object, such as one of the data elements 202A or transformations 204A shown in FIG. 2A. In some examples, each instance of the data structure 300 represents a node that can be displayed in a lineage diagram, such as the diagrams 200A - 200E shown in FIGS. 2A - 2E.

[0042] In use, the system server 102 stores each data structure 300 at a memory location 302 specific to the data structure. Each data structure 300 typically points to the memory locations of other data structures.

[0043] The data structure 300 is composed of several fields. A field is a collection of data, for example, a subset of bits that create an instance of the data structure 300. The identification information field 310 contains data representing unique identification information regarding an instance of the data structure 300. The type field 312 contains data representing the type of the metadata object represented by the corresponding instance of the data structure 300. In some examples, the type can be "data element", "transformation", etc. In some examples, the type field 312 also indicates the number of forward and backward edges included within an instance of the data structure 300. The property field 314 represents various properties of the metadata object represented by the corresponding instance of the data structure 300 respectively. Examples of the property field 314 include a "name" field containing a text label identifying the metadata object, and a "subtype" field indicating the subtype of the metadata object, for example, whether the metadata object represents a file object, an executable file object, a database object, or another subtype. Other types of properties can also be used. Generally, the type field 312 and the property field 314 can be customized for a specific instance of the system server 102 and are not limited to the examples given herein.

[0044] The data structure also includes fields representing forward edges 316A - C and backward edges 316D - F. These edge fields 316A - F enable the system server 102 to collect data from the data structure as it "walks" from data structure to data structure to gather system metadata. In the broadest sense, referring to "collecting" partial data means identifying a portion of the data as relevant to future actions (such as transmitting the collected data). Collecting a portion of the data may include duplicating the data, for example, duplicating the data into a buffer or queue for future use.

[0045] Each of the edge fields 316A - F includes pointer fields 320A - B. The pointer fields 320A - B store the addresses of respective memory locations 322A - B. Generally, the memory locations 322A - B referred to by the pointer fields 320A - B point to the portion of memory that stores another instance of the data structure 300. In this way, one instance of the data structure representing a metadata object is "linked" to one or more other instances of the data structure representing other metadata objects. Thus, edges 316A - D may correspond to the relationships between metadata objects shown, for example, in examples 200A - E of the lineage diagrams of FIGS. 2A - 2E. For example, forward edge 316A represents the influence that a metadata object (such as the metadata object represented by this instance of the data structure 300) has on another metadata object (such as the metadata object represented by an instance of the data structure at memory location 322A). As another example, backward edge 316D represents the influence that another metadata object (such as the metadata object represented by an instance of the data structure at memory location 322B) has on the metadata object of this instance of the data structure 300.

[0046] Each edge field 316A - F also includes one or more flags 324. The flag 324 is an indicator of information regarding its associated edge. In some examples, one of the flags 324 may indicate the type of the associated edge selected from a plurality of possible types. There can be many types of edges. For example, some types of edges are input / output edges (representing output from one object and input to another object), element / dataset edges (representing the relationship between an element and the dataset to which the element belongs), and application / parent edges (representing the relationship between an executable application and a container such as a container that also includes a dataset related to the application).

[0047] Many of the elements of the data structure 300 typically use a relatively small amount of data. For example, the data associated with the identification information field 310, the type field 312, and the property field 314 may together be only a few bytes, e.g., 32 bytes. These fields typically encode information used in only a few bits. For example, if there are only 8 possible types for a node, the type field 312 can be only 3 bits long. More complex data such as a text string representing the type of the node need not be used. Further, the data associated with the memory locations 322A - C is typically the same amount of data as the length of the memory address related to the type of computer system that executes the software that instantiates the data structure 300. Thus, compared to the data used by other techniques for storing lineage metadata, most or all instances of the data structure 300 can use a relatively small amount of data in total.

[0048] Figure 4 shows an example of a walk plan 400. As described above with respect to Figure 1C, the walk plan 400 is typically stored by the metadata server 104. In use, the metadata server 104 provides the walk plan to the lineage server 102 when requested for lineage metadata.

[0049] The walk plan 400 describes information used by the lineage server 102 when traversing the stored data structure 134 of the lineage server 102. Generally, when receiving a query for lineage metadata, for example, lineage metadata related to a specific metadata object, it is not necessary to return all types of lineage metadata in response. In some cases, depending on the query, it may not be necessary to return because some types of lineage metadata related to some edges are not in response to the query.

[0050] Accordingly, the walk plan 400 includes records 402A - C for each type of edge that can be among the types of edges represented by the lineage metadata stored by the lineage server 102. Record 402A includes an edge type field 404 that contains data indicating the type of edge corresponding to record 402A. Record 402A also includes a follow flag 406, a node collection flag 408, and an edge collection flag 409 with respect to the forward direction 410, and a follow flag 412, a node collection flag 414, and an edge collection flag 415 with respect to the backward direction 416.

[0051] The follow flags 406, 412 indicate whether the lineage server 102 should follow edges of this type of edge when traversing its own data structure 134. In other words, the follow flag 406 in the forward direction 410 indicates whether the lineage server 102 should access the memory location 322A identified by the pointer field 320A of the forward edge field 316A of the instance of the data structure 300 with reference to FIG. 3. Similarly, the follow flag 412 in the backward direction 416 indicates whether the lineage server 102 should access the memory location 322b identified by the pointer field 320b of the backward edge field 316D of the instance of the data structure 300 with reference to FIG. 3.

[0052] The node collection flags 408, 414 indicate whether the system server 102 should collect an instance of a data structure 300 (FIG. 3), sometimes called a "node", pointed to by the type of edge when the system server 102 traverses its own data structure 134. When referring to collecting an instance of the data structure 300, it means that data of the instance (or node) is added to the data returned in response to the system server 102 (FIG. 1A) processing a query. Thus, when a node is collected, data related to the metadata object represented by the instance of the data structure 300 is among the system metadata returned by the system server 102.

[0053] The edge collection flags 409, 415 indicate whether the system server 102 should collect an edge (e.g., corresponding to the pointer field 320A of an instance of the data structure 300). When an edge is collected, the data representing the edge is among the system metadata returned by the system server 102. In some implementations, if the edge does not represent a data flow between nodes, the edge may not need to be collected. For example, an edge can represent a relationship between a data object (represented by a node) and a container of a data object (represented by another node). In this way, by using the node collection flags 408, 414 and the edge collection flags 409, 415 within the walk plan 400, nodes that may or may not be collected for inclusion in the system metadata can be related to each other in various ways, and the nodes can represent various data that may or may not be collected for inclusion in the system metadata.

[0054] In some implementations, during use, the walk plan 400 can be represented in one or more XML (Extensible Markup Language) document formats. An XML document is a collection of parts separated by "tags". A tag typically includes a label (e.g., a label identifying the type of tag) and may also include one or more attributes. Tags are in the form of start tags and end tags and may be provided so that a start tag is paired with a corresponding end tag. In this way, tags can be hierarchical, so that for example, by placing a tag between the start tag and end tag pair of another tag, the tag is "nested" within another tag.

[0055] An example of a walk plan in XML document format is shown below: <lineageserverplan direction=""both”conditionalOnArg="!autoFilterEnabled”replacesQueries="walk”"> <useedge name=""DE-Tr”"> <condition special=""ExelnterfaceCallStack”" / > <condition special=""ControlFilter”" / > <condition special=""Summarization”" / > < / useedge> <useedge name=""Tr-DE”"> <condition special=""ExelnterfaceCallStack”" / > <condition special=""ControlFilter”" / > <condition special=""Summarization”" / > < / useedge> <useedge name=""DE-DS”direction="forward”collectEdge="false”" conditionalOnArg=""walkDSlevel”" / > <useedge name=""Tr-Exe”direction="forward”collectEdge="false”" conditionalOnArg=""walkDSlevel”" / > <useedge name=""DS-Exe”conditionalOnArg="walkDSlevel”"> <condition special=""ExeInterfaceCallStack”" / > <condition special=""DSLevelIfNoDE”" / > < / useedge> <useedge name=""Exe-DS”conditionalOnArg="walkDSlevel”"> <condition special=""ExeInterfaceCallStack”" / > <condition special=""DSLevelIfNoDE”" / > < / useedge> <useedge name=""DE-DS”direction="backward”collectEdge="false”"> <condition special=""DSLevelIfNoDE”" / > < / useedge> <useedge name=""Tr-Exe”direction="backward”collectEdge="false”"> <condition special=""DSLevelIfNoDE”" / > < / useedge>

[0056] In this example, the "useEdge" tag specifies information about a given type of edge. Each "useEdge" tag can correspond to a record (e.g., records 402A - C of the walk plan 400). The "name" attribute specifies the type of edge (e.g., edge type 404), the "direction" attribute specifies the direction (e.g., forward direction 410 or backward direction 416), and the "collectEdge" attribute specifies whether to collect the edge (e.g., collection flags 408, 414). Other tags can also be used. For example, the "condition special" tag shown in the above example is used to specify a custom rule that is executed when an edge of the specified edge type is traversed. In some examples, the custom rule can specify conditions for determining whether an edge should be traversed and / or collected.

[0057] FIG. 5A shows a flowchart representing a procedure 500 for storing lineage metadata in a format defined by a dedicated data structure, such as the data structure 300 shown in FIG. 3. The procedure 500 can be executed by components of the lineage server 102 shown in FIG. 1A, for example.

[0058] This procedure requests lineage metadata from a metadata source (502). For example, the metadata source can be the metadata database 108 shown in FIG. 1A. This request can be made at regular or semi-regular intervals, such as hourly, every 10 minutes, every minute, or any other interval. In some examples, this request can be made in response to an event, such as a notification that new metadata is available at the metadata source.

[0059] Lineage metadata typically describes nodes and edges such that each node represents a metadata object, and the edges each represent a unidirectional influence of one node on another node, and thus each edge has a single direction.

[0060] In some examples, for instance, if the lineage server 102 has not yet generated an initial set of data structures representing the lineage metadata, this request is for all the lineage metadata stored by the data source. In some examples, for instance, if the lineage server 102 is updating an existing set of stored data structures, this request is for the lineage metadata that has been added or changed since the previous request.

[0061] This procedure receives data, such as lineage metadata, from the metadata source (504). For example, the lineage metadata can be data representing metadata objects and the relationships between the metadata objects.

[0062] This procedure generates (506) an instance of a data structure, such as the data structure 300 shown in FIG. 3. For example, the data structure may include information corresponding to data received from a metadata source. In some examples, each instance of the data structure corresponds to a respective node received from the metadata source. The data structure may include a field for an identification value, such as an identification value that identifies the node corresponding to the instance of the data structure. The data structure may also include a property field that represents a property of the node corresponding to the instance of the data structure. The data structure may also be capable of pointers to the identification values of other nodes, such that the pointers represent edges to the nodes corresponding to the respective identification values.

[0063] This procedure stores (508) the data structure. For example, the data structure can be stored in a memory, such as the memory 135 shown in FIG. 1C. In some examples, the data structure is stored in a random access memory. Since the data structure is used to store lineage metadata, any data not related to lineage (e.g., other types of metadata stored in the metadata source) can be omitted, reducing the amount of data required to store the data structure.

[0064] During use, this procedure returns to requesting (502) lineage metadata from the metadata source, for example, at the next periodically scheduled interval.

[0065] FIG. 5B shows a flowchart representing a procedure 520 for displaying lineage metadata. Procedure 520 may be performed by components of the lineage server 102 shown in FIG. 1A, for example. Generally, the lineage server is configured to return a response to a query that includes metadata describing the lineage of a particular data element, such as a metadata object. In some examples, the metadata describes a sequence of nodes and edges, and one of the nodes in the sequence represents a particular data element. In some examples, procedure 520 is used to access the lineage metadata stored by procedure 500 described above with respect to FIG. 5A.

[0066] This procedure receives (522) a query, e.g., a query for lineage metadata. In some examples, this query identifies the metadata object for which lineage metadata is requested.

[0067] In some implementations, this query includes a walk plan that identifies the type of lineage and which types of edges are associated with the identified type of lineage. In some examples, the walk plan includes conditions for traversing or collecting edges based on one or more characteristic values that represent respective characteristics of the corresponding nodes. An example of a walk plan 400 is shown in FIG. 4.

[0068] This procedure collects (524) lineage metadata. For example, it can access and collect the nodes representing the metadata objects of the received query, and can traverse edges (e.g., pointers to memory locations) to collect other nodes. The collection of lineage metadata will be described in more detail below with respect to FIG. 6.

[0069] This procedure transmits (526) the collected lineage metadata. For example, the collected lineage metadata can be transmitted to the computer system that issued the query.

[0070] After transmission of the collected lineage metadata, the collected lineage metadata may be displayed (528) on a computer system, e.g., on the user terminal 116 shown in FIG. 1A. For example, the lineage metadata can be displayed in the form of a lineage diagram such as the lineage diagrams 200A - 200E shown in FIGS. 2A - 2E.

[0071] FIG. 6 shows a flowchart representing a procedure 600 for traversing lineage metadata stored in the form of an instance of a dedicated data structure, e.g., the data structure 300 shown in FIG. 3. The procedure 600 can be executed by components of the lineage server 102, e.g., as shown in FIG. 1A.

[0072] This procedure receives (602) a query and a walk plan, such as query 114 and walk plan 130 shown in FIG. 1C. This procedure accesses (604) an initial node (e.g., an instance of data structure 300 shown in FIG. 3) that represents the metadata object referenced by query 114. For example, the initial node can be identified by identification information field 310 (FIG. 3) that stores data related to the metadata object. Then that initial node is used as the "current" node, and the recursive part of this process begins where the current node is selected from the queue and an operation is applied to the current node. In other words, the initial node is placed in the queue as the first node in the queue, and as this procedure is executed, other nodes are subsequently added to the queue.

[0073] This procedure determines (606) whether there is a forward edge pointer remaining in the current node (e.g., a forward edge pointer that has not yet been accessed). If there is such a forward edge pointer, this procedure accesses (608) the next pointer that has not yet been accessed, e.g., accesses the memory location of the pointer and retrieves the data stored at that memory location. This procedure determines (610) whether to "walk" (e.g., process) the node at that pointer based on the walk plan (described above with respect to FIG. 4) according to the type of edge associated with the pointer. If not walking, this procedure accesses (608) another pointer. If walking, this procedure determines (611) whether to collect the node at that pointer. If collecting, this procedure stores (612) the data of the node to be returned in response to the query, and then places (614) that node in the queue so that the pointer can be accessed. If not collecting, this procedure simply places the node in the queue.

[0074] When all the forward edge pointers within the current node have been traversed, this procedure determines (616) whether there are any backward edge pointers remaining within the current node. If there are such backward edge pointers, this procedure accesses (608) the next backward edge pointer.

[0075] If there are no forward or backward edge pointers remaining, this procedure determines (618) whether there are any nodes remaining in the queue. If there are, this procedure accesses (620) the next node in the queue, and uses the next node in the queue as the current node to perform the above operations. If there are no nodes remaining, this procedure prepares (622) the data that has been collected for transmission to another system. For example, the collected data can be configured in a specific format since it is to be transmitted. As another example, the encoded data within the collected data can be decoded. For example, a data field containing an encoded value can be converted to a text string corresponding to the value.

[0076] When the data is ready for transmission, it can be transmitted as described, for example, with respect to data transmission 526 of FIG. 5B.

[0077] The systems and techniques described herein can be implemented, for example, using a programmable computing system that executes appropriate software instructions, can be implemented by appropriate hardware such as a field programmable gate array (FPGA), or can be implemented in some hybrid form. For example, in a programmed approach, software can include procedures within one or more computer programs that are executed on one or more programmed computing systems or programmable computing systems (which can be of various architectures such as distributed, client / server, or grid). Such a computing system includes at least one processor, at least one data storage system (including volatile and / or non-volatile memory and / or storage elements), and at least one user interface (for receiving input using at least one input device or port and for providing output using at least one output device or port). The software can include, for example, one or more modules of a larger program that provides services related to the design, configuration, and execution of data flow graphs. The program modules (e.g., elements of a data flow graph) can be implemented as data structures or other organized data that conform to a data model stored in a data repository.

[0078] Software can be stored in a non-transitory form using the physical characteristics of a medium (such as surface pits and lands, magnetic domains, charges) over a period of time (e.g., the time during the refresh period of a dynamic memory device such as dynamic RAM) in a volatile memory medium, a non-volatile memory medium, or any other arbitrary non-transitory medium. In preparation for loading the instructions, the software can be provided on a tangible non-transitory medium such as a CD-ROM or other computer-readable medium (e.g., readable by a general-purpose or special-purpose computing system or computing device), or can be delivered over a network communication medium to a tangible non-transitory medium of the computing system on which the software is to be executed (e.g., encoded in a propagated signal). Some or all of the processing can be executed on a dedicated computer or using dedicated hardware such as a coprocessor, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc. The processing can be implemented in a distributed manner, in which case different parts of the calculations specified by the software are executed by different computing elements. Each such computer program, when the storage device medium is read by a computer to configure and operate the computer for executing the processes described herein, is preferably stored on or downloaded to a computer-readable storage medium of a storage device accessible by a general-purpose or special-purpose programmable computer (such as a solid-state memory or medium, a magnetic medium, an optical medium). The system of the present invention may be implemented as a tangible non-transitory medium composed of a computer program, and the medium so configured causes the computer to operate in a specific and predefined manner to execute one or more of the processing steps described herein.

[0079] Some embodiments of the present invention have been described. Nevertheless, it should be understood that the above description is for the purpose of illustration rather than limiting the scope of the present invention as defined by the appended claims. Accordingly, other embodiments are also within the scope of the following claims. For example, various modifications can be made without departing from the scope of the present invention. In addition, some of the above steps may not be order-dependent and can therefore be executed in a different order than described.< / lineageserverplan>

Claims

1. A method executed by a data processing system for updating a data structure having a lineage from a database, the updating comprising enhancing the efficiency of determining the lineage, the method comprising: The method comprises: storing in a memory of the data processing system a data structure that identifies a lineage of a given data item or a given transformation, the lineage identifying data items or transformations that affect or are affected by the given data item or the given transformation; occasionally receiving from the database data representing the lineage of the given data item or the given transformation; identifying, based on the received data, data items or transformations that affect or are affected by the given data item or the given transformation; updating the data structure to have a reference to a memory location of another data structure representing the identified data items or the identified transformations; updating the lineage specified by the data structure; A method comprising the above.

2. The method according to claim 1, wherein updating the data structure comprises updating the data structure to have a pointer to a memory location of another data structure representing the identified data items or the identified transformations.

3. generating a second data structure that identifies the lineage of the identified data items or the identified transformations; storing the second data structure in the memory of the data processing system; The method according to claim 1, further comprising the above.

4. The method according to claim 1, further comprising receiving the data representing the lineage of the identified data items or the identified transformations from the database at regular intervals or at scheduled intervals.

5. The method according to claim 1, wherein the data representing the lineage of the identified data items or the identified transformations is received in response to a request transmitted to the database.

6. receiving a query including a lineage request; accessing the data structure and at least one other data structure to determine the requested lineage based on the query; generating a response to the query, the response including data representing the requested lineage; The method according to claim 1, further comprising the above.

7. accessing a walk plan that includes an instruction to determine the required lineage; collecting data from the data structure and the at least one other data structure according to the walk plan to determine the required lineage; The method according to claim 6, further comprising.

8. A system for updating a data structure having a lineage from a database, the updating enhancing the efficiency of determining the lineage, the system comprising: the system is at least one processor; at least one computer-readable medium storing instructions executable by the at least one processor to perform operations, the operations are storing in memory a data structure that identifies a lineage of a given data item or a given transformation, the lineage identifying data items or transformations that affect or are affected by the given data item or the given transformation; occasionally receiving from the database data representing the lineage of the given data item or the given transformation; identifying, based on the received data, data items or transformations that affect or are affected by the given data item or the given transformation; updating the data structure to have a reference to a memory location of another data structure representing the identified data items or the identified transformations; updating the lineage specified by the data structure; A system comprising.

9. The system according to claim 8, wherein updating the data structure includes updating the data structure to have a pointer to a memory location of another data structure representing the identified data items or the identified transformations.

10. The at least one computer-readable medium stores instructions executable by the at least one processor to perform operations, the operations are generating a second data structure that identifies the lineage of the identified data items or the identified transformations; storing the second data structure in the memory; The system according to claim 8, further comprising.

11. The at least one computer-readable medium stores instructions executable by the at least one processor to perform operations, The operation The system of claim 8, further comprising receiving, at regular intervals or at scheduled intervals, the data representing the identified data item or the identified sequence of transformations from the database. Claim 12 The system of claim 8, wherein the data representing the identified data item or the identified sequence of transformations is received in response to a request transmitted to the database. Claim 13 The at least one computer-readable medium stores instructions executable by the at least one processor to perform an operation, The operation receiving a query including a query for a sequence, accessing the data structure and at least one other data structure to determine the requested sequence based on the query, generating a response to the query, the response including data representing the requested sequence, The system of claim 8, further comprising. Claim 14 The at least one computer-readable medium stores instructions executable by the at least one processor to perform an operation, The operation accessing a walk plan including instructions to determine the requested sequence, collecting data from the data structure and the at least one other data structure according to the walk plan to determine the requested sequence, The system of claim 13, further comprising. Claim 15 At least one non-transitory computer-readable medium storing instructions executable by at least one processor to perform an operation, The operation storing in memory a data structure that identifies a given data item or a sequence of transformations for a given transformation, the sequence identifying data items or transformations that affect or are affected by the given data item or the given transformation, occasionally receiving from a database data representing the sequence of the given data item or the given transformation, identifying, based on the received data, data items or transformations that affect or are affected by the given data item or the given transformation, Updating the data structure having a reference to a memory location of the identified data item or another data structure representing the identified transformation, updating the lineage specified by the data structure, and at least one non-transitory computer-readable medium including the same. **Claim 16** The updating of the data structure includes updating the data structure having a pointer to a memory location of the identified data item or another data structure representing the identified transformation, the at least one non-transitory computer-readable medium according to claim 15. **Claim 17** The at least one non-transitory computer-readable medium stores instructions executable by the at least one processor to perform operations, the operations including generating a second data structure specifying a lineage of the identified data item or the identified transformation, storing the second data structure in the memory, and the at least one non-transitory computer-readable medium according to claim 15 further including the same. **Claim 18** The at least one non-transitory computer-readable medium stores instructions executable by the at least one processor to perform operations, the operations including further receiving, at regular intervals or at scheduled intervals, the data representing a lineage of the identified data item or the identified transformation from the database, the at least one non-transitory computer-readable medium according to claim 15. **Claim 19** The data representing a lineage of the identified data item or the identified transformation is received in response to a request transmitted to the database, the at least one non-transitory computer-readable medium according to claim 15. **Claim 20** The at least one non-transitory computer-readable medium stores instructions executable by the at least one processor to perform operations, the operations including receiving a query including a lineage request, accessing the data structure and at least one other data structure to determine the requested lineage based on the query, and generating a response to the query, the response including data representing the requested lineage, and the at least one non-transitory computer-readable medium according to claim 15 further including the same.

Citation Information

Patent Citations

  • Graphic representation of data relationships

    JP2011517352A

  • Data management program, node, and distributed database system

    JP2013077063A

  • Updating cached database query results

    JP2015531129A

  • Related Information Dissemination System

    JP2015537294A

  • Data lineage transformation analysis

    US20150012478A1