Data object graph enhancement framework
By identifying the relationship between the data object in the lineage diagram and the candidate data object, the problem of inefficient identification of complex data object relationships in the prior art is solved, and more efficient data processing and potential relationship discovery are achieved.
Patent Information
- Application Number
- CN202411651695.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-21
- Filing Date
- 2024-11-19
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to effectively identify and understand the relationships between complex data objects, resulting in inefficiency in processing and undiscovered potential relationships.
By providing a technique and solution, the relationship between the data object in the lineage graph and the candidate data object can be identified, and the relationship can be established or suggested by comparing attribute types, attribute values or related information to determine whether the relationship criteria are met.
It realizes effective identification and understanding of the relationships between data objects, improves processing efficiency, and may discover potential useful relationships.
Smart Images

Figure CN120030065A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to identifying relationships between data objects. Background Art
[0002] Software is increasingly integral to business processes, from manufacturing processes including supply chain management to logistics processes, sales processes, planning and accounting processes. Large amounts of data may be involved, which may be distributed across thousands of data objects, each of which often has hundreds of individual attributes. Data objects often have complex interrelationships, where for example a higher-level data object may contain data from one or more lower-level data objects. Further complexity is introduced by the rapid growth of data volumes, and because the requirements placed on business processes may change over time.
[0003] Data objects can be distributed across different platforms. For example, data can be stored in a relational database or in an unstructured or semi-structured format, such as an "object store" that can store data in JSON (or more generally, CSON) format, or in a flexible NoSQL database. Other types of data can include unstructured text or media files, such as images, audio, or video.
[0004] As another example, data from a transactional system (such as using OLTP) can be transformed and used in an analytical (OLAP) scenario. "ETL" processes and other types of import / export functions can be used to share data between systems. The design of ETL processes can be complex, especially when ensuring data accuracy and consistency at the same time. Computing systems can have business processes involving data flows within a given system and between systems.
[0005] Data can be stored in various "layers", such as a relational database layer (physical storage) referenced by a virtual data model (such as CDS views or ABAP views of SAP SE in Walldorf, Germany), representing an intermediate level of abstraction, while logical data objects (such as BUSINESSOBJECTS of SAP SE) are used at a higher level of abstraction. Higher-level data objects can provide representations of semantically meaningful documents that are easy to process, such as a logical data object representing a sales order, where the data object can provide access to data associated with the sales order and provide operations that can be performed with or on the sales order.
[0006] Understanding relationships between data can be a difficult task. Relationships between data may have been created for a specific purpose, but a lack of overall understanding of the data can lead to inefficient processing. This lack of understanding can be particularly problematic considering that business processes often benefit from ongoing monitoring and adjustments. Furthermore, because data relationships are often enforced on the data, there may be latent relationships between the data that, if discovered, may actually be useful. Therefore, there is room for improvement. Summary of the invention
[0007] This summary is provided to introduce some concepts in a simplified form, which are further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0008] Techniques and solutions for identifying relationships between data objects are provided. In particular, the present disclosure provides techniques that can identify data objects related to data objects in a lineage graph and can be used to identify relationships between existing lineage graph data objects. A lineage data object is compared to a candidate data object. For example, attribute types, attribute values, or information about these attribute types or values can be compared between the lineage data object and the candidate data object. If it is determined that the relationship criteria are met, a relationship can be established or suggested between the lineage data object and the candidate data object. A display of a lineage graph showing the relationship between the lineage data object and the candidate data object is rendered. A user can define edges between the lineage data object and the candidate data object.
[0009] In one aspect, the present disclosure provides a process for identifying at least an inferred relationship between a lineage data object and a candidate data object. A definition of a lineage graph of a plurality of lineage data objects is received. The definition includes identifiers of lineage data objects in the plurality of lineage data objects forming nodes of the lineage graph, and identifiers of relationships between the plurality of lineage data objects forming edges of the lineage graph, wherein the lineage data objects include one or more attributes.
[0010] Metadata of a plurality of candidate data objects is received. The metadata includes a candidate data object identifier and an identifier of an attribute defined for the corresponding candidate data object. A lineage data object and a candidate data object are selected. The lineage data object is compared with the candidate data object.
[0011] A determination is made that the lineage data object and the candidate data object satisfy relationship criteria. At least an inferred relationship between the candidate data object and the lineage data object is established. An updated lineage graph including at least the inferred relationship is rendered for display.
[0012] The present disclosure also includes computing systems and tangible, non-transitory computer-readable storage media configured to perform the above-described methods or including instructions for performing the above-described methods. As described herein, various other features and advantages may be incorporated into the technology as desired. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 A diagram of a database schema that shows the technical relationships between at least a portion of the database tables in the schema.
[0014] Figure 2 A graph of a computing environment in which the disclosed technology may be implemented, wherein a graph enhancement framework may be used to identify relationships between objects in a lineage graph and candidate data objects, and relationships between objects or candidate data objects in a lineage graph.
[0015] Figure 3 An example lineage graph and a graph enhancement framework including information for a plurality of candidate data objects are shown.
[0016] Figure 4 is a diagram illustrating how properties of a lineage data object may be compared to properties of a candidate data object, wherein these techniques may also be used to compare two lineage data objects or two candidate data objects.
[0017] Figure 5 Example pseudo code for a process of comparing a table in a lineage graph to a table corresponding to a candidate data object is provided.
[0018] Figure 6 A lineage graph including suggested relationships to candidate data objects is shown, including indicators providing relative strength or confidence of the relationships.
[0019] Figure 7 is a flow chart of a process for enhancing a lineage graph with candidate data objects.
[0020] Figure 8 is a flow chart of a process for identifying at least an inferred relationship between a lineage data object and a candidate data object.
[0021] Fig. 9 is a diagram of an example computing system in which some described embodiments may be implemented.
[0022] Fig.10 is an example cloud computing environment that can be used in conjunction with the techniques described herein. DETAILED DESCRIPTION
[0023] Example 1 - Overview
[0024] Software is increasingly integral to business processes, from manufacturing processes including supply chain management to logistics processes, sales processes, planning and accounting processes. Large amounts of data may be involved, which may be distributed across thousands of data objects, each of which often has hundreds of individual attributes. Data objects often have complex interrelationships, where for example a higher-level data object may contain data from one or more lower-level data objects. Further complexity is introduced by the rapid growth of data volumes, and because the requirements of the business may change over time.
[0025] Data objects can be distributed across different platforms. For example, data can be stored in a relational database or in an unstructured or semi-structured format, such as an "object store" that can store data in JSON (or more generally, CSON) format, or in a flexible NoSQL database. Other types of data can include unstructured text or media files, such as images, audio, or video.
[0026] As another example, data from a transactional system (such as using OLTP) can be transformed and used in an analytical (OLAP) scenario. "ETL" processes and other types of import / export functions can be used to share data between systems. The design of ETL processes can be complex, especially when ensuring data accuracy and consistency at the same time. Computing systems can have business processes involving data flows within a given system and between systems.
[0027] Data may be stored in various "layers", such as a relational database layer (physical storage) referenced by a virtual data model (such as CDS views or ABAP views of SAP SE in Walldorf, Germany), representing an intermediate level of abstraction, while logical data objects (such as BUSINESSOBJECTS of SAP SE) are in turn used at a higher level of abstraction. Higher level data objects may provide convenience for handling representations of semantically meaningful documents, such as a logical data object representing a sales order, where the data object may provide access to data associated with the sales order and provide operations that may be performed with or on the sales order.
[0028] Understanding relationships between data can be a difficult task. Relationships between data may have been created for a specific purpose, but a lack of overall understanding of the data can lead to inefficient processing. This lack of understanding can be particularly problematic considering that business processes often benefit from ongoing monitoring and adjustments. Furthermore, because data relationships are often enforced on the data, there may be latent relationships between the data that, if discovered, may actually be useful. Therefore, there is room for improvement.
[0029] The present disclosure provides techniques and solutions that can be used to discover relationships between data stored in data objects. As used herein, a data object refers to a collection of related attribute values stored as a collection. A collection can be an "object" in the sense of being an instance of an abstract or composite data type. However, an "object" also refers to a data set that may not have a strictly enforced structure. For example, a "data object" includes information stored in a JSON object or an XML document.
[0030] In some scenarios, the relationships between data and its associated data objects can be used to improve data processing, such as reducing the amount of data processed or reducing the number of processing steps. In addition, these relationships can suggest new insights that can be obtained using already existing data.
[0031] There is usually a data model or schema for a particular data or data set. As described, this can be part of a formal "definition", or can be provided in computing code, i.e., code that provides a JSON object or XML document in a defined manner. Thus, a "data model" or "schema" can refer to an organization of objects in code that is not part of an explicit object definition in the sense of an abstract or composite data type. A data model can be used for a single abstraction layer (which may include physical storage), or for a single data source, such as a specific database system. A data model can be defined with respect to multiple data sources or abstraction layers.
[0032] A data model includes data objects, such as tables or views of a relational database, or entities or views of a virtual data model. As discussed, data typically "flows" between data objects based on the relationships between them. For example, a view can be defined as a projection of a specific data object, or as a combination of data from multiple data objects, such as a join operation in a relational database. For a particular data model, some data objects may be related to each other, while other data objects may not be related. Even when data objects are related, they may have a direct relationship or an indirect relationship. For example, a view may be built on a lower-level view, which in turn is defined with respect to one or more tables.
[0033] Relationships between data objects can be of various types. For example, a view can be defined with respect to selection conditions (including join operations), where the selection or join involves specific attributes. Foreign key relationships can be used for various purposes, including facilitating joins, and also for purposes of enforcing referential integrity, enabling cascading operations, improving query performance (where the relationship provides a hint to the query optimizer about related objects), facilitating understanding of schemas, or enforcing data rules (such as data rules associated with computational processes (such as computational processes that implement business processes)), or for controlling access to data.
[0034] The data model contains data objects such as tables or views of a relational database, or entities and views within a virtual data model (VDM). Within the realm of the VDM, associations serve as predefined links between entities or views, encapsulating business logic and common access paths, while calculated fields derived from the attributes of VDM entities can also create implicit relationships. Direct mappings define how data in the physical database layer is acquired and presented in the VDM. Additionally, annotations as an overlay of metadata provide insights into how VDM entities or views are associated with database objects.
[0035] The transformation between VDM and business objects is usually facilitated by service exposure, where entities in the VDM can be exposed as consumable services of business objects. Business semantics extracted from annotations or metadata within the VDM provide interpretations that business objects can then use to process data.
[0036] For ETL processes and inter-system data flows, the transformation logic during ETL operations dictates how data from source tables should be shaped, aggregated, or enriched, creating a link between source and transformed data. The keys and indexes introduced after extraction are consistent with the source system, maintaining a relationship with the transformed data. ETL operations often use audit logs that track data lineage, forming a connection between the original data and its final location after loading.
[0037] Outside of ETL, data replication tools that copy data from one system to another inherently form links between the source and target, typically maintained via replication logs or timestamps. When systems exchange data via APIs, the subsequent request-response dynamics supported by unique identifiers such as transaction IDs create explicit links between data objects in those systems. Messaging protocols, where systems can offload data to topics or queues for retrieval by other systems, exhibit publisher-consumer relationships, establishing links between data objects. In some database scenarios, objects in one system can be related to objects in an entirely different system using foreign key references.
[0038] In data modeling, schema mappings act as guides for the transformation and alignment between source and target schemas. These mappings, arranged in the form of transformation rules or transformation functions, connect different data models. In relational databases, they determine how the tables, columns, and relationships in the source schema relate to the tables, columns, and relationships in the target schema. For example, a column with a specific data type and constraints in one database may correspond to another column in a different system with adjustments to its definition or constraints, all governed by these mapping rules.
[0039] Schema mappings are not limited to tables and columns, but also provide logic for the transformation of stored procedures, views, and triggers. Entities in one schema can undergo transformation based on these mappings to appear as views in another schema, possibly combined or adjusted based on business needs.
[0040] In more complex scenarios such as object-relational mapping (ORM), schema mapping acts as an intermediary between object-oriented application code and a relational database. In this context, class definitions and object structures in the application are mapped to tables, rows, and columns in the database. This mapping affects how data is stored, retrieved, and aspects such as data integrity and relationship enforcement.
[0041] When dealing with data lakes or non-relational storage systems, schema mapping helps convert structured data sets to flat files or convert flat files to structured data sets and create connections between hierarchical data formats (such as JSON or XML) and tabular database structures. Schema mapping can be used in data migration and integration projects, especially when connecting various data models. Their detailed design includes conditional logic, data type conversion, and exception handling, which helps provide data consistency and accuracy across different systems.
[0042] In one aspect, the present disclosure uses a set of relationships as a starting point to identify new relationships between objects in the set of relationships and between other objects not connected by relationships.While various types of relationship representations can be used with the techniques of the present disclosure, specific scenarios are described with respect to lineage graphs.
[0043] In the field of data management, lineage diagrams are used as a visualization tool that illustrates the trajectory and transformation of data as it flows through different systems and processes. A lineage diagram visually tracks data from its origin or from a source to its final destination or target, highlighting the transformation steps in between. As used herein, a lineage diagram is defined as a collection of data objects from multiple different software abstraction layers or multiple data sources, wherein at least a portion of the edges between data objects in the lineage diagram connect data objects from different software layers or data sources, and wherein at least a portion of the edges indicate specific data transformations performed as the data flows through the edges.
[0044] In some cases, data sources may be "remote" or otherwise operationally separated from one another. As a specific example, one data source may be able to obtain data from other data sources using data federation. Thus, an edge represents a relationship between data objects, but may also represent a "barrier" between different data sources, including a "physical barrier" where the data sources may be connected via a network.
[0045] Within a lineage graph, nodes typically represent various data objects. They can range from tables in a relational database to entities in a virtual data model, or even more abstract constructs like business objects. On the other hand, the edges connecting these nodes depict the relationships or operations acting on the data, showing the paths of its flow and transformation.
[0046] Take, for example, a scenario involving a relational database. A lineage graph may depict tables as primary nodes. Edges between these nodes may indicate operations such as SQL join conditions or foreign key constraints. When a view is introduced that aggregates or shapes data from multiple tables, it may also appear as a node connected to its constituent tables with edges indicating the nature of the data transformation or selection.
[0047] In a lineage graph, edges serve as representations of data flows and transformations between different data objects. Each edge has an origin as its starting point and a destination marking its endpoint. This establishes connections between specific data objects. In addition, edges often carry transformation metadata that clarifies the nature of the data manipulation. This includes the type of transformation—whether it is aggregation, filtering, joining, or calculation—as well as detailed details such as “aggregate using SUM” or “filter where age>30”.
[0048] Some edges can also include time information, such as indicating the exact time when the data was moved or transformed and how often such an operation occurred. Edges can have an associated confidence or quality score, especially in automatic lineage graphs. This score provides insight into the system's certainty about a particular transformation.
[0049] Another attribute that can be included as edge information is data volume, which indicates the amount of data that has been transferred from the source to the target. Operational metadata attached to the edge can also indicate the tool or system used for the transformation, such as a specific ETL tool or database, and pinpoint the party responsible for the data flow. The edge can also indicate documentation or annotations for added context and clarity about the relationship or transformation.
[0050] The visual depiction of the edge, its type or style may also be included as edge information and may vary based on its importance. For example, a solid line may represent direct data flow, while a dashed line may represent a less frequent or periodic flow.
[0051] As a more concrete example, the following code represents a sample Python edge class for a lineage graph:
[0052] )
[0054] print(edge1)#Outputs:From TableA to ViewX via join:joined on column_ID
[0055] print(edge2)#Outputs:From ViewX to FinalView via filter:filteredwhere value>10
[0056] When the virtual data model is also considered, the complexity of the lineage graph may grow. Nodes representing entities in the VDM may be linked to nodes in the underlying relational database. These links or edges may indicate mapping or transformation logic that directs data from physical storage devices to a more abstract logical layer.
[0057] Incorporating BusinessObjects into lineage diagrams, these high-level logical constructs may appear as distinct nodes, linked to VDM entities or even directly to database tables. Connecting edges can represent not only data flows, but also operational interactions—such as how a Sales Order BusinessObject can retrieve, modify, or store data in an associated table or entity.
[0058] With the introduction of ETL processes, further complexity can arise, which can be depicted as transformation nodes or edges on a lineage graph. Data flowing from one system, through the ETL process, and received by another system can be represented visually, capturing the transformations, filtering, or aggregations applied during the data migration.
[0059] Schema mappings can also be represented in lineage diagrams, especially in scenarios involving data integration or migration. These mappings, which define the dependencies between source and target schemas, can be shown as dedicated connectors or annotated edges, helping users understand how data attributes in one system align with those in another.
[0060] Information about another set of data objects, referred to as candidate objects in a candidate set, may be obtained, such as objects in schemas and data sources that are sources of objects and relationships, or from schemas and data objects that are not currently used in a lineage graph. Information about a candidate set of candidate data objects may include relationships between candidate data objects, candidate object types, candidate object metadata, and candidate object data in the candidate set. For example, attributes and attribute data types may be identified for a particular candidate data object. Data values associated with instances of candidate data objects may be retrieved.
[0061] Data objects of a lineage graph (referred to as lineage data objects) may be compared to candidate data objects. For example, names of attributes in a lineage data object may be compared to names of attributes in a candidate data object. Individual values associated with a lineage data object may be compared to individual values in a candidate data object. Information about a set of values in a lineage data object (such as values of a particular attribute) may be compared to corresponding information of a candidate data object. More generally, a comparison may be made between a source object and a target object, where both compared objects may be lineage data objects or may be candidate data objects.
[0062] In addition to comparing a single aspect of a lineage data object to an aspect of a candidate object, multiple aspects may also be compared. In one example, the values of two attributes in one data object may be combined, such as mathematically or by concatenation, and compared to the value of a single attribute in another data object being compared. Similarly, scores for multiple attributes or other aspects of data objects may be compared. Sets of data objects may also be used for comparison purposes. For example, if two candidate data objects are identified as being related to two lineage data objects, and the candidate data objects are related in a manner corresponding to the relationship between the two lineage data objects, this may be more indicative of a corresponding relationship between the lineage data objects and the candidate data objects.
[0063] If it is determined that the lineage data object is sufficiently related to the candidate data object, a link can be established or proposed between the objects. For example, if the columns of the lineage data object are sufficiently related to the columns of the candidate data object, a foreign key relationship can be defined (or proposed) between the columns, or a connection can be defined (or proposed) between the two data objects using matching columns as a connection condition. In the case of a connection operation, defining a connection between data objects can include modifying an existing connection of a lineage graph, such as switching a connection between two lineage data objects to a connection between one of the lineage data objects and the candidate data object, or using a different lineage data object than was used in the original connection. In general, the disclosed techniques can be used to establish or propose a type of relationship described above between data objects, whether of the same type or different types, or at the same system or different systems.
[0064] In some aspects, once a possible relationship between a candidate data object and a lineage data object is identified, the possible relationship can be visually displayed to a user. The user can then choose whether to make changes to the lineage graph. For example, the user can formalize the relationship by including details about transformations, combinations, filtering, or aggregations between the lineage data object and the candidate data object that is added as a new lineage data object to the lineage graph.
[0065] In some scenarios, establishing a relationship between a lineage data object and a candidate data object can improve computing efficiency. In one example, the newly identified relationship can result in a smaller amount of data being transmitted or processed, or fewer data objects being used for a particular data stream. For example, a possible relationship between a first lineage data object and a second lineage data object or a candidate data object can be used to remove or modify the relationship between a first lineage data object and a third lineage data object. In particular, the identified possible relationship can provide a basis for a connection between a first lineage data object and a data object that is at a "higher level" in the hierarchy than the second lineage data object, which can be more efficient than the original connection with the second data object. In another example, data can be obtained from multiple sources, and the newly identified relationship can allow data to be obtained more efficiently, such as by obtaining data locally rather than from a remote system, or by obtaining data in a manner that requires no or less data conversion.
[0066] Although the disclosed techniques have been described as applied between lineage data objects and candidate data objects, they can also be used to identify new relationships between lineage data objects, between candidate data objects, or between a group of lineage data objects and candidate data objects that already have identified relationships with the group of lineage data objects.
[0067] Example 2 - Example relationships between data objects
[0068] Database systems typically include an information repository that stores information about a database schema, which is a specific example of object metadata that can be used in the disclosed technology. For example, PostgreSQL includes an INFORMATION_SCHEMA that includes information about tables in the database system, as well as certain table components, such as attributes (or fields) and their associated data types (e.g., varchar, int, float). Other database systems or query languages include similar concepts. However, as described above, these types of repositories typically only store technical information about database components, not semantic information.
[0069] Other database systems or applications or frameworks that operate using the database layer may include a repository for storing semantic information of data. For example, SAP SE of Walldorf, Germany provides the ABAP programming language that can be used in conjunction with a database system. ABAP provides the ability to develop database applications that are independent of the properties (including vendors) of the underlying relational database management system. In part, this ability is implemented using a data dictionary. The data dictionary may include at least some information similar to the information maintained in the information model. However, the data dictionary may include semantic information about the data and may optionally include additional technical information.
[0070] In addition, the data dictionary can include textual information about fields in the tables, such as a human-readable description of the purpose or use of the fields (sometimes in different languages, such as English, French, or German). In at least some cases, the textual information can be used as semantic information to the computer. However, other types of semantic information do not necessarily need to be (at least easily) understandable to humans, but can be more easily processed by computers than parsing textual information that is primarily intended for human use. The data dictionary can also contain or express relationships between data dictionary objects through various attributes (which can be reflected in metadata), such as having the data dictionary reflect that dictionary objects are assigned to packages and therefore have relationships to each other through the package assignments. The information in the data dictionary can correspond to metadata, which can be retrieved by the target system from the source system according to the techniques previously described in this disclosure.
[0071] As used herein, "technical information" (or technical metadata) refers to information that describes data as data, which is information such as the type of value that can be used to interpret the data, and it can affect how the data is processed. For example, the value "6453" can be interpreted (or converted) as an integer, a floating point number, a string, or a character array, as well as various possibilities. In some cases, a value can be processed differently, depending on whether it is a number such as an integer or a floating point number, or whether it is considered a collection of characters. Similarly, technical information can specify acceptable values for data, such as the length or number of decimal places allowed. Technical information can specify the properties of data without considering what the data represents or "means". However, of course, the designer of a database system can choose specific technical properties for specific data with the understanding of the semantic properties of the data - for example, "If I want to have a value that represents a person's name, I should use a string or a character array instead of a floating point number". On the other hand, in at least some cases, the data type may be a type that the database administrator or user does not expect. For example, instead of using a person's name to identify data associated with that person, a separate numeric or alphanumeric identifier is used, which may be counterintuitive based on the "meaning" of the data (for example, "I don't think of myself as a number").
[0072] As used herein, "semantic information" (or semantic metadata) relates to information that describes the meaning or purpose of data, which can be to a person or to a computer process. As an example, technical data information may specify that data having a value of the format "XXX-XX-XXXX" be obtained, where X is an integer between 0 and 9. This technical information can be used to determine how the data should be processed, or whether a particular value is valid (e.g., "111-11-1111" is valid, but "1111-11-1111" is not valid), but does not indicate what the value represents. Semantic information associated with the data may indicate whether the value is a social security number, a telephone number, a routing address, etc.
[0073] Semantic information can also describe how to process or display data. For example, "knowing" that the data is a phone number can cause the value to be displayed in one part of the GUI but not in another part of the GUI, or a specific processing rule may be invoked or not, depending on whether the rule is active for "phone number". In at least some cases, "semantic information" can include other types of information that can be used to describe the data or how the data should be used or processed. In certain cases, data can be associated with one or more of the following: a label, such as a human-understandable description of the data (e.g., "phone number"), documentation, such as a description of what information should be included in a field with a label (e.g., "Enter an 11-digit phone number including area code"), or information that can be used in a help screen (e.g., "Enter your home phone number here").
[0074] Typically, technical information must be provided for data. For example, in the case of fields of a database table, it is typically necessary to provide names or identifiers for the fields and data types. The name or identifier of a field may or may not be used to provide semantic information. That is, a database designer may choose the name "Employee_Name", "EMPN", or "3152". However, since the name or identifier is used to locate / distinguish the field from another field, it is considered technical information rather than semantic information in the context of the present disclosure, even though it can easily convey meaning to a person. In at least some implementations, the use of semantic information is optional. For example, even with a data dictionary, some fields used in database objects (such as tables, but also other objects, where such other objects are typically associated with one or more tables in the underlying relational database system) may be specified without using semantic information, while other fields are associated with semantic information.
[0075] Figure 1is an example entity-relationship (ER) type diagram illustrating a data schema 100 or metadata model related to driver accident history. The schema 100 (which may be part of a larger schema, other components of which are not shown in the diagram) Figure 1 ) may include table 108 associated with a license holder (e.g., an individual with a driver's license), table 116 associated with the license, table 112 representing accident history, and table 104 representing cars (or other vehicles).
[0076] Each of the tables 104, 108, 112, 116 has multiple attributes 120 (although in some cases, a table may have only one attribute). For a particular table 104, 108, 112, 116, one or more of the attributes 120 may be used as a primary key—uniquely identifying a particular record in a tuple and designated as the primary method of accessing tuples in the table. For example, in table 104, the Car_Serial_No attribute 120a is used as the primary key. In table 116, the combination of attributes 120b and 120c together is used as the primary key.
[0077] A table can reference records associated with the primary key of another table by using foreign keys. For example, the license number table 116 has an attribute 120d for Car_Serial_No in table 116 that is a foreign key and is associated with the corresponding attribute 120a of table 104. The use of foreign keys can be used for a variety of purposes. Foreign keys can link specific tuples in different tables. For example, a foreign key value of 8888 for attribute 120d will be associated with a specific tuple in table 104 that has this value for attribute 120a. Foreign keys can also act as constraints, where a record cannot be created with (or changed to have) a foreign key value that does not exist as a primary key value in the referenced table. Foreign keys can also be used to maintain database consistency, where changes to a primary key value can be propagated to the table for which the attribute is a foreign key.
[0078] Tables may have other attributes or combinations of attributes that can be used to uniquely identify tuples but are not primary keys. For example, table 116 has an alternate key formed by attribute 120c and attribute 120d. Thus, unique tuples in table 116 may be accessed using the primary key (e.g., as a foreign key in another table) or by association with the alternate key.
[0079] Schema information is typically maintained in a database layer, such as a software layer that maintains table values (e.g., in an RDBMS), and typically includes identifiers of tables 104, 108, 112, 116 and the names 126 and data types 128 of their associated attributes 120. Schema information may also include at least some information conveyable using flags 130, such as whether a field is associated with a primary key, or indicates a foreign key relationship. However, other relationships, including more informal associations, may not be included in a schema associated with a database layer (e.g., INFORMATION_SCHEMA for PostgreSQL).
[0080] Example 3 - Example computing environment with graph enhancement framework
[0081] Figure 2 An example computing environment 200 is shown in which the disclosed techniques may be implemented. The computing environment 200 includes a graph enhancement framework 210. Although the disclosed techniques may be implemented using a graph data structure, the disclosed techniques may be implemented in a variety of ways as long as a set of data objects (nodes) are available and relationships (edges) are defined between the data objects. Although this example 3 describes a specific example of a graph enhancement framework 210 for enhancing a lineage graph, the described techniques may be applied to data objects that are not technically part of a lineage graph, including data objects that are related by fewer than all relationship types that might typically be associated with a lineage graph, including relationships between data objects that are located within a common computing system rather than representing different computing systems.
[0082] The graph enhancement framework 210 includes a metadata manager 214. The metadata manager 214 retrieves information from various data sources 218 (shown as data sources 218a-218c). The information may include metadata 222 and data 226. The metadata 222 may be stored in or associated with a schema 228 such as a data dictionary, information schema, or other type of metadata data catalog.
[0083] In some cases, the comparison between the lineage data object and the candidate data object can be made using information describing the data object, such as descriptions of the table, the table's identifier, and the names and data types of the data columns. Other types of metadata can include semantic information, such as a more detailed name or description of the table or table columns, or information about the data in the table, such as the number of records in the table (table cardinality) or the number of unique values for a particular column (attribute cardinality).
[0084] Information about columns can also include variance and entropy. Variance, a statistical measure, quantifies the dispersion of data points relative to their mean. Mathematically, the variance of a data set X with n observations is given by:
[0085]
[0086] Here x i represents each individual data point in the collection, and is the average of all observed values. In the context of a relational database or similar data, the variance within a column (such as a column containing sales figures) describes the consistency or dispersion of those figures. A high variance indicates wide dispersion, meaning that the sales figures vary significantly from the mean. Conversely, a low variance means that most values fluctuate around the mean, indicating consistency in the sales values.
[0087] Entropy (such as that used in information theory) measures the level of unpredictability or randomness associated with data. 1 ,x 2 ,…,x n The entropy of a discrete random variable and its respective probabilities of possible values can be defined as:
[0088]
[0089] When applied to a database column, entropy measures the degree of diversity or uniqueness among its values. A column with varying values (such as a unique transaction ID) will have an entropy close to its maximum possible value, indicating high unpredictability. On the other hand, a column with repeated values will exhibit lower entropy. For example, a column that captures a predefined set of categories, such as order status ("shipped", "pending", "cancelled"), may show low entropy due to its limited set of possible values.
[0090] Data 226 may include data values for a particular lineage data object or candidate data object. In some scenarios, a comparison of actual data values between two data objects may be used to determine their similarity.
[0091] Metadata manager 214 may include one or more connectors 230. Connectors 230 may be used to retrieve data from one or more data sources 218. In the case of local data sources 218, connectors 230 may include query functionality for retrieving metadata 222, data 226, or schema 228. Data on remote data sources 218 may be accessed using technologies such as data federation or REST or other types of APIs.
[0092] The graph enhancement framework 210 also includes a comparator 234. The comparator 234 may include one or more comparison techniques 236. In general, comparison techniques may be used to compare metadata or data between two data objects. In a specific example that will be further described, columns of data or similar data organizations (such as key values for specific keys in data stored in a JSON representation) may be compared. The comparison may be performed on an element-by-element basis, or entire columns of information may be compared.
[0093] As described, the comparison can be made using information such as column (or key) names, data object names or identifiers, data types, object types, or by comparing individual data values between columns. The information can also include data object or data attribute cardinality, or statistics of data objects or columns, such as minimum or maximum values. However, various other metrics can also be used.
[0094] In data analysis, the Jaccard distance is a useful metric for assessing the similarity between datasets. It is derived from the Jaccard coefficient, which quantifies the overlap between two sets. Given two sets A and B, the Jaccard coefficient, denoted J(A,B), is calculated as follows:
[0095]
[0096] Here |A∩B| denotes the intersection of sets, and |A∪B| is their union.
[0097] Based on this coefficient, the Jaccard distance that captures the dissimilarity between two sets is:
[0098] D(A,B)=1-J(A,B)
[0099] In the context of relational databases, the Jaccard distance has two main applications: comparing single values and comparing summarized values. If each value in a column is considered to be a collection (of characters or tokens), the Jaccard distance can measure the similarity between the values in two columns. For example, when comparing the strings "apple" and "appetite", the Jaccard formula can be used to measure their similarity. This can be useful when dealing with text data that may have minor variations or alternative representations.
[0100] For more extensive column comparisons, summary sets can be created that capture unique values or attributes. Using the Jaccard distance on these sets provides a measure of overall content similarity between columns. As an example, for a column storing product categories, the unique categories from each column can be grouped into sets. Then, calculating the Jaccard distance between these sets gives an idea of how similar the product categories are between the columns.
[0101] Another measure that can be used to compare the values between a lineage data object and a candidate data object is the Euclidean distance. The Euclidean distance, derived from the Pythagorean theorem, gives a direct measure of the "straight line" distance between two points in Euclidean space.
[0102] Mathematically, for two points P(x 1 ,y 1 ) and Q(x 2 ,y 2 ), the Euclidean distance d is given by:
[0103]
[0104] This formula can be extended to higher dimensions. 1 ,x 2 ,…x n ) and (y 1 ,y 2 ,…y n ) in n-dimensional space, the distance between two points P and Q is:
[0105]
[0106] If each row in a column represents a point in a multidimensional space (e.g., a feature of a product or an attribute of a customer), then the Euclidean distance can measure the dissimilarity between rows across the columns. To assess the overall dissimilarity between two columns, you can calculate an aggregate measure (such as the mean or centroid) for each column and then determine the Euclidean distance between these aggregate points.
[0107] Cosine similarity can also be used for data comparison. Mathematically, given two vectors AA and BB, their cosine similarity is calculated as the dot product of the vectors divided by the product of their magnitudes:
[0108] Cosine Similarity
[0109] where A·B denotes the dot product of vectors A and B, and ‖A‖ and ‖B‖ denote their respective magnitudes.
[0110] When interpreting single values, especially text data represented in vector form, cosine similarity can assess the proximity of these values. Consider two text strings in a column, such as "database" and "data". By converting these strings to vector representations (usually using techniques such as TF-IDF or word embeddings), the cosine similarity between the vectors can illustrate their semantic proximity. This becomes particularly useful in identifying close matches or when the text data carries subtle differences.
[0111] For a broader level of comparison, aggregate vector representations can be extracted. The aggregate vector representation can be the centroid of the vectors for the individual entries or another form of aggregation. Computing cosine similarities between these aggregate vectors can help indicate the overall similarity of content or context between columns. For example, in a column that includes document topics, generating centroid vectors for the content of each column and measuring cosine similarity between these vectors can provide insights into how consistent the topics of two columns are.
[0112] Another data comparison technique uses the Levenshtein distance. Also known as the "edit distance," the Levenshtein distance quantifies the dissimilarity between two strings by calculating the minimum number of single-character edits (insertions, deletions, or substitutions) required to transform one string into the other.
[0113] Conceptually, the Levenshtein distance operates by constructing a matrix where one string is placed along the top row and the other string is placed along the side columns. Each cell of this matrix represents the partial edit distance between the substrings. The process starts with an initial setup where the value of each cell depends on its neighboring cells, representing the number of operations taken so far to match the two substrings. Working through this matrix, one eventually reaches the bottom right cell which provides the total edit distance between the two complete strings.
[0114] The Levenshtein distance can be used to compare individual strings in two columns. For example, in a column storing names, the Levenshtein distance can pinpoint and quantify small changes, typos, or inconsistencies between names in the columns. In the disclosed technology, the Levenshtein distance can be used to determine whether two columns include identical values or similar values. If the values in one column always differ by a single edit from the values in the other column, this may indicate that the values in the columns are equal, but one of the columns simply includes additional characters that may not affect the semantic meaning of the data.
[0115] Aggregation of Levenshtein distances is also useful. For example, computing the average Levenshtein distance between all pairs of strings from two columns provides a measure of the overall textual similarity or difference between them.
[0116] Some types of data, such as unstructured text or media files (images, video, or audio), can benefit from other types of comparison techniques. In particular, it may be useful to submit values to an autoencoder and then compare the encoded values.
[0117] An autoencoder is a neural network designed for unsupervised learning. Its architecture consists of two main components: an encoder and a decoder. The encoder compresses the input data into a lower dimensional representation called a latent space or encoding, which is usually of lower dimensionality than the original data. The decoder reconstructs the input data from the encoded representation. During training, the autoencoder minimizes the reconstruction error and learns to capture the essential features and patterns in the input data.
[0118] Unstructured data, including text and images, presents challenges in various machine learning tasks, including data comparison. To make unstructured data suitable for analysis, it is converted into a fixed-size vector representation. Autoencoders are trained on unstructured data, such as text or images, and learn to create meaningful and concise data representations. After training, the encoder converts unstructured data into fixed-size vectors that capture essential information.
[0119] The trained autoencoder enables the application of traditional vector-based comparison techniques such as cosine similarity or Euclidean distance to measure the similarity between documents or text snippets. This vector representation facilitates efficient and scalable comparison of unstructured data.
[0120] Return to Figure 2 , the graph generator 250 may include conditional logic 252. The conditional logic 252 may be used to evaluate the different similarity measures described above. For example, the conditional logic 252 may specify a Jaccard distance threshold distance that will be used to determine whether two data objects, data object components (such as columns), or data values are sufficiently similar to be considered related.
[0121] In some scenarios, a similarity may be used instead of or in addition to a binary result of determining similarity. The similarity may be a raw metric value, or may be derived from one or more metric values. In a similar manner, a binary determination of similarity may be made using raw values, or may be derived from one or more metric values. Even if two objects are sufficiently similar that an edge is suggested between a lineage data object and a candidate data object, a similarity value may be reported, which may help a human or computational process understand how strong a relationship may be, which may be used, for example, to modify a data flow based on a candidate data object identified as being related to a lineage data object.
[0122] exist Figure 2 , table 260 provides an example representation of conditional logic 252. Table 260 includes a column 262a identifying a particular comparison metric, a column 262b identifying a threshold for determining whether the comparison metric indicates a relationship, and a column 262c associating a priority with the particular comparison metric. That is, at least in some cases, a variety of techniques may be used to evaluate a given pair of data objects. For example, priority information may be used to weight different binary results of a comparison technique, where the combined results are then used to make a final determination of whether two data objects are related, or a measure of the strength of a relationship.
[0123] Note that table 260 includes a condition for level distance. Level distance can indicate the degree of indirection between two data objects. For example, given an existing lineage graph, relationships between lineage data objects that may not yet be reflected in the lineage graph can be determined. A relationship can be considered more direct than an existing indirect relationship. However, a relationship can be more strongly indicated as a lower level of indirection. When a candidate data object is identified as being related to a lineage data object, a similar evaluation can be considered, where additional lineage data objects can be considered for possible relationships to the candidate data object.
[0124] The graph generator 250 of the graph enhancement framework 210 may be used to generate a graph based on values generated using the conditional logic 252. For example, the graph generator 250 may use rules 270 to determine whether a particular candidate data object should be added to the graph, which objects in the graph it should be connected to, and optionally provide a weight or other metric (which may also be referred to as a score) to indicate the strength of the relationship.
[0125] Optionally, the graph generator 250 may call the relationship generator 274 to determine the type of relationship, or to suggest a specific implementation of the relationship, or to actually implement the relationship. As an example, consider a lineage data object that is identified as having a relationship with a candidate data object. Both objects may have multiple attributes such as columns. Assuming that both objects have multiple columns, it may be determined that the objects are related based on one or more of the specific columns, where other columns may not have a corresponding "match" in the other data object. In the case of a relational database table or view, the relationship generator 274 may suggest or generate a connection based on one or more related columns, or may suggest a foreign key relationship.
[0126] For example, metadata 222 associated with a data object can be used to identify primary key columns in the data object. Such information can also be used to suggest relationships at the outset, such as determining that a column of a data object is a more likely connection or foreign key relationship if the column forms the primary key of the object.
[0127] In the case of connecting or establishing another type of relationship (such as a relationship that results in data import of a candidate data object), all data from the candidate data object (which can now be considered a lineage data object) can be connected or imported. Alternatively, more complex logic can be used to retrieve a selected portion of such data.
[0128] Information for the graph enhancement framework 210, such as metadata 222, data 226, schema 228, or relationships between data objects established using graph generator 250 or relationship generator 274, may be stored in a data store 278. In particular, data store 278 is shown to include a graph 280, which may correspond to a graph definition—the nodes and edges of a particular graph, which may be an “initial” graph or a graph generated during or at the end of the disclosed graph enhancement process.
[0129] The graph enhancement framework 210 can communicate with the data flow component 284. The data flow component 284 can store information used to generate a lineage graph and the lineage graph generated therefrom. The data flow component 284 can also receive and use updated graphs generated by the graph enhancement framework 210.
[0130] The data flow component 284 is shown to include data model metadata 286, where the data model metadata can include information about various types of relationships between data objects (including lineage graph objects). Such information can include the definition of an ETL process 288 or a data model 290 that describes the data in the relevant data source 218 (where the data model can correspond to the schema 228 of the data source).
[0131] The data flow component 284 may include a schema 292 for a data source of a graph 294 (eg, a lineage graph). The data flow component 284 may include a user interface 296, which may allow a user to visualize and interact with (including modify) the graph 294.
[0132] Example 4 - Example Lineage Diagram
[0133] Figure 3 An example lineage diagram 300 is shown. The lineage diagram 300 includes a plurality of lineage data objects 310. The lineage data objects 310 may be of the same type or may be of different types and may be associated with the same system or different systems. The lineage data objects 310 may be data such as relational database tables or views, entities or views in a virtual data model, data stored in OLAP objects, data stored in logical data objects (such as BUSINESSOBJECTS of SAP SE of Walldorf, Germany), or data in a semi-structured or unstructured format (including data stored in XML or JSON format).
[0134] Typically, a lineage data object 310 includes one or more attributes (not shown) that can be used to identify the lineage data object 310. Figure 1 The same method as described above is implemented. Figure 1 As described, and in the discussion of Example 1, lineage data objects can have various relationships to each other, including through the use of attributes.
[0135] Lineage data objects 310 may be connected via edges 320. Edges 320 indicate a relationship between two data objects, where typically edges are directed. Figure 3 As shown, the lineage data object 310 at the top of the lineage graph can be “built” on lower-level lineage data objects. For example, a lineage data object 310 in the form of a view can be defined with respect to other views or relational database tables.
[0136] The relationships between lineage data objects 310 may be specified at various levels of detail. In some cases, an edge 320 may simply represent a relationship, optionally with a direction. In other cases, an edge 320 may include information such as the nature of the relationship (such as a join or foreign key relationship) or specific details about how the relationship is implemented. In the case of a join, the details may include a join condition, which includes attributes involved in the join. In the case of a foreign key relationship, the details may include attributes used as a primary key and attributes used as a foreign key.
[0137] In lineage graph 300, dependency nodes 328 indicate when a lineage data object 310 has a direct or indirect dependency on one or more other lineage data objects. For example, lineage data object 310a Products_View has a direct dependency on data object 310b Products, and has an indirect dependency on lineage data object 310c ProductTexts. In turn, lineage data object 310b also has a direct dependency on lineage data object 310c. Although not shown, dependency nodes 328 are typically associated with specific relationship types (such as foreign key relationships or connections) and specific attributes used in the relationship type.
[0138] As discussed, lineage graph 300 can be used as a starting point for graph enhancement. That is, lineage data objects 310 can be evaluated as being consistent with Figure 2 The graph enhancement framework provides the relationships of the candidate data objects 340. In addition, the lineage data objects 310 themselves can be evaluated for relationships that are not captured in the existing lineage graph.
[0139] Example 5 - Example comparison of data object attributes
[0140] Figure 4 A process 400 of comparing a lineage data object to a candidate data object is shown. Although this example is described with respect to a table, comparisons can be performed between other types of data objects, including between different types of data objects. That is, generally a data object can be considered a collection of one or more attributes, where generally the attributes have a specific semantic meaning and a specific data type. Furthermore, as previously described, the techniques of this Example 5 can be applied generally between pairs of data objects, even if one or more of the data objects being compared are not associated with a lineage graph.
[0141] In particular, process 400 is performed using a first data object 410 and a second data object 420. The first data object 410 and the second data object 420 have respective attributes 412, 422. The values 416, 426 of the attributes 412, 422 are arranged in a set of specific instances 414, 424 of the attributes, such as corresponding to rows of a relational database table or specific instances of a JSON object having a specific schema.
[0142] The values 416, 426 of specific instances 414, 424 can be compared, as will be further described. However, instances 414, 424 can optionally be filtered first. For example, in processes 430, 432, the values 416, 426 of specific data object instances 410, 420 can be filtered using the criteria of those terms such as variance or entropy defined in Example 3. That is, not all attributes provide the same value or meaning in this comparison. Attributes with very high variance or entropy can indicate a wide distribution of values or a high degree of randomness, respectively. These attributes may obscure meaningful comparisons by drowning them with noise. By filtering out attributes based on entropy or variance thresholds, comparisons between data objects (including comparisons between value sets of specific data object attributes) can be focused on more stable or predictable attributes. This can simplify comparative analysis and highlight more substantial differences or similarities between two data sets. In addition, attributes with extremely high variance or entropy may indicate data quality issues, such as errors in missing value filling or data collection.
[0143] Relatedly, when there is significant internal variability within a set of values for instances of a particular attribute, the attribute itself may not have a consistent or clear meaning within its own data set. Thus, when the attribute is compared to another attribute from a different data set, the comparison may become less meaningful. The inherent variability or randomness within one data set may make it challenging to discern whether differences observed between data sets are due to real differences or are simply the result of high variability in the original data.
[0144] Optionally, other criteria can be used to filter the attributes to be compared. For example, in some scenarios, some attributes may not be semantically meaningful. For example, Boolean values may not be as meaningful as numeric or string values.
[0145] In addition to or in lieu of being used as a filtering criterion, metrics such as variance and entropy or information such as attribute data type may be used when evaluating two sets of attributes for semantic equivalence or semantic relationships. For example, two sets of attributes having the same data type and similar entropy or variance may be more semantically similar than two sets of attributes having different data types or having significantly different entropies or variances.
[0146] In some cases, in operation 440, summary information, such as variance, entropy, data type, name, cardinality, or other descriptive information, is compared between the attributes of data object 410 and one or more attributes (and in some cases, all filtered attributes) of data object 420. In addition or alternatively, individual attribute values may be compared. For example, for two attributes being compared, the values of the attributes of data object 410 may be compared to all values of the attributes of data object 420 in pairs. If it is known that the values in the attributes of data objects 410, 420 are ordered in a common manner, or it is otherwise known which attribute values are expected to correspond, a single pair comparison may be performed between the attribute values.
[0147] Figure 4 Attribute 412a of first data object 410 is shown being compared to attributes 422a-422d of second data object 420. In this case, information about attribute 412a may be compared to information about attributes 422a-422d, such as their names, data types, entropy, variance, attribute cardinality, or instance (record) cardinality.
[0148] Figure 4 Also shown is a specific set of pairwise comparisons between value 416a of attribute 412a of the first data object and each value 426a-426c of attribute 422 of the second data object 420. Similar comparisons may be made for each value 416a of attribute and value 426 of attributes 422b-422d.
[0149] The comparison can include determining whether two values are identical or related, including using the techniques described in Example 3. For example, if the attribute value being compared is associated with an attribute having a string value, the Levenshtein distance between two particular attribute values can be determined. In the case of numerical data, the equality of two attribute values can be tested, or the percentage difference between the two values can be calculated. In some cases, numerical values that differ by less than a threshold amount can be classified as identical. Similarly, the Euclidean distance between two values can be calculated, and if the distance is less than a defined threshold, the values can be considered identical. For unstructured data, an autoencoder can be used to process the attribute values and compare the resulting encoded vectors, including as described for numerical values.
[0150] The information calculated using the attribute values 416, 426 can also be used to compare the summary information of the attributes 412a, 422a in operation 440. For example, the values 416, 426 of the compared attributes 412, 422 can be used to calculate the Jaccard distance or cosine similarity. Alternatively, the aggregated values of the attribute comparison can be compared. For example, for string data or character arrays, the Levenstein distance between the attribute values 416, 426 of a particular attribute pair 412, 422 can be calculated. The average Levenstein distance can be calculated for each attribute 412a, 422a, and those averages can be compared. Other types of statistical measurements can be used in a similar manner, such as by comparing the standard deviation of the Levenstein values of two attributes 412a, 422a.
[0151] In some scenarios, attribute 412 may correspond to multiple attributes 422. For example, value 416 may correspond to a combination of values 426 of multiple attributes 422. The opposite scenario may also occur, where values 416 of multiple attributes 412 correspond to value 426 of a single attribute 422. The disclosed technology can analyze data for these types of more complex relationships.
[0152] Figure 4 A specific example of a more complex scenario of these types is shown, where comparison 448 compares the value 416 of attribute 412b to a value formed by concatenating the values 426 from attributes 422c, 422d. For attributes with numeric values, Figure 4 It is shown how a modified attribute value 460 is produced by applying an operator to the value 416 of the attribute 422a. The value 460 can then be compared with the value of the second data object 420 (such as the value 426 of the attribute 422d), which can be performed as described above. A similar combination / "scaling" operation can optionally be performed on the summary data of the attributes 412b, 422d, such as scaling the average attribute value.
[0153] Figure 5 Example pseudocode 500 is provided for a process of comparing data objects (in this case, database tables) to determine whether a relationship should be established (or reported) between two tables. In pseudocode 500, a relationship between tables is identified as long as at least one column of the second table (which may correspond to a candidate data object) is sufficiently similar to a column of the first table (which may correspond to a lineage data object). However, the technique may be implemented in another manner. For example, a relationship may be identified when multiple columns of the second table match columns of the first table. As described, the sufficiently similar columns may optionally be used not only to identify the relationship, but also to establish the basis and type of the relationship, such as corresponding to a foreign key relationship using specific columns of the first and second tables.
[0154] Figure 6A lineage graph 300 is shown with a table 610 identified as being related to a lineage data object 310d using the disclosed techniques added thereto. An edge 620 connecting the table 610 and the lineage data object 310d is shown. The edge 620 is again shown with a value indicating the strength of the relationship. The strength of the relationship may be determined in various ways, such as increasing the strength value if a greater number of attributes are identified as similar between the table 610 and the lineage data object 310d, or based on a particular similarity value calculated for the table and the lineage data object. For example, if a greater number of attribute values match between the table 610 and the lineage data object 310d, or if a metric (such as a comparison of averages) or a distance (such as Levenshtein, Jaccard, or Euclidean distance) indicates a stronger relationship (such as by having a shorter distance), the edge 620 may be weighted more heavily or scored higher.
[0155] In some cases, edge 620 may be converted to a "formal" relationship - a relationship that adds table 610 to lineage graph 300. Converting edge 620 may include defining a transformation, combination, filter, or aggregation associated with lineage data object 310d and table 610, and may also include assigning directionality to the edge. For example, edge 620 may be defined to include a value of the example LineageEdge class defined in Example 1.
[0156] In some cases, these operations can be performed automatically using defined rules. As a simple example, a rule can be defined that results in the combination of data from lineage data object 310d and table 610 (such as via a join operation), where it is assumed that the data in table 610 should be represented in data object 310d. In other cases, a user can manually determine the conversion of edge 620 to a formal relationship, including defining properties associated with the connection between lineage graph objects.
[0157] Example 6 - Example Graph Enhancement Process
[0158] Figure 7 A flow chart of a process 700 of the present disclosure for identifying relationships between data objects is provided. Metadata for the data objects is collected at 705. Collecting metadata may include collecting information about available data objects, which may include candidate data objects as well as lineage data objects. Metadata may include information such as object identifiers, object attributes, attribute data types, and information about values in the data objects, including their specific attributes. For example, information about values in the data objects may include cardinality information, entropy, variance or maximum, minimum or mean values, or information about the distribution of values present in the data for a given data object.
[0159] The metadata may be filtered at 710. Filtering the metadata may include removing from consideration data objects or data object attributes that do not meet specific data type requirements or whose entropy or variance does not meet threshold criteria. Filtering the metadata at 710 may also include filtering the metadata using provided processing criteria, such as identifying specific data sources or patterns to be considered or not considered in the relationship identification process.
[0160] At 715, an initial relationship graph may be retrieved or generated. The graph may be, for example, a lineage graph. The lineage graph may have been defined or may be generated based on the relationships identified between lineage data objects of the lineage graph.
[0161] At 720, data of the data objects or specific attributes thereof are compared, where the data may include metadata or actual data values of instances of the data objects. The comparison may be performed as described and may be used to generate various similarity measures. Relationships between data objects may be established as part of the operation at 720.
[0162] A graph, such as a lineage graph or a graph of data using a lineage graph, is iteratively updated at 725. Iteratively updating the graph may include adding data objects to the graph based on the operation at 720, and the updated graph may then be analyzed again using operation 720. Relationships may be proposed or generated at 730. That is, the graph may be updated to include the additional data objects without formally establishing relationships between the data objects. The operation at 730 may include, for example, defining a connection or foreign key relationship between two data objects, or proposing such a definition to a user, who may then confirm whether the relationship should be implemented.
[0163] Optionally, an enhanced graph is displayed at 735 , where the enhanced graph may correspond to the original lineage graph with any added data objects or data object relationships determined as part of process 700 .
[0164] Example 7 - Example Relationship Identification Process
[0165] Figure 8 is a flow chart of a process 800 for identifying at least an inferred relationship between a lineage data object and a candidate data object. At 810, a definition of a lineage graph of a plurality of lineage data objects is received. The definition includes identifiers of lineage data objects in the plurality of lineage data objects forming nodes of the lineage graph and identifiers of relationships between the plurality of lineage data objects forming edges of the lineage graph, wherein the lineage data objects include one or more attributes.
[0166] At 820, metadata for a plurality of candidate data objects is received. The metadata includes a candidate data object identifier and an identifier of an attribute defined for the respective candidate data object. A lineage data object is selected at 830, and a candidate data object is selected at 840. At 850, the lineage data object is compared to the candidate data object. At 860, it is determined that the lineage data object and the candidate data object satisfy a relationship criterion. At 870, at least an inferred relationship between the candidate data object and the lineage data object is established. At 880, an updated lineage graph including at least the inferred relationship is rendered for display.
[0167] Example 8 - Calculation System
[0168] Fig. 9 A general example of a suitable computing system 900 is depicted in which the described innovations may be implemented. Computing system 900 is not intended to suggest any limitation as to the scope of use or functionality of the present disclosure, as the innovations may be implemented in various general-purpose or special-purpose computing systems.
[0169] refer to Fig. 9 , the computing system 900 includes one or more processing units 910, 915 and memories 920, 925. Fig. 9 , the basic configuration 930 is included within the dashed line. Processing units 910, 915 execute computer executable instructions, such as for implementing the database environment and associated methods described in Examples 1-7. The processing unit can be a general-purpose central processing unit (CPU), a processor in an application-specific integrated circuit (ASIC), or any other type of processor. In a multi-processing system, multiple processing units execute computer executable instructions to increase processing power. For example, Fig. 9 A central processing unit 910 and a graphics processing unit or co-processing unit 915 are shown. Tangible memory 920, 925 may be volatile memory (e.g., registers, cache, RAM), non-volatile memory (e.g., ROM, EEPROM, flash memory, etc.), or some combination of both accessible by the processing units 910, 915. The memory 920, 925 stores software 980 implementing one or more innovations described herein in the form of computer-executable instructions suitable for execution by the processing units 910, 915.
[0170] The computing system 900 may have additional features. For example, the computing system 900 includes storage 940, one or more input devices 950, one or more output devices 960, and one or more communication connections 970. An interconnection mechanism (not shown) such as a bus, controller, or network interconnects the components of the computing system 900. Typically, operating system software (not shown) provides an operating environment for other software executed in the computing system 900 and coordinates the activities of the components of the computing system 900.
[0171] Tangible storage 940 may be removable or non-removable and include disks, tapes or cassettes, CD-ROMs, DVDs, or any other medium that can be used to store information in a non-transitory manner and that can be accessed within computing system 900. Storage 940 stores instructions for software 980 implementing one or more innovations described herein.
[0172] The input device 950 may be a touch input device (such as a keyboard, mouse, pen, or trackball), a voice input device, a scanning device, or another device that provides input to the computing system 900. The output device 960 may be a display, a printer, a speaker, a CD burner, or another device that provides output from the computing system 900.
[0173] Communication connection 970 enables communication with another computing entity (such as another database server) via a communication medium. The communication medium transmits information such as computer executable instructions, audio or video input or output, or other data in a modulated data signal. A modulated data signal is a signal that encodes the information in the signal by setting or changing one or more of its characteristics. By way of example and not limitation, the communication medium may use electrical, optical, RF or other carriers.
[0174] The innovations may be described in the general context of computer executable instructions, such as those included in program modules that are executed in a computing system on a target real or virtual processor. Typically, program modules or components include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules may be combined or split between program modules as desired. Computer executable instructions for program modules may be executed within a local or distributed computing system.
[0175] The terms "system" and "device" are used interchangeably herein. Unless the context clearly indicates otherwise, neither term implies any limitation on the type of computing system or computing device. In general, a computing system or computing device may be local or distributed, and may include any combination of special-purpose hardware and / or general-purpose hardware with software that implements the functionality described herein.
[0176] For purposes of presentation, the detailed description uses terms such as "determine" and "use" to describe computer operations in a computing system. These terms are high-level abstractions for operations performed by a computer and should not be confused with actions performed by a human. The actual computer operations corresponding to these terms vary depending on the implementation.
[0177] Example 9 - Cloud computing environment
[0178] Fig.10 An example cloud computing environment 1000 is depicted in which the described techniques may be implemented. Cloud computing environment 1000 includes cloud computing services 1010. Cloud computing services 1010 may include various types of cloud computing resources, such as computer servers, data storage repositories, network resources, etc. Cloud computing services 1010 may be centrally located (e.g., provided by a data center of an enterprise or organization) or distributed (e.g., provided by various computing resources located in different locations, such as different data centers and / or located in different cities or countries).
[0179] The cloud computing service 1010 is utilized by various types of computing devices (e.g., client computing devices), such as computing devices 1020, 1022, and 1024. For example, the computing devices (e.g., 1020, 1022, and 1024) can be computers (e.g., desktop or laptop computers), mobile devices (e.g., tablet computers or smart phones), or other types of computing devices. For example, the computing devices (e.g., 1020, 1022, and 1024) can utilize the cloud computing service 1010 to perform computing operations (e.g., data processing, data storage, etc.).
[0180] Example 10 - Implementation
[0181] Although the operations of some of the disclosed methods are described in a particular sequential order for ease of presentation, it should be understood that this description encompasses rearrangement unless the specific language set forth herein requires a particular ordering. For example, the operations described in sequence may be rearranged or performed simultaneously in some cases. In addition, for simplicity, the accompanying drawings may not show the various ways in which the disclosed methods may be used in conjunction with other methods.
[0182] Any disclosed method may be implemented as computer executable instructions or a computer program product stored on one or more computer readable storage media (such as tangible, non-transitory computer readable storage media) and executed on a computing device (e.g., any available computing device, including a smartphone or other mobile device that includes computing hardware). Tangible computer readable storage media is any available tangible media that can be accessed within a computing environment (e.g., one or more optical media disks such as DVDs or CDs, volatile memory components (such as DRAM or SRAM), or non-volatile memory components (such as flash memory or a hard drive)). As an example and with reference to Fig. 9 , computer-readable storage media include memories 920 and 925 and storage 940. The term computer-readable storage medium does not include signals and carrier waves. In addition, the term computer-readable storage medium does not include communication connections (eg, 970).
[0183] Any computer executable instructions for implementing the disclosed technology and any data created and used during the implementation of the disclosed embodiments can be stored on one or more computer readable storage media. The computer executable instructions can be, for example, a dedicated software application or a part of a software application accessed or downloaded via a web browser or other software application (such as a remote computing application). Such software can be executed, for example, on a single local computer (e.g., any suitable commercially available computer) or in a network environment using one or more network computers (e.g., via the Internet, a wide area network, a local area network, a client-server network (such as a cloud computing network) or other such networks).
[0184] For the sake of clarity, only some selected aspects of the software-based implementation are described. Other details known in the art are omitted. For example, it should be understood that the disclosed technology is not limited to any specific computer language or program. For example, the disclosed technology can be implemented by software written in C++, Java, Perl, JavaScript, Python, Ruby, ABAP, structured query language, Adobe Flash or any other suitable programming language, or in some examples, by a markup language such as html or XML or a combination of suitable programming languages and markup languages. Similarly, the disclosed technology is not limited to any specific computer or hardware type. Some details of suitable computers and hardware are known and do not need to be elaborated in this disclosure.
[0185] In addition, any software-based embodiments (including, for example, computer-executable instructions for causing a computer to perform any disclosed method) can be uploaded, downloaded, or remotely accessed via suitable communication means. Such suitable communication means include, for example, the Internet, the World Wide Web, an intranet, a software application, cables (including fiber optic cables), magnetic communications, electromagnetic communications (including RF, microwave, and infrared communications), electronic communications, or other such communication means.
[0186] The disclosed methods, apparatus, and systems should not be construed as limiting in any way. Rather, the present disclosure is directed to all novel and non-obvious features and aspects of the various disclosed embodiments, individually and in various combinations and sub-combinations with each other. The disclosed methods, apparatus, and systems are not limited to any particular aspect or feature or combination thereof, nor do the disclosed embodiments require the presence of any one or more specific advantages or problems to be solved.
[0187] The technology from any example can be combined with the technology described in any one or more other examples. In view of the many possible embodiments to which the principles of the disclosed technology can be applied, it should be recognized that the illustrated embodiments are examples of the disclosed technology and should not be considered as limiting the scope of the disclosed technology. On the contrary, the scope of the disclosed technology includes what is covered by the scope and spirit of the appended claims.
Claims
1. A computing system comprising: at least one memory; one or more hardware processor units coupled to the at least one memory; as well as One or more computer-readable storage media storing computer-executable instructions that, when executed, cause the computing system to perform operations including: receiving a definition of a lineage graph of a plurality of lineage data objects, the definition comprising identifiers of lineage data objects of the plurality of lineage data objects forming nodes of the lineage graph and identifiers of relationships between the plurality of lineage data objects forming edges of the lineage graph, wherein the lineage data objects comprise one or more properties; receiving metadata for a plurality of candidate data objects, the metadata comprising candidate data object identifiers and identifiers of attributes defined for the respective candidate data objects; Select the lineage data object; selecting candidate data objects; comparing the lineage data object to the candidate data object; determining that the lineage data object and the candidate data object satisfy a relationship criterion; establishing at least an inferred relationship between the candidate data object and the lineage data object; as well as An updated lineage graph including the at least inferred relationship is rendered for display.
2. The computing system of claim 1, wherein: Comparing the lineage data object with the candidate data object includes: Attribute definitions of attributes of the lineage data object are compared with attribute definitions of attributes of the candidate data object.
3. The computing system of claim 1, wherein: Comparing the lineage data object with the candidate data object includes: The summary information of the attribute of the lineage data object is compared with the summary information of the attribute of the candidate data object.
4. The computing system of claim 3, wherein: The summary information includes the entropy or variance of values in the attributes of the lineage data object and the attributes of the candidate data object.
5. The computing system of claim 3, wherein: The summary information includes the number of values of the attribute in the instance of the lineage data object and the number of values of the attribute in the instance of the candidate data object.
6. The computing system of claim 3, wherein: The summary information includes a number of unique values for the attribute in the instance of the lineage data object and a number of unique values for the attribute in the instance of the candidate data object.
7. The computing system of claim 1, wherein: Comparing the lineage data object to the candidate data object includes comparing a first value of a metric generated at least in part from a value of at least one attribute of the lineage data object to a second value of the metric generated at least in part from a value of at least one attribute of the candidate data object.
8. The computing system of claim 7, wherein: The first value of the metric is a statistical value generated from values of instances of a corresponding lineage data object or candidate data object.
9. The computing system of claim 7, wherein: The first value of the metric is a statistical value generated by comparing respective attribute values of respective instances of the lineage data object and the candidate data object.
10. The computing system of claim 1, wherein: Comparing the lineage data object to the candidate data object includes comparing a property value of a property of an instance of the lineage data object to a combination of values of a plurality of properties of an instance of the candidate data object.
11. The computing system of claim 1, wherein: Comparing the lineage data object to the candidate data object includes comparing an attribute value of an attribute of an instance of the lineage data object to a result of performing a mathematical operation on the attribute value of the instance of the candidate data object.
12. The computing system of claim 1, wherein: The at least inferred relationship comprises a foreign key relationship or connection.
13. The computing system of claim 1, wherein: At least a portion of the relationships between the plurality of lineage data objects corresponds to a foreign key relationship, connection, or association.
14. The computing system of claim 1, wherein: The plurality of lineage data objects are associated with a plurality of software programs or layers.
15. The computing system of claim 1, the operations further comprising: After establishing the at least inferred relationship, the candidate data object is considered as part of the lineage graph, and one or more of the following additional operations are performed: selecting a lineage data object, selecting a candidate data object, and comparing the lineage data object to the candidate data object.
16. The computing system of claim 1, wherein: The relationship criteria includes a relationship threshold defined with respect to a first metric that compares information about one or more attributes of the lineage data object to information about one or more attributes of the candidate data object.
17. The computing system of claim 1, wherein: Establishing at least a putative relationship between the candidate data object and the lineage data object includes assigning a score to the at least putative relationship.
18. The computing system of claim 1, the operations further comprising: User input defining a lineage graph edge corresponding to the inferred relationship is received, the user input comprising at least one filtering, aggregation, data combining, or data transformation operation.
19. A method implemented in a computing system, the computing system comprising at least one hardware processor and at least one memory coupled to the at least one hardware processor, the method comprising: receiving a definition of a lineage graph of a plurality of lineage data objects, the definition comprising identifiers of lineage data objects of the plurality of lineage data objects forming nodes of the lineage graph and identifiers of relationships between the plurality of lineage data objects forming edges of the lineage graph, wherein the lineage data objects comprise one or more properties; receiving metadata for a plurality of candidate data objects, the metadata comprising candidate data object identifiers and identifiers of attributes defined for the respective candidate data objects; Select the lineage data object; selecting candidate data objects; comparing the lineage data object to the candidate data object; determining that the lineage data object and the candidate data object satisfy a relationship criterion; establishing at least an inferred relationship between the candidate data object and the lineage data object; as well as An updated lineage graph including the at least inferred relationship is rendered for display.
20. One or more computer-readable storage media comprising: Computer executable instructions that, when executed by a computing system including at least one hardware processor and at least one memory coupled to the at least one hardware processor, cause the computing system to receive a definition of a lineage graph of a plurality of lineage data objects, the definition including identifiers of lineage data objects of the plurality of lineage data objects forming nodes of the lineage graph and identifiers of relationships between the plurality of lineage data objects forming edges of the lineage graph, wherein the lineage data objects include one or more properties; computer-executable instructions that, when executed by the computing system, cause the computing system to receive metadata for a plurality of candidate data objects, the metadata comprising candidate data object identifiers and identifiers of attributes defined for the respective candidate data objects; computer-executable instructions that, when executed by the computing system, cause the computing system to select a lineage data object; computer-executable instructions that, when executed by the computing system, cause the computing system to select a candidate data object; computer-executable instructions that, when executed by the computing system, cause the computing system to compare the lineage data object to the candidate data object; computer-executable instructions that, when executed by the computing system, cause the computing system to determine that the lineage data object and the candidate data object satisfy a relationship criterion; computer-executable instructions that, when executed by the computing system, cause the computing system to establish at least an inferred relationship between the candidate data object and the lineage data object; as well as Computer-executable instructions that, when executed by the computing system, cause the computing system to render for display an updated lineage graph including the at least inferred relationship.