Implementing metadata

Through the metadata management system, the metadata model is used to identify and operate data items within the enterprise, and the problem of low metadata management and operation efficiency in the existing technology is solved, and automated operations and efficient data processing are realized.

CN119948476AActive Publication Date: 2025-05-06AB INITIO TECHNOLOGY LLC
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202380068199.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-31
Filing Date
2023-08-09
Publication Date
2025-05-06
Estimated Expiration
2043-08-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively manage and operate metadata within the enterprise, making it difficult for users to understand and access data, and data processing operations are easily interrupted.

Method used

Through the metadata management system, it uses the metadata model to identify and operate data items, including identifying data items and their physical metadata, accessing the metadata model, traversing edges to identify the parent node, determining and applying operations, storing transformed data items, and updating the metadata model.

Benefits of technology

It realizes the automated operation and management of metadata, improves data processing efficiency, enhances the robustness and integrity of data strategies, and reduces the need for frequent access and processing of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948476A_ABST
    Figure CN119948476A_ABST
Patent Text Reader

Abstract

A method for performing operations on a data item using a metadata model, where the metadata model comprises a parent node and a child node connected by an edge, where the parent node specifies logical metadata and the child node specifies physical metadata representing the data item, and where the edge specifies a relationship between the nodes. The method includes identifying a given data item and physical metadata for the given data item, accessing the metadata model, identifying a child node representing the physical metadata for the given data item in the metadata model, traversing one or more edges in the metadata model to identify a parent node for the child node, and transmitting the parent node to the metadata model. One or more operations to be performed on the given data item are determined from logical metadata associated with the identified parent node, the one or more operations are applied to the given data item to transform the data item, and the transformed data item is stored.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] priority

[0002] This application claims priority to and the benefit of U.S. patent application No. 18 / 104,066, filed on January 31, 2023, which claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 400,333, filed on August 23, 2022, the entire contents of both patent applications are incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure relates to techniques for making metadata actionable within a data enterprise. Background Art

[0004] In order to conduct business, organizations maintain increasingly large and complex data sets. In many cases, an organization's data is stored in a way that promotes efficient use of computing resources rather than human comprehension. As a result, it can be difficult for users to understand what data an organization has, or how to access that data for various tasks. Even when a user is able to identify a piece of relevant data, generating code to perform the desired operations on the data can be difficult and time-consuming, particularly for non-technical users. In addition, changes to an organization's data can disrupt or otherwise render existing data processing operations obsolete. Summary of the invention

[0005] In general, in a first aspect, a method implemented by a metadata management system uses a metadata model to identify one or more data items and perform one or more operations on them, wherein the metadata model includes one or more parent nodes and one or more child nodes, wherein the one or more operations are defined relative to the one or more parent nodes in the metadata model and applied to data represented by the one or more child nodes in the metadata model, wherein the one or more parent nodes specify logical metadata, and wherein the one or more child nodes specify physical metadata representing the one or more data items. The method includes: identifying a given data item and physical metadata for the given data item; accessing the metadata model by the metadata management system, wherein edges connect nodes, wherein the edges specify a relationship between two nodes; identifying a child node in the metadata model that represents the physical metadata of the given data item; traversing one or more edges in the metadata model to identify one or more parent nodes of the child node; determining one or more operations to be performed on the given data item from metadata associated with the identified one or more parent nodes; applying the one or more operations to the given data item by the metadata management system to transform the given data item; and storing the transformed data item in a memory.

[0006] Generally speaking, in a second aspect combinable with the first aspect, identifying a child node representing physical metadata of a given data item comprises matching the physical metadata of the given data item with physical metadata represented by the child node in the metadata model.

[0007] Generally speaking, in a third aspect that can be combined with the first aspect and the second aspect, the method also includes accessing one or more metadata transformations to determine one or more operations to be performed on a given data item, each metadata transformation specifying at least one operation to be performed on the data and at least one condition for performing the at least one operation.

[0008] Generally speaking, in a fourth aspect which can be combined with any one of the first to third aspects, determining the one or more operations to be performed on a given data item includes: selecting a metadata transformation from the one or more metadata transformations; determining whether the logical metadata associated with the one or more parent nodes satisfies at least one condition of the selected metadata transformation; and in response to determining that the logical metadata associated with the one or more parent nodes satisfies at least one condition of the selected metadata transformation, determining that the one or more operations to be performed on the given data item include at least one operation specified by the selected metadata transformation.

[0009] Generally speaking, in a fifth aspect which can be combined with any of the first to fourth aspects, identifying a given data item includes identifying a given data item accessed according to a processing specification, and applying the one or more operations to the given data item includes inserting the one or more operations into the processing specification, and executing the processing specification to apply the one or more operations to the given data item.

[0010] Generally speaking, in a sixth aspect which can be combined with any of the first to fifth aspects, the processing specification includes a specification for a dataflow graph, and applying the one or more operations to a given data item includes executing the dataflow graph, wherein executing the dataflow graph applies the one or more operations to the given data item.

[0011] Generally speaking, in a seventh aspect which may be combined with any of the first to sixth aspects, the method further comprises updating a metadata model based on the transformed data item.

[0012] Generally speaking, in an eighth aspect which may be combined with any of the first to seventh aspects, updating the metadata model comprises adding one or more nodes to the metadata model, the added one or more nodes representing metadata of the transformed data item.

[0013] Generally speaking, in a ninth aspect which may be combined with any of the first to eighth aspects, updating the metadata model comprises adding one or more edges to existing nodes in the metadata model to connect the existing nodes in the metadata model to the added nodes.

[0014] Generally speaking, in a tenth aspect which can be combined with any of the first to ninth aspects, some of the added nodes are copies of some of the nodes already existing in the metadata model before the update.

[0015] Generally speaking, in an eleventh aspect which may be combined with any of the first to tenth aspects, the updated metadata model is output for further processing of the data item.

[0016] Generally speaking, in the twelfth aspect which can be combined with any aspect from the first to the eleventh aspect, the method also includes: identifying the physical metadata of the transformed data item; identifying a child node representing the physical metadata of the transformed data item in the updated metadata model; traversing one or more edges in the updated metadata model to identify one or more parent nodes of the child node representing the physical metadata of the transformed data item; and determining from the logical metadata associated with the one or more parent nodes identified by traversing the one or more edges in the updated metadata model that the one or more operations performed on the given data item will not be performed on the transformed data item.

[0017] Generally speaking, in the thirteenth aspect which can be combined with any of the first to twelfth aspects, determining that one or more operations performed on a given data item will not be performed on the transformed data item includes determining that logical metadata associated with one or more parent nodes identified by traversing one or more edges in the updated metadata model is different from the logical metadata associated with one or more parent nodes identified by traversing one or more edges in the metadata model.

[0018] Generally speaking, in the fourteenth aspect which can be combined with any aspect from the first to the thirteenth aspect, the method also includes: identifying the physical metadata of the transformed data item; identifying a child node representing the physical metadata of the transformed data item in the updated metadata model; traversing one or more edges in the updated metadata model to identify one or more parent nodes of the child node representing the physical metadata of the transformed data item; determining one or more second operations to be performed on the transformed data item from the logical metadata associated with the one or more parent nodes identified by traversing the one or more edges in the updated metadata model; applying the one or more second operations to the transformed data item by the metadata management system to further transform the transformed data item; and storing the further transformed data item in a memory.

[0019] Generally speaking, in a fifteenth aspect combinable with any of the first to fourteenth aspects, applying the one or more operations to a given data item comprises discarding data fields associated with the data item from further processing by the computer program.

[0020] Generally speaking, in a sixteenth aspect which may be combined with any of the first to fifteenth aspects, applying the one or more operations to a given data item comprises adding a data field to the data item or a data set associated with the data item.

[0021] Generally speaking, in a seventeenth aspect combinable with any of the first to sixteenth aspects, applying the one or more operations to a given data item comprises filtering data records associated with the data item.

[0022] Generally speaking, in the eighteenth aspect which can be combined with any aspect from the first to the seventeenth aspect, the one or more operations performed on the given data item include a tokenization operation, and applying the one or more operations to the given data item to transform the given data item includes applying the tokenization operation to the given data item to tokenize one or more fields of the given data item.

[0023] Generally speaking, in a nineteenth aspect which may be combined with any of the first to eighteenth aspects, the logical metadata associated with the one or more parent nodes comprises metadata received from a user through interaction with a metadata management system.

[0024] Generally speaking, in the twentieth aspect which can be combined with any one of the first to nineteenth aspects, the metadata received from the user is a personally identifiable information (PII) specification specified by the user through a graphical user interface, the graphical user interface including a visualization of a metadata model, and the PII classification is specified by the user interacting with the visualization of the metadata model to select metadata items associated with one or more parent nodes in the parent node for which the PII specification will be specified in the model, and the method also includes: after receiving the PII specification, updating the model to associate logical metadata with the one or more parent nodes, including indicating in the model that the one or more parent nodes in the parent node associated with the selected metadata items are associated with the PII specification, wherein the one or more operations applied to the given data item to transform the given data item include tokenizing one or more fields of the given data item.

[0025] Generally speaking, in the twenty-first aspect which can be combined with any one of the first to twentieth aspects, multiple data items stored in a hardware storage device are accessed; for each data item in the multiple data items, physical metadata and logical metadata corresponding to the data item are identified; a metadata model is generated based on the physical metadata and logical metadata identified for each data item in the multiple data items; access to the metadata model is provided for a first application and a second application; the metadata model is accessed by at least one of the first application or the second application, wherein each of the first application and the second application is configured to: access a given data item to identify the physical metadata of the given data item; identify a child node in the metadata model that presents the identified physical metadata of the given data item; traverse one or more edges in the metadata model to identify the one or more parent nodes of the identified child node; determine at least one operation to be performed on the given data item from the logical metadata associated with the one or more identified parent nodes; apply the at least one operation to the given data item to transform the given data item; and store the transformed data item.

[0026] Generally speaking, in aspect 22 which can be combined with any aspect from aspect 1 to aspect 21, the metadata model includes a first data structure corresponding to a parent node, the data structure including logical metadata and at least a first pointer and a second pointer, wherein the first pointer points to a second data structure corresponding to a child node representing physical metadata of a given data item, and the second pointer points to a third data structure corresponding to another child node representing physical metadata of another data item different from the given data item.

[0027] In general, in a twenty-third aspect, a system includes at least one processor and a memory storing instructions executable by the at least one processor to perform the operations of any one of the first to twenty-second aspects.

[0028] In general, in a twenty-fourth aspect, a non-transitory computer-readable medium stores instructions executable by at least one processor to perform the operations of any one of the first to twenty-second aspects.

[0029] These aspects may include one or more of the following advantages.

[0030] By using metadata to automate the identification of operations and their application to data, operations can be automatically applied across multiple data items within a data enterprise without requiring users to define operations on a data item-by-data item basis. This not only improves the efficiency of processing large amounts of continuously changing data (e.g., because operations do not need to be defined for each individual data item), but also enhances the robustness and integrity of data policies (e.g., data security policies, data governance policies, data quality policies, etc.) because operations can be automatically adjusted to account for changes in the underlying data, including the addition of new data. In addition, the incorporation of new operations is facilitated, such as to account for new data types or new data policies.

[0031] The technology described herein also provides for more efficient processing of data with reduced memory. Existing systems have stored metadata in a read-only format, which has become a major source of delay and inefficiency. This is because metadata is only used by analysts to understand the data, but they then manually obtain the information and use it in different projects as needed. Rather than treating metadata as read-only, the described technology uses metadata as the initial starting point in a connection chain for implementing the metadata using other applications to define the necessary processing of selected data (e.g., data sets) and then perform the processing. For example, a metadata management system can be connected to one or more other systems, such as a first system for writing a data flow diagram (and / or other computer program) that defines data processing operations, and a second system that is a data governance system. Under the old system, users can view metadata for various data sets. However, if a user wants to access the data, such as for defining processing to be performed in a first system and / or controlling processing to be performed by a second system, the data will have to be accessed and imported twice (once for each system) in order to access the metadata, which is in turn used to define data processing operations in the first system and specify control operations in the second system. Now, each data set only needs to be accessed once, rather than accessing the data multiple times (once for each application that needs metadata to define processing, control, etc.), thereby reducing the time and resources required to access the data set. The metadata is identified and provided with semantic meaning. A metadata model is then generated and continuously updated as additional metadata is created. The data processing system enables the metadata in the metadata model to be accessible to each of the applications or systems, resulting in a system in which the metadata of a data set only needs to be read once, and semantic discovery only needs to be performed once on the read metadata (or data corresponding to the read metadata) for use in multiple applications or systems. The single read, access, and processing of the metadata (via semantic discovery) results in a metadata model that is continuously updated with new metadata (e.g., by adding new nodes to the metadata model), and the metadata model can be accessed by various systems and applications for defining processing.

[0032] According to a preferred aspect, a user can specify operations to be performed on data within a data enterprise without defining (e.g., coding) means for accessing the data or performing the operations. For example, a user (e.g., a non-technical user) can specify that a particular item of logical metadata (e.g., an SSN) is a form of personally identifiable information (PII) without knowing which item(s) of physical data within the enterprise corresponds to the SSN, or how to access those data items. Based on the metadata definition, data items corresponding to the SSN are automatically identified by the system (e.g., by traversing a metadata model) and processed to obfuscate the data (e.g., by executing or inserting in a computer program operations that mask or tokenize the data), without the user having to generate code to perform these operations.

[0033] According to a preferred aspect, metadata generated by applying an operation to the data is propagated or copied to a metadata model to reduce further consumption of computing resources during subsequent processing based on the metadata model. That is, the technology described herein provides increased efficiency in processing data. This is because the metadata model is continuously updated with the results of previous processing. For example, if a new data item is generated due to an operation (e.g., tokenized SSN data is generated due to a tokenization operation), the system can update the metadata model to include metadata for the new data item. That is, nodes and edges are added to the metadata model, wherein the added nodes and edges represent the tokenized SSN data (e.g., the meaning of the tokenized data, the relationship between the tokenized data and other data or metadata, the storage location and other access parameters of the tokenized data, etc.). Therefore, if the data processing system needs the tokenized data at a later point in time, the data processing system does not need to re-tokenize the original data. Instead, the data processing system uses the metadata model to identify and access the tokenized SSN data. Saving metadata representing the tokenized SSN data in the metadata model saves computing resources because the data processing system can simply look up the tokenized data instead of having to recalculate it based on the original data.

[0034] Previously discovered or generated metadata may also be propagated or copied to new data items. That is, new edges may be added to the metadata model to associate nodes representing new data items with existing nodes in the model. For example, if a particular data item is associated with PII such as an SSN in a metadata model, and a new data item is created based on the particular data item (e.g., due to a copy operation), the system may automatically propagate the SSN association to the new data item in the metadata model (e.g., by adding an edge between the SSN node and the node representing the new data item). Therefore, if the system accesses the new data item at a later point in time, the system will know that the new data item is associated with the SSN, and if necessary, perform appropriate data quality and / or data security operations based on the association. By updating the metadata model in this manner, the system leverages existing work, such as work done to discover or generate metadata, in order to reduce the amount of computing resources (e.g., memory, processing cycles, etc.) required to perform subsequent operations on the data. In addition, propagating metadata in the metadata model ensures that data policies (e.g., data security policies, data governance policies, data quality policies, etc.) are followed when data changes are created within the system.

[0035] The details of one or more implementations are described in the accompanying drawings and the detailed description below.Other features, objects, and advantages of the techniques described herein will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 An example of a metadata management system is illustrated.

[0037] FIG. 2A to FIG. 2B An example metadata management system configured to implement metadata is illustrated.

[0038] FIG. 3A to FIG. 3D An example metadata management system configured to implement metadata is illustrated.

[0039] FIG. 4A to FIG. 4D An example metadata management system configured to implement metadata to discard fields is illustrated.

[0040] Figure 5A and Figure 5B An example metadata management system configured to implement metadata to add fields is illustrated.

[0041] Fig. 6A and Figure 6B An example metadata management system configured to implement metadata to filter rows is illustrated.

[0042] Fig. 7A and Figure 7B An example metadata management system configured to enforce metadata to adapt to changes in data is illustrated.

[0043] Figure 8 is an example procedure for using metadata to automate operations on data. DETAILED DESCRIPTION

[0044] The present disclosure relates to the use of metadata to automate the identification of operations and their application to data within a data enterprise. In some examples, a data processing system (sometimes referred to as a metadata management system) uses metadata of data stored within a data enterprise to generate a metadata model. The system can also enrich the metadata model with metadata specified by a user. When accessing data within a data enterprise, the system uses the metadata model to automatically identify operations and apply the operations to the data. In this way, users can specify operations to be performed on data within an enterprise without having to understand the technical details of the enterprise or generate code to perform the operations. In addition, because operations are defined at the metadata level, they can be automatically applied across multiple data items that are accessed or otherwise processed by multiple different applications without requiring users to define operations item by item in each application. Defining operations at the metadata level also enables the system to automatically adapt to changes in the underlying data, including the addition of new data. In some examples, updating the metadata model due to operations will improve the efficiency of subsequent data processing.

[0045] Generally speaking, a metadata management system may perform various processes to obtain metadata for data stored on one or more data sources within a data enterprise. For example, the system may discover physical metadata describing the attributes of data (e.g., where the system data resides, schema, table, field name, data type, data format, etc.) and the relationships between data (e.g., primary-foreign key relationships, entity relationships, etc.). The data processing system may also generate logical metadata based on the content of the stored data (or a selected subset of the stored data). Logical metadata provides detailed information about how data is linked together to form a larger set. It also outlines how data flows through systems and processes from creation, to storage, transformation, and consumption. Logical metadata may establish a roadmap on the path of data through the data supply chain, including its use and changes over time. For example, the system may determine that the stored data includes a data set related to information about a customer, and that the data set contains a data field that holds the social security number (SSN) of each customer.

[0046] Using the metadata, the system generates a metadata model that describes the physical and logical relationships and other properties of the stored data. Generally speaking, the metadata model may include nodes representing physical metadata items and logical metadata items, where edges represent relationships between nodes. The system can enrich the metadata model with metadata specified by the user (sometimes referred to as user-specified metadata). For example, a user can specify that an SSN is a form of personally identifiable information (PII) that should be protected by the system. Based on this specification, the system can update the metadata model to indicate that the node representing the SSN is associated with the PII.

[0047] The system then uses the metadata model to automatically perform operations on the data within the data enterprise. For example, when (e.g., by a computer program) accessing a data item within the enterprise, the system identifies a node in the metadata model corresponding to the metadata (e.g., physical metadata) of the data item being accessed. The system can then traverse the metadata model to find other nodes related to the identified node, as described in detail below. Based on the metadata (e.g., logical metadata) associated with the relevant node, the system determines one or more operations to be performed on the data item being accessed. For example, if the data item being accessed is related to an SSN node (which has been marked as PII as described above), the system can determine to automatically insert or otherwise perform an obfuscation or tokenization operation on the data item in a computer program. This enhances data security. In addition, the system can automatically perform operations (e.g., tokenization) on the data item based on the user's high-level metadata specification, without the user identifying the specific data item to be processed or generating code to perform the operation. This improves the efficiency of data processing.

[0048] In some examples, the system can update the metadata model based on the operations performed on the data items. For example, if a new data item is generated as a result of the tokenization operation described above, the system can update the metadata model to include metadata for the new data item. In some examples, previously discovered or generated metadata can also be propagated to the new data item, as described in detail below. By updating the metadata model in this manner, the system leverages existing work to reduce the amount of computing resources (e.g., memory, processing cycles, etc.) required to perform subsequent operations on the data.

[0049] refer to Figure 1, a system 100 for implementing metadata is shown. In this example, system 100 includes a metadata management system 102, storage systems 104, 106, and client devices 108, 110. Generally speaking, metadata management system 102 is a data processing system that uses metadata to automate the identification of operations and their application to data stored, for example, in storage system 104. After processing the data, metadata management system 102 may store the data in storage system 106, provide the data to client device 110 for display, or both. Although shown as separate entities, in some examples, storage system 104 may be the same storage system as storage system 106. Similarly, in some examples, client device 108 may be the same client device as client device 110.

[0050] The metadata management system 102 includes a metadata discovery engine 112 and a metadata-based processing engine 114. The discovery engine 112 includes program instructions and / or executable logic for discovering or otherwise obtaining metadata of data stored in the storage system 104. For example, the discovery engine 112 may perform a discovery process to obtain physical metadata describing attributes of the data (e.g., field names, data types, data formats, etc.) and relationships between the data (e.g., primary-foreign key relationships, entity relationships, etc.). In an example, the data is prepared for processing by the discovery engine 112 using format information. Data including records having values ​​of fields is received through an input device or port. A target record format for processing the data is determined. Multiple records are analyzed according to a validation test to determine whether the data matches a candidate record format. Each candidate record format specifies a format for each field, and each validation test corresponds to at least one candidate record format. In response to receiving the result of the validation test, the target record format is associated with the data based on at least one of the following: a candidate record format that is determined to be at least partially matched according to at least one validation test, a parsed record format selected according to a data type associated with the data, and a constructed record format generated from an analysis of data characteristics. Other examples of such discovery processes are described in US Patent Application No. 12 / 945,094, entitled "Managing record format information," which is incorporated herein by reference in its entirety.

[0051] The discovery engine 112 may also perform a semantic discovery process on the data (or a selected subset of the data) to generate logical metadata representing the semantic meaning of the data, etc. For example, the discovery engine 112 may identify a field included in one or more data sets, the field having an identifier. For the field, the discovery engine 112 profiles the data value of the field to generate a data profile, accesses multiple label proposal tests, and generates a set of label proposals by applying the multiple label proposal tests to the data profile. The discovery engine 112 then determines the similarity between the label proposals and selects a classification. The discovery engine 112 identifies one of the label proposals as identifying a semantic meaning. The discovery engine 112 stores an identifier of the field of the identified one of the label proposals having the identifying semantic meaning in the label proposal. Other examples of such semantic discovery are described in U.S. Patent Application No. 16 / 794,361, entitled "Discovering a semantic meaning of data fields from profile data of the data fields", the entire contents of which are incorporated herein by reference.

[0052] The discovery engine 112 passes the physical metadata and the logical metadata to a metadata-based processing engine 114, which uses the metadata to generate a metadata model 116 having a plurality of nodes 117 and edges 118 representing relationships between the nodes. The metadata-based processing engine 114 also enriches the metadata model 116 with user-defined metadata received from, for example, the client device 108. In some examples, the metadata model 116 may include a set of objects or data structures, each of which represents a node. Each object or data structure may include data elements representing physical metadata, logical metadata, and / or user-defined metadata of the corresponding node, as well as pointers to other objects or data structures representing other nodes connected to the corresponding node by edges. In some examples, the metadata-based processing engine uses the metadata model and other data to generate a data catalog 120 of data stored in the storage system 104 (or a selected subset of the stored data). The data catalog 120 may include one or more data objects that contain metadata and other information identifying data or data groups stored in the storage system 104. A user may interact with the data catalog 120 to define properties of an object or select an object for data processing, etc. For example, a user may associate objects in the data catalog 120 with one or more metadata-driven transformations 122, as discussed below. As another example, a development environment (not shown) that is part of or in communication with the metadata management system 102 may include a user interface with a representation of the catalog 120, and a user may select an object from the catalog to use, for example, as input to a data flow diagram or other computer program. Techniques for generating, maintaining, and using data catalogs are described in U.S. Patent No. 9,977,659, entitled "Managing Data Set Objects," which is incorporated herein by reference in its entirety.

[0053] The metadata-based processing engine 114 may also store or otherwise access a plurality of processing specifications 124. In general, the processing specifications 124 may define or include program instructions and / or executable logic for processing data stored in the storage system 104. In some examples, each of the processing specifications 124 may be or otherwise define a computer program, such as a data flow graph. The data flow graph may include: a plurality of vertices representing computational processes, each vertex having an associated access method; and a plurality of links, each link connecting at least two vertices to each other and representing a flow of data between the connected vertices. When executed, a system executing the graph (e.g., the metadata management system 102 or another data processing system) may prepare the graph for execution by performing graph transformation steps until each vertex is in a runnable state, and each link is associated with at least one communication method that is compatible with the access method of the vertices connected by the link; each link is initiated by creating, by means of the execution system, a combination of a communication channel and / or a data store suitable for the link communication method; and each process is initiated by invoking execution of the process on the execution system. Additional details regarding specific implementations of such graph-based computations are described in U.S. Patent No. 5,966,072, entitled “Executing Computations Expressed as Graphs,” the entire contents of which are incorporated herein by reference. The processing specification 124 may define or include operations for accessing data in the data catalog 120 (e.g., from the storage system 104) for any of a variety of reasons, such as ingesting the data into the storage system 106 or generating views of the data for presentation on the client device 110. Regardless of the specific process defined by the processing specification 124, the metadata-based processing engine 104 may use the metadata model 116 to automatically perform operations (e.g., metadata-driven transformations 122) on the data accessed from the storage system 104, as described herein.

[0054] Figure 2A and Figure 2B A system 200 for implementing metadata according to aspects of the present disclosure is illustrated. Figure 2A , system 200 includes a metadata management system 202, storage systems 204, 206, and client devices 208, 210. Figure 1, the metadata management system 202 includes a metadata discovery engine 212 and a metadata-based processing engine 214. In this example, the metadata discovery engine 212 receives a data set 230 stored in the storage system 204 and processes the data set 230 to obtain corresponding metadata. Specifically, the metadata discovery engine 212 determines that the data set 230a includes the following items of physical metadata: name "Cust_data", type "table", and fields "cust_fnln" and "cust_ssn".

[0055] A user may access the metadata management system 202 (e.g., via an application executing on a client device 208) to view and interact with a graphical user interface 232 that includes physical metadata for data sets 230. Within the user interface 232, the user may select 234 data sets 230 to add to a data cart 236 based on the physical metadata. The metadata discovery engine 212 may then perform semantic discovery or other processing on the data sets 230 within the cart 236 to generate logical metadata that represents the semantic meaning of the data sets 230. For example, the metadata discovery engine 212 may determine that the data set 230a includes information about a customer, that the field "cust_fnln" includes the name of the customer, and that the field "cust_ssn" includes the SSN of the customer. In some examples, the user may approve the meaning associated with the data set 230a and its fields, and submit the data set for entry into a data catalog 220 maintained by a system (not shown).

[0056] The metadata 238 discovered by the metadata discovery engine 212 is passed to the metadata-based processing engine 214. As described above, the discovered metadata 238 may include physical metadata and logical metadata (sometimes referred to as logical item associations) for the selected data set 230. The metadata-based processing engine 214 uses the metadata 238 to generate a metadata model 216a for the data set 230. In this example, the metadata model 216a includes oval nodes 217a, 217b, and 217c representing physical metadata for the data set 230a, and rectangular nodes 217d, 217e, and 217f representing logical metadata for the data set 230a. The metadata model 216a also includes edges 218a, 218b representing the physical relationship between the dataset 230a (e.g., node 217a) and its fields (e.g., nodes 217b, 217c), edge 218c representing the logical relationship between the dataset 230a and its corresponding logical entity (e.g., the logical dataset "Customer"), and edges 218d, 218e representing the logical relationship between each of the fields of the dataset 230a and its corresponding logical entity (e.g., the logical data elements "Name" and "SSN").

[0057] The metadata-based processing engine 214 also receives user-specified metadata 240 from a user of the client device 210. In some examples, the metadata 240 may be specified with respect to a physical metadata item or a logical metadata item in the metadata model 216a. For example, in this example, the user-specified metadata 240 indicates that a logical metadata item (e.g., a logical data element SSN) representing an SSN is associated with PII. Once received, the metadata-based processing engine 214 may incorporate the user-defined metadata 240 into the metadata model 216a. For example, the metadata-based processing engine 214 may update the metadata model 216a to indicate 219 that a node 217f representing an SSN is associated with PII. This association may be performed by, for example, including a metadata definition or other indication identifier with the SSN logical metadata, or by adding another node linked by an edge to the node representing the SSN to the metadata model 216a, etc.

[0058] refer to Figure 2B, the metadata-based processing engine 214 can use the metadata model 216a to automatically perform operations on the data sets 230 stored in the storage system 204. In this example, the metadata-driven data transformation 222 associated with the data in the directory 220 (including the data set 230) specifies that if the data contains PII, the data should be tokenized to enhance data security. Therefore, when the metadata-based processing engine 214 executes the processing specification 224, it will use the metadata model 216a (and the transformation 222) to determine whether some or all of the data sets 230 being accessed are associated with PII, and if so, perform a tokenization operation on the data. For example, in this example, one of the processing specifications 224 specifies a data pipeline configured to ingest the data set 230a into the storage system 206 (e.g., a data pipeline configured to read the data set 230a from the storage system 106 and write the data set to the storage system 104). Therefore, the metadata-based processing engine 214 determines the physical metadata of the data set 230a and identifies the node in the metadata model 216a that corresponds to the physical metadata of the data set 230a. For example, the metadata-based processing engine 214 can determine that the data set 230a has a name "Cust_data" and can match the name with the node 217a in the metadata model 216a. The metadata-based processing engine 214 can then traverse the metadata model 216a to determine whether the data set 230a contains PII. In this example, the metadata-based processing engine 214 determines that the field "cust_ssn" (represented by node 217c) of the data set 230a is associated with the logical data element SSN (represented by node 217f) that has been defined as PII. Therefore, the metadata-based processing engine 214 tokenizes the data in the "cust_ssn" field of the data set 230a to create a "Cust_data_tokenized" data set 242 for ingestion into the storage system 206.

[0059] In this example, another of the processing specifications 224 specifies an operation configured to generate a catalog view of the data set 230a for presentation to a user of the client device 210. The metadata-based processing engine 214 traverses the metadata model 216a as described above to determine that the field "cust_ssn" of the data set 230a includes PII. Accordingly, the metadata-based processing engine 214 tokenizes the data in the "cust_ssn" field before sending the catalog view data 244 to the client device 210. Once received, the client device 210 can present the data 244 in a graphical user interface 246 that allows the user to view the tokenized data records of the data set 230a and use the tokenized data as a component in subsequent processing.

[0060] In some examples, the metadata-based processing engine 214 can update the metadata model 216a to generate a metadata model 216b that incorporates the metadata of the newly defined "cust_data_tokenized" data set 242. For example, the metadata-based processing engine 214 can add nodes 217g, 217h, and 217i that represent the physical metadata of the data set 242 and edges 218f, 218g that represent the relationship between the nodes 217g, 217h, and 217i. Since the "cust_ssn" field of the data set 242 has been tokenized and no longer represents an SSN, the metadata-based processing engine 214 can also add a logical node 217j (e.g., "tokenized SSN") that accurately describes the content of the tokenized field. The edge 218h can then be inserted to indicate the logical relationship between the nodes 217i and 217j. The metadata-based processing engine 214 can also copy the existing metadata in the metadata model 216b (see, e.g., nodes 217h, 217i) to populate the metadata of the new data set 242. Specifically, metadata-based processing engine 214 may insert edge 218i representing a logical relationship between data set 242 and its corresponding logical entity (e.g., logical data set "customer"), and edge 218j representing a logical relationship between a "cust_fnln" field of data set 242 and its corresponding logical entity (e.g., logical data element "name"). By updating the metadata model in this manner, the computational resources required for subsequent processing of data set 242 may be reduced because the model has already been populated with metadata for data set 242 when some of metadata 238 discovered for data set 242 is replicated (thereby eliminating the need to perform metadata discovery on data set 230), and because the model has been updated to reflect that the "cust_ssn" field of data set 242 has already been tokenized (and therefore does not need to be re-tokenized during subsequent processing).

[0061] Figure 3AA system 300 for implementing metadata according to aspects of the present disclosure is illustrated. In this example, the system 300 includes a metadata management system 302, storage systems 304, 306, and client devices 308, 310. Similar to the metadata management systems 102 and 202, the metadata management system 302 includes a metadata discovery engine 312 and a metadata-based processing engine 314. In this example, the metadata discovery engine 312 receives data sets 330a-330d stored in the storage system 304 and processes the data sets 330a-330d to obtain corresponding physical metadata and logical metadata 332. Specifically, the metadata discovery engine 312 determines that the data set 330a has a name "Cust_data" and includes fields "cust_id" (which serves as a primary key, as indicated by the key symbol) and "fname". The metadata discovery engine 312 also determines, based on a semantic analysis of the dataset 330a and its data, that the dataset 330a "Cust_data" represents information about a "customer", that the field "cust_id" represents a "customer ID" as part of a broader group "DB identifier", and that the field "fname" represents the "name" of the customer as part of a broader group "customer identity". The metadata discovery engine 312 also determines that the dataset 330a is related to the dataset 330b (e.g., through a primary-foreign key relationship), and that the dataset 330b has the name "SSN_table" and includes the fields "I95" and "S14". Based on a semantic analysis of the dataset 330b and its data, the metadata discovery engine 312 determines that the dataset 330b represents a "customer SSN", that the field "I95" represents a "customer ID" as part of the group "DB identifier", and that the field "S14" represents the "SSN" of the customer as part of the group "customer identity". The metadata discovery engine 312 also determines that the dataset 330b is related to the dataset 330c (e.g., through a primary-foreign key relationship), and that the dataset 330c has a name of "Loan_data" and includes fields "cust_ssn" and "amt." Based on the semantic analysis of the dataset 330c, the metadata discovery engine 312 determines that the dataset 330c represents "Customer Loan," the field "cust_ssn" represents the customer's "SSN" as part of the group "Customer Identity," and the field "amt" represents "Amount" within the group "Loan Information." The metadata discovery engine 312 determines that the dataset 330c is related to the dataset 330d (e.g., through a primary-foreign key relationship), and that the dataset 330d has a name of "Cust_acct" and fields "c_id" and "level."The metadata discovery engine 312 also determines based on semantic analysis of dataset 330d and its data that dataset 330d represents information about a "customer account", field "c_id" represents a "customer ID" as part of the group "DB identifier", and field "level" represents the customer's "account level".

[0062] The metadata discovery engine 312 passes the discovered metadata 332 to the metadata-based processing engine 314. The metadata-based processing engine 314 uses the metadata 332 to generate a metadata model 316 including a plurality of nodes 317 and edges 318. In this example, the oval nodes 317 represent physical metadata items of the data sets 330a to 330d, and the rectangular nodes 317 represent logical metadata items of the data sets 330a to 330d. In some examples, the metadata-based processing engine 314 also uses the metadata 332 to add the data sets 330a to 330d to the data catalog 320. The data sets 330a to 330d may also be associated (e.g., within the data catalog 320 or otherwise) with one or more metadata-driven data transformations 322. The metadata-based processing engine 314 also stores or otherwise accesses a plurality of processing specifications 324. In this example, the processing specification 324 includes a data pipeline 324 a - 324 d in the form of a data flow graph that is configured to access (eg, read) data sets 330 a - 330 d from the storage system 306 and store (eg, write) them in the storage system 304 .

[0063] refer to Figure 3B , the metadata-based processing engine 314 can receive a metadata specification 334 from a user of the client device 308. Specifically, the user can use the client device 308 to access a graphical user interface 336 including a visualization 336a of the metadata model 316. The user can interact with the visualization 336a of the metadata model to select a metadata item for which metadata will be specified. In this example, the user has selected 336b a metadata item corresponding to "customer identity". After selecting the metadata item, the user can specify metadata for the metadata item. For example, the user can specify a PII classification 336c for the metadata item, which is a form of PII specification. In this example, the user has specified that "customer identity" has a PII classification 336c of "level 2 (tokenization)". Once completed, the user can choose to submit 336d the metadata specification 334 to the metadata-based processing engine 314.

[0064] After receiving the metadata specification 334, the metadata-based processing engine 314 updates the model 316. For example, the metadata-based processing engine 314 can update the model 316 to indicate 319 that the node representing "customer identity" is associated with the level 2 PII classification. As shown in the visualization 322a of the metadata-driven transformation 322, the metadata-based processing engine 314 is configured to tokenize the fields associated with the node when the related node (e.g., the parent node) has a PII classification of "level 2". Incorporating the metadata specification 334 into the metadata model 316 in this way reduces the amount of space required to store the metadata model 316. This is because the metadata specification 334 is defined only once in the model 316 (e.g., by adding an element with the specification 334 in the object or data structure representing "customer identity", or by adding a single new object or data structure with the specification 334 pointing to the object or data structure representing "customer identity"), rather than defining the specification at each applicable node (which can amount to thousands of definitions or more). Thus, only a single data element (or object or data structure) is added to the model 316 (as opposed to many data elements), which reduces the storage and memory requirements of the model. A similar reduction in storage is also achieved by linking multiple nodes representing physical metadata (e.g., "cust_ssn" and "S14" nodes) to a single node representing logical metadata (e.g., "SSN" node), relative to defining and storing a separate logical node for each physical node.

[0065] refer to Figure 3C, the metadata-based processing engine 314 uses the metadata model 316 to apply the metadata-driven data transformation 322 to the processing specification 324, and therefore to the data sets 330a to 330d. To do this, the metadata-based processing engine 314 can determine the physical metadata of the data item being accessed, and can use the determined physical metadata to identify the node in the metadata model 316 that corresponds to the data item. For example, when executing the processing specification 324a, the metadata-based processing engine 314 can identify the node in the metadata model 316 that corresponds to the name of the data set 330a (e.g., "Cust_data") and / or the name of the field of the data set 330a (e.g., "cust_id", "fname"). In some examples, the identified nodes are referred to as child nodes. The metadata-based processing engine 314 then traverses the metadata model 316 according to the metadata-driven transformation 322 to determine the operations to be performed on the data set 330a (if any). For example, the metadata-based processing engine 314 may traverse the metadata model 316 to identify one or more parent nodes of the identified child nodes, such as nodes corresponding to "customer", "customer ID", "DB identifier", "name", and "customer identity". In this example, the metadata-based processing engine 314 determines based on the metadata model 316 that the node representing "customer identity" associated with the "level 2" PII classification is the parent node of the node representing "fname" (e.g., through the "name" node), but is not the parent node of the node representing "cust_id". Therefore, the metadata-based processing engine 314 inserts operation 325a into the processing specification 324a to tokenize "fname". The metadata-based processing engine 314 may perform a similar process to insert tokenization operations 325b, 325c into the processing specifications 324b, 324c. However, a tokenization operation is not added to the processing specification 324d because there is no parent node associated with the PII classification of "level 2" for the node corresponding to the dataset "Cust_acct" or its fields "c_id" and "level".

[0066] After applying the metadata-driven transformation 322, the metadata-based processing engine 314 can execute the processing specification to generate data sets 338a to 338d (where applicable) with tokenized PII data. Each of the data sets 338a to 338d can be sent to the storage system 306 for storage. The metadata-based processing engine 314 can also apply similar processing to generate a tokenized catalog view data 340 (etc.) of the data set 330a. The catalog view data 340 can be sent to the client device 310 for presentation in a graphical user interface 342 that allows a user to view each of the tokenized data records of the data set 330a.

[0067] By leveraging metadata in this manner, system 300 enables users to automatically implement operations such as data security operations across an entire data enterprise (e.g., across multiple data items, such as data sets 338a to 338d and other data items) through a single global metadata specification (e.g., metadata specification 334) without requiring users to identify the specific data items to which these operations apply, and without requiring users to generate code on a data item-by-data item basis. Therefore, relative to systems that do not leverage metadata in this manner, system 300 provides more efficient implementation of data policies across an enterprise, thereby reducing latency in implementing such policies. In addition, by automatically identifying and applying operations based on metadata, system 300 enhances data security (and implementation of other data policies) by reducing the likelihood that data that should be subject to policy will be overlooked. Defining operations based on metadata also enables the system to automatically adapt to changes in the underlying data (e.g., changes in data names, data storage locations, data keys, etc.).

[0068] The metadata-based processing engine 314 may also update the metadata model 316 after applying the metadata-driven transformation 322, such as Figure 3D . For example, the metadata-based processing engine 314 can update the metadata model 316 to incorporate metadata for the newly defined "cust_data_tokenized" data set 338a. For example, the metadata-based processing engine 314 can add nodes 317a, 317b, and 317c representing the physical metadata of the data set 338a, and edges 318a, 318b representing the relationship between the nodes 317a, 317b, and 317c. Since the "cust_ssn" field of the data set 338a has been tokenized and no longer represents an SSN, the metadata-based processing engine 314 can also add logical nodes 317d, 317e that accurately describe the contents of the tokenized field (e.g., "tokenized name" and "tokenized customer identity"). Edges 318c, 318d can then be inserted to indicate the logical relationship between nodes 317c, 317d, and 317e.

[0069] The metadata-based processing engine 314 may also propagate or copy existing metadata in the metadata model 316 to populate the metadata of the new data set 338a. Specifically, the metadata-based processing engine 314 may insert an edge 318e representing a logical relationship between the data set 338a and its corresponding logical entity (e.g., the logical data set "Customer"), and an edge 318f representing a logical relationship between the "c_id" field of the data set 338a and its corresponding logical entity (e.g., the logical data element "Customer ID"). Similar operations may be performed to update the metadata model 316 to incorporate metadata of data sets 338b to 338d (not shown).

[0070] By updating the metadata model in this manner, the computational resources required for subsequent processing of the data sets 338a-338d can be reduced because the model is already populated with the metadata for the data sets 338a-338d, thereby avoiding the need for metadata discovery for those data sets. Additionally, because the model has been updated to reflect that the "cust_ssn" field of the data set 338a (as well as other fields of other data sets) has been tokenized, the metadata-based processing engine 314 can use previously tokenized data during subsequent processing instead of re-tokenizing.

[0071] Figure 4A A system 400 is illustrated, which is FIG. 3A to FIG. 3D . In this example, system 400 includes a metadata management system 402, storage systems 404, 406, and client devices 408, 410. Similar to metadata management systems 102, 202, and 302, metadata management system 402 includes a metadata discovery engine 412 and a metadata-based processing engine 414. Metadata-based processing engine 414 stores or otherwise accesses multiple processing specifications 424 including processing specifications 424a. In this example, processing specifications 424a illustrate a data flow graph that defines a directory view of multiple data sets (e.g., data sets 330a to 330c) in a data directory 420 from the perspective of a "Cust_data" data set 330a. In other words, data set 330a is a root node in the directory view defined by processing specification 424a, where other data sets are connected to data set 330a based on the relationships discovered between data sets 330a to 330c.

[0072] The metadata-based processing engine 414 uses the techniques described herein to automatically convert metadata-driven transformations 322 (in Figure 3B ) is applied to the processing specification 424a. Specifically, the metadata-based processing engine 414 uses the metadata model 316 (shown in Figure 3B ) to determine that the fields "fname", "S14", and "cust_ssn" should be tokenized, and insert the tokenization operation 425a into the processing specification 424. After applying the metadata-driven transformation 322, the metadata-based processing engine 414 can execute the processing specification 424a to generate tokenized catalog view data 430. The catalog view data 430 can be sent to the client device 410 for presentation in a graphical user interface 432, which allows the user to view the tokenized data records in a customer view. The metadata-based processing engine 414 can also generate a customer view data set 434 (e.g., a wide record of a customer view) that is sent to the storage system 406 for storage.

[0073] In some examples, the metadata-based processing engine 414 can receive a new metadata-driven transformation 422a (or a modification to an existing metadata-driven transformation) from a user of the client device 408, such as Figure 4B . Specifically, a user can use a client device 408 to access a graphical user interface 436 for defining a metadata-driven transformation 422a. The user can interact with the user interface 436 to enter a name 436a of the metadata-driven transformation 422a, one or more conditions 436b for applying the transformation, and one or more expressions 436c specifying the operation to be performed when the condition 436b is met. In this example, the user has defined a metadata-driven transformation 422a named "Administrator View", which causes the field to be discarded (e.g., removed to avoid further processing) when the PII classification of the parent node is equal to "Level 2" and the PII viewing permission level of the parent node is greater than "User". In some examples, the user can also define the item or item group (not shown) to which the transformation 422a is applied in the data directory 420. After defining the metadata-driven transformation 422a, the user can choose to submit 436d the metadata-driven transformation 422a to the metadata-based processing engine 414. The metadata-based processing engine 414 may then incorporate the metadata-driven transformation 422a into the transformation set 422 maintained by the system 400 , as shown in the visualization 422b of the metadata-driven transformation 422 .

[0074] refer to Figure 4C , the metadata-based processing engine 414 may receive a metadata specification 438 from a user of the client device 408. Specifically, the user may use the client device 408 to access a metadata model 316 (eg, Figure 3B 4 (shown). A user may interact with the visualization 440a of the metadata model to select metadata items for which metadata will be specified. In this example, the user has selected 440b the metadata item corresponding to "Customer Identity." After selecting the metadata item, the user may specify metadata for the metadata item. For example, a user may specify PII viewing permissions 440d for the metadata item, which is a form of PII specification. In this example, the user has specified that "Customer Identity" has a PII viewing permission 440d of "Administrator Level" (in addition to the previously specified PII classification 440c of "Level 2 (Tokenization)"). Once completed, the user may choose to submit 440e the metadata specification 438 to the metadata-based processing engine 414. After receiving the metadata specification 438, the metadata-based processing engine 414 updates the model 316 ( Figure 3B) to generate model 416. For example, metadata-based processing engine 414 can update model 416 to indicate 419 that a node representing "customer identity" is also associated with "administrator-level" PII viewing permission.

[0075] refer to Figure 4D , the metadata-based processing engine 414 can use the metadata model 416 to automatically apply the metadata-driven transformation 422, including the newly defined metadata-driven transformation 422a, to the processing specification 424 and the underlying data. In this example, the metadata-based processing engine 414 identifies the node 417a corresponding to the "Cust_data" data set accessed in the processing specification 424a. The metadata-based processing engine 414 then traverses the metadata model 416 as shown by the directional arrows to identify nodes 417b, 417c corresponding to the fields of the "Cust_data" data set and nodes 417d to 417h corresponding to the logical parent nodes of the physical nodes. In this example, the metadata-based processing engine 414 determines that the field "fname" corresponding to node 417c should be discarded (e.g., from further processing) because one of its parent nodes (e.g., node 417f, representing "Customer Identity") is associated with a "Level 2" PII classification (a first condition of metadata-driven transformation 422a), and is also associated with an "Administrator Level" PII viewing permission that is greater than a "User Level" viewing permission (a second condition of metadata-driven transformation 422a). The determination to discard the field "fname" is made in Figure 4D 417b, and the traversal of the metadata model 416 to make this determination is shown by the dashed edge with a directional arrow. Because the field "fname" is associated with PII (through its parent node 417f), the metadata-based processing engine 414 can also determine that the data contained within this field should be based on the "tokenization" metadata-driven transformation 422 ( Figure 4B However, the metadata-based processing engine 414 may determine that the tokenized data within the field is inconsistent with completely discarding the field, and may decide that the discard operation has priority based on rules maintained by the system 400, such as optimization rules.

[0076] On the other hand, the metadata-based processing engine 414 determines that the field "cust_id" should not be discarded because it does not have any parent nodes that satisfy the two conditions specified by the metadata-driven transformation 422a. Similarly, the metadata-based processing engine 414 determines that the data within the field "cust_id" should not be tokenized because the data does not have any parent nodes with a "level 2" PII classification. The determination to retain (and not tokenize) the field "cust_id" is made in Figure 4D417a. The determination of the metadata model 416 is indicated by the check mark and thick outline of node 417c, and the traversal of the metadata model 416 to make this determination is indicated by the thick edge with a directional arrow. Similar analysis of the metadata model 416 results in the determination to retain the fields "I95" and "amt", and to discard the fields "cust_ssn" and "S14". Note that the nodes for the dataset "Cust_acct" and related metadata are downplayed in this example because they are not relevant to the processing specification 424a.

[0077] The metadata-based processing engine 414 may update the processing specification 424a using the information determined by traversing the metadata model 416. For example, because the metadata-based processing engine 414 determined from the metadata-driven transformation 422a that the fields "fname", "S14", and "cust_ssn" should be discarded, the metadata-based processing engine 414 may update the processing specification to remove or modify operations that access those fields. In some examples, after applying these modifications, the metadata-based processing engine 414 may optimize the processing specification 424a to remove redundant or unnecessary operations. For example, in this example, the metadata-based processing engine 414 determines that the "read Cust_data" operation 425b and the join operation 425c are no longer necessary, and may optimize these operations as a result (as indicated by the "X" for the corresponding operations).

[0078] After applying the metadata-driven transformation 422, the metadata-based processing engine 414 can execute the processing specification 424a to generate tokenized catalog view data 442. The catalog view data 442 can be sent to the client device 410 for presentation in a graphical user interface 444, which allows the user to view the tokenized data records in a customer view. The metadata-based processing engine 414 can also generate a customer view data set 446 (e.g., a wide record of the customer view) that is sent to the storage system 406 for storage. As shown in the view data 442 and the data set 446, the "fname", "S14", and "cust_ssn" fields (along with other redundant or unnecessary fields) have been discarded.

[0079] Figure 5A A system 500 is illustrated, which is FIG. 4A to FIG. 4D. In this example, system 500 includes a metadata management system 502, storage systems 504, 506, and client devices 508, 510. Similar to other metadata management systems described herein, metadata management system 502 includes a metadata discovery engine 512 and a metadata-based processing engine 514. In this example, metadata-based processing engine 514 receives a new metadata-driven transformation 522a from a user of client device 508. Specifically, a user can use client device 508 to access a graphical user interface 530 for defining metadata-driven transformation 522a. The user can interact with user interface 530 to enter a name 530a of metadata-driven transformation 522a, one or more conditions 530b for applying the transformation, and one or more expressions 530c specifying the operation to be performed when condition 530b is met. In this example, the user has defined a metadata-driven transformation 522a named "Add Field", which causes a field to be inserted into a data set when the PII classification of the parent node is equal to "Level 1" or "Level 2". After defining the metadata-driven transformation 522a, the user may choose to submit 530d the metadata-driven transformation 522a to the metadata-based processing engine 514. The metadata-based processing engine 514 may then incorporate the metadata-driven transformation 522a into the transformation collection 522 maintained by the system 500, as shown in the visualization 522b of the metadata-driven transformation 522.

[0080] refer to Figure 5B , the metadata-based processing engine 514 can use the metadata model 516 to automatically apply the metadata-driven transformation 522 to the processing specification 524 and the underlying data. In this example, the metadata-based processing engine 514 identifies the node 517a corresponding to the "Cust_data" data set accessed in the processing specification 524a. The metadata-based processing engine 414 then traverses the metadata model 516 as shown by the directional arrows to identify the nodes 517b, 517c corresponding to the fields of the "Cust_data" data set and the nodes 517d to 517h corresponding to the logical parent nodes of the physical nodes. In this example, the metadata-based processing engine 514 determines that the field "fname" corresponding to the node 517c has a parent node 517f (representing "customer identity") associated with the "level 2" PII classification. Therefore, the metadata-based processing engine 514 determines that the PII field should be added to the data set (e.g., "Cust_data") containing "fname" according to the metadata-driven transformation 522a. Similarly, the metadata-based processing engine 514 also determines that the data within the "fname" field should be transformed based on the "tokenization" metadata driven transformation (in Figure 5BThe determination of adding the PII field and tokenizing the data in the field "fname" is made in Figure 5B The determination is indicated by the check mark and thick outline of node 517b, and the traversal of metadata model 516 to make this determination is indicated by the thick edge with a directional arrow.

[0081] On the other hand, the metadata-based processing engine 514 determines from the metadata-driven transformation 522a that the PII field should not be added to the data set containing the field "cust_id" because the data set does not have any parent node associated with PII. For the same reason, the metadata-based processing engine 514 determines that the data within "cust_id" does not need to be tokenized. The determination not to add the PII field and not to tokenize the data within the "cust_id" field is made in Figure 5B 517c, and the traversal of metadata model 516 to make this determination is indicated by the dashed edge with a directional arrow. Similar analysis of metadata model 516 results in a determination to add PII fields and tokenize the data within the "cust_ssn" and "S14" fields, and not to add PII fields and not to tokenize the data within the "I95" and "amt" fields. Note that the nodes for the dataset "Cust_acct" and related metadata are downplayed in this example because they are not relevant to processing specification 524a.

[0082] The metadata-based processing engine 514 can update the processing specification 524a using the information determined by traversing the metadata model 516. For example, because the metadata-based processing engine 514 determines that a PII field should be added to the data set containing the "fname", "S14", and "cust_ssn" fields and the data within these fields should be tokenized, the metadata-based processing engine 514 can update the processing specification to include a tokenize operation 525a for tokenizing these fields and an add PII operation 525b for adding a PII field with a value of "yes" to each corresponding data set. In some examples, the metadata-based processing engine 514 can optimize the processing specification 524a to remove redundant or unnecessary operations, such as by adding only a single PII field to the generated data set (e.g., Figure 5B ).

[0083] After applying the metadata driven transformation 522, the metadata based processing engine 514 can execute the processing specification 524a to generate the tokenized catalog view data 532. The catalog view data 532 can be sent to the client device 510 for presentation in the graphical user interface 534, which allows the user to view the tokenized data records in the customer view. The metadata based processing engine 514 can also generate a customer view data set 536 (e.g., a wide record of the customer view) that is sent to the storage system 506 for storage. As shown in the view data 532 and the data set 536, the PII field with the value "yes" has been added.

[0084] Fig. 6A A system 600 is illustrated, which is FIG. 5A to FIG. 5B 6. In this example, system 600 includes a metadata management system 602, storage systems 604, 606, and client devices 608, 610. Similar to other metadata management systems described herein, metadata management system 602 includes a metadata discovery engine 612 and a metadata-based processing engine 614. In this example, metadata-based processing engine 614 receives a new metadata-driven transformation 622a from a user of client device 608. Specifically, the user can use client device 608 to access a graphical user interface 630 for defining metadata-driven transformation 622a. The user can interact with user interface 630 to enter a name 630a of metadata-driven transformation 622a, one or more conditions 630b for applying the transformation, and one or more expressions 630c specifying an operation to be performed when condition 630b is met. In this example, the user has defined a metadata-driven transformation 622a named "PII View Permissions" that causes rows (e.g., data records) to be filtered from a data set when the PII classification of the parent node is equal to "Level 2" and the customer's account level is greater than the user's level (e.g., the level of the user accessing the data). After defining the metadata-driven transformation 622a, the user may choose to submit 630d the metadata-driven transformation 622a to the metadata-based processing engine 614. The metadata-based processing engine 614 may then incorporate the metadata-driven transformation 622a into the transformation set 622 maintained by the system 600, as shown in a visualization 622b of the metadata-driven transformation 622.

[0085] refer to Figure 6B, the metadata-based processing engine 614 can use the metadata model 616 to automatically apply the metadata-driven transformation 622 to the processing specification 624 and the underlying data. In this example, a user with a non-administrator user level has requested access to the customer view defined by the processing specification 624a. Because the metadata-driven transformation 622a requires a runtime comparison of the user's level with the customer's account level to determine whether to filter rows, the metadata-based processing engine 614 adds a read operation 625a for reading the "Cust_acct" data set and a join operation 625b for joining the "Cust_acct" data set to the "Loan_data" data set. In some examples, the metadata-based processing engine 614 can use the metadata included in the metadata model 616 (or data catalog) to determine how to access the "Cust_acct" data set and join it with an existing data set (e.g., the "Loan_data" data set). The metadata-based processing engine 614 can also insert a filter row operation 625c to filter rows when the condition "Cust_acct.level>user's level" is met. The metadata-based processing engine 614 also adds a tokenization operation 625d to tokenize the data according to a “tokenization” metadata-driven transformation 622, as described herein.

[0086] After applying the metadata-driven transformation 622, the metadata-based processing engine 614 can execute the processing specification 624a to generate the catalog view data 632. The catalog view data 632 can be sent to the client device 610 for presentation in a graphical user interface 634, which allows the user to view the data records in the customer view when the filter condition is met. The metadata-based processing engine 614 can also generate a customer view data set 636 (e.g., a wide record of the customer view) that is sent to the storage system 606 for storage. As shown in the data set 636, when the filter condition is met, the row is filtered out (as shown by the row running through). It is noted that although the metadata-based processing engine 614 has added an operation 625a to access the "Cust_acct" data set for the purpose of the filter condition, the data associated with the "Cust_acct" data set is not output according to the processing specification 624a.

[0087] Fig. 7A A system 700 is illustrated, which is FIG. 6A to FIG. 6B In this example, system 700 includes a metadata management system 702, storage systems 704, 706, and client devices 708, 710. Similar to other metadata management systems described herein, metadata management system 702 includes a metadata discovery engine 712 and a metadata-based processing engine 714.

[0088] In this example, metadata discovery engine 712 receives new dataset 730 from storage system 704. Metadata discovery engine 712 processes the dataset as described herein to discover or otherwise obtain metadata 732 for new dataset 730. Specifically, metadata discovery engine 712 determines that dataset 730 has a name "Cust_DOB" and includes fields "I95" (which serves as a primary key, as indicated by the key symbol) and "D55". Metadata discovery engine 712 also determines, based on semantic analysis of dataset 730 and its data, that dataset 730 "Cust_DOB" represents information about "Customer DOB", field "I95" represents "Customer ID" as part of group "DB Identifier", and field "D55" represents customer "DOB" as part of group "Customer Identity". Metadata discovery engine 712 also determines that dataset 730 is related to dataset 330a (e.g., via a primary-foreign key relationship).

[0089] The metadata discovery engine 712 passes the discovered metadata 732 to the metadata-based processing engine 714. The metadata-based processing engine 714 uses the metadata 732 to update the metadata model 716. Specifically, the metadata-based processing engine 714 can add nodes representing items of physical metadata and logical metadata 732 of the data set 730, and add edges representing the relationships between the nodes. These additions are Fig. 7A The metadata model 716 is shown in bold.

[0090] refer to Figure 7B , the metadata-based processing engine 714 uses the metadata model 716 to update the processing specification 724. Specifically, the metadata-based processing engine 714 uses the metadata model 716 to update the processing specification 724a, which can specify the generation of a wide data record including all connected data sets related to "Cust_data". In this example, the metadata-based processing engine 714 uses the metadata in the metadata model 716 (e.g., metadata describing access parameters of the data set) to update the processing specification 724a to include an access operation 725a configured to read the new data set 730. The metadata-based processing engine 714 also uses the metadata in the metadata model 716 (e.g., metadata describing primary-foreign key relationships) to include a connection operation 725b configured to connect the data set 730 to the data set "Cust_data". The metadata-based processing engine 714 can also use the metadata model 716 to apply the metadata-driven transformation 722 to the processing specification 724a that now includes the new data set 730. By leveraging metadata in this manner, system 700 can automatically adjust the processing of data to account for changes in the underlying data (including the addition of new data, such as data set 730) without requiring a user to redefine or recode the underlying specification.

[0091] Figure 8 A flowchart of an example process 800 for automating the identification of operations and their application to data using metadata is illustrated. The process 800 can be implemented by one or more of the systems and components described herein (e.g., a metadata management system or components thereof, such as a metadata discovery engine, a metadata-based processing engine, etc.), including a system configured to implement a reference Figure 1 One or more computing systems that implement the techniques described in FIG. 7 .

[0092] The operations of process 800 include identifying 802 a given data item and physical metadata for the given data item. In some examples, identifying the given data item includes identifying the given data item accessed according to a processing specification. After identifying the given data item and the physical metadata for the data item, a metadata model is accessed 804. Generally speaking, a metadata model may include a parent node and a child node connected by an edge, wherein the parent node specifies logical metadata and the child node specifies physical metadata representing the data item, and wherein the edge specifies a relationship between the nodes. In some examples, the metadata model includes a first data structure corresponding to the parent node, the data structure including logical metadata and at least a first pointer and a second pointer, wherein the first pointer points to a second data structure corresponding to a child node representing physical metadata for the given data item, and the second pointer points to a third data structure corresponding to another child node representing physical metadata for another data item different from the given data item.

[0093] A child node representing physical metadata for the given data item is identified 806 in the metadata model. In some examples, identifying the child node representing physical metadata for the given data item includes matching the physical metadata for the given data item with the physical metadata represented by the child node in the metadata model.

[0094] One or more edges in the metadata model are traversed 808 to identify one or more parent nodes of the child node. Based on the logical metadata associated with the identified one or more parent nodes, one or more operations to be performed on the given data item are determined 810. In some examples, the metadata associated with the one or more parent nodes (e.g., logical metadata) includes metadata received from a user through interaction with a metadata management system. For example, a user may access a graphical user interface including a visualization of a metadata model, and may interact with the metadata model to select metadata items associated with one or more nodes (e.g., parent nodes) and specify metadata for those nodes. For example, the metadata received from the user may be a personally identifiable information (PII) specification, which is specified by the user by interacting with the visualization of the metadata model to select metadata items associated with one or more parent nodes in the parent node for which the PII specification will be specified in the model. In some examples, after receiving the metadata specification (e.g., PII specification), the metadata model may be updated to associate logical metadata with the one or more parent nodes, such as indicating in the model that one or more parent nodes in the parent node associated with the selected metadata item are associated with the PII specification. In some examples, the one or more operations applied to the given data item to transform the given data item include tokenizing one or more fields of the given data item based on, for example, a PII specification.

[0095] In some examples, one or more metadata transformations are accessed to determine one or more operations to be performed on a given data item, wherein each metadata transformation specifies at least one operation to be performed on the data and at least one condition for performing the at least one operation. The metadata transformations can then be used in conjunction with the metadata to determine the one or more operations to be performed on the given data item. For example, a metadata transformation can be selected from the one or more metadata transformations, and it is determined whether the logical metadata associated with the one or more parent nodes satisfies at least one condition of the selected metadata transformation. When it is determined that the logical metadata associated with the one or more parent nodes satisfies the at least one condition of the selected metadata transformation, it can be determined that the one or more operations to be performed on the given data item include at least one operation specified by the selected metadata transformation.

[0096] The one or more operations are applied 812 to a given data item to transform the data item. Generally speaking, applying one or more operations to a given data item may include transforming the data item, discarding the data item or a data field associated with the data item (e.g., to avoid further processing by a computer program), adding a data field to the data item or a data set associated with the data item, or filtering a data record associated with the data item, etc. For example, the one or more operations performed on the given data item include a tokenization operation, and applying the one or more operations to the given data item to transform the given data item includes applying the tokenization operation to the given data item to tokenize one or more fields of the given data item. In some examples, applying the one or more operations to the given data item includes inserting the one or more operations into a processing specification, and executing the processing specification to apply the one or more operations to the given data item. In some examples, the processing specification is a specification for a data flow graph, and applying the one or more operations to the given data item includes executing the data flow graph, wherein executing the data flow graph applies the one or more operations to the given data item. After the one or more operations are applied to the data item, the data item may be stored 814 (e.g., in a memory or another hardware storage device), displayed to a user, or both.

[0097] In some examples, the metadata model is updated based on the transformed data item. For example, one or more nodes representing the metadata of the transformed data item may be added to the metadata model, wherein the added one or more nodes represent the metadata of the transformed data item. Some of the added nodes may be copies of some of the nodes that already existed in the metadata model before the update. As another example, one or more edges may be added to the metadata model to propagate existing metadata to the transformed data item. For example, one or more edges may be added to an existing node in the metadata model to propagate or connect an existing node in the metadata model to the added node of the transformed data item. The updated metadata model may be output for further processing of the data item.

[0098] In some examples, operation 800 may also include identifying physical metadata of the transformed data item; identifying, in the updated metadata model, a child node representing the physical metadata of the transformed data item; traversing one or more edges in the updated metadata model to identify one or more parent nodes of the child node representing the physical metadata of the transformed data item; and determining from logical metadata associated with the one or more parent nodes identified by traversing the one or more edges in the updated metadata model that the one or more operations performed on the given data item will not be performed on the transformed data item. Determining that the one or more operations performed on the given data item will not be performed on the transformed data item may include determining that the logical metadata associated with the one or more parent nodes identified by traversing the one or more edges in the updated metadata model is different from the logical metadata associated with the one or more parent nodes identified by traversing the one or more edges in the metadata model.

[0099] In some examples, operation 800 may also include identifying physical metadata of the transformed data item; identifying a child node representing the physical metadata of the transformed data item in the updated metadata model; traversing one or more edges in the updated metadata model to identify one or more parent nodes of the child node representing the physical metadata of the transformed data item; determining one or more second operations to be performed on the transformed data item from logical metadata associated with the one or more parent nodes identified by traversing the one or more edges in the updated metadata model; applying the one or more second operations to the transformed data item to further transform the transformed data item; and storing the further transformed data item.

[0100] In some examples, operation 800 may also include accessing a plurality of data items stored in a hardware storage device; for each of the plurality of data items, identifying physical metadata and logical metadata corresponding to the data item; generating a metadata model based on the physical metadata and logical metadata identified for each of the plurality of data items; and providing access to the metadata model for the first application and the second application. The first application and / or the second application may be a data flow graph or other computer program executed on, for example, a metadata management system or a client device and other data processing systems. At least one of the first application or the second application may access the metadata model, and each of the first application and the second application may be configured to: access a given data item to identify physical metadata for the given data item; identify a child node in the metadata model that presents the identified physical metadata of the given data item; traverse one or more edges in the metadata model to identify the one or more parent nodes of the identified child node; determine at least one operation to be performed on the given data item from the logical metadata associated with the one or more identified parent nodes; apply the at least one operation to the given data item to transform the given data item; and store the transformed data item.

[0101] The specific implementation of the subject matter and operations described in this specification, including the data ingestion system and its components, can be implemented in digital electronic circuits, or in computer software, firmware or hardware, including the structures disclosed in this specification and their structural equivalents, or in a combination of one or more thereof. The specific implementation of the subject matter described in this specification can be implemented as one or more computer programs (also called data processing programs) (i.e., one or more modules of computer program instructions, which are encoded on a computer storage medium for execution by a data processing device or control of the operation of the data processing device). The computer storage medium can be or can be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. The computer storage medium can also be or be included in one or more separate physical components or media (e.g., multiple CDs, disks or other storage devices). The subject matter can be implemented on computer program instructions stored on a non-transitory computer storage medium.

[0102] The operations described in this specification may be implemented as operations performed by a data processing system or device on data stored on one or more computer-readable storage devices or received from other sources. The term "data processing system" covers all kinds of devices, equipment and machines for processing data, including, for example, programmable processors, computers, systems on a chip, or multiple items of the foregoing items or a combination of the foregoing items. The system may include a dedicated logic circuit (e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit)). In addition to hardware, the system may also include code that provides an execution environment for related computer programs (e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them). The system and execution environment can implement various different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0103] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on one computer or on multiple computers, which are located at one site or distributed across multiple sites and interconnected by a communications network.

[0104] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. These processes and logic flows can also be performed by special-purpose logic circuits (such as FPGAs (field programmable gate arrays) or ASICs (application-specific integrated circuits)), and the apparatus can also be implemented as special-purpose logic circuits.

[0105] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more large-capacity storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operably coupled to receive data from it or transmit data to it or both, however, the computer does not need to have such devices. In addition, the computer can be embedded in another device (e.g., a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive)). Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0106] Implementations of the subject matter described in this specification may be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., autonomous peer-to-peer networks).

[0107] A computing system may include a user and a server. The user and the server are usually remote from each other and usually interact through a communication network. The relationship between the user and the server is due to the computer program running on the corresponding computer, and they have a user-server relationship with each other. In some specific implementations, the server transmits data (e.g., an HTML page) to the user device (e.g., for the purpose of displaying data to a user interacting with the user device and receiving user input from the user). Data generated at the user device (e.g., the result of user interaction) can be received from the user device at the server.

[0108] Although this specification contains many specific implementation details, these should not be interpreted as limitations on the scope of any specific implementation or content that may be claimed, but rather as descriptions of features specific to a particular specific implementation. Certain features described in this specification in the context of separate specific implementations may also be implemented in combination in a single specific implementation. Conversely, the individual features described in the context of a single embodiment may also be implemented in multiple specific implementations individually or in any suitable sub-combination. In addition, although the features described above may function in certain combinations, and may even be initially claimed as such, in some cases, one or more features in the claimed combination may be deleted from this combination, and the claimed combination may be a variation of a sub-combination or sub-combination.

[0109] Similarly, although the drawings show each operation in a specific order, this should not be understood as such operations must be performed in the specific order shown or in a sequential order, or that all of the operations shown must be performed to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system components in the above-mentioned specific implementations should not be understood as requiring such separation in all specific implementations, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0110] Other implementations are within the scope of the following claims.

Claims

1. A method implemented by a metadata management system for using a metadata model to identify which one or more operations to perform when processing one or more data items, wherein the metadata model includes one or more parent nodes and one or more child nodes, wherein the one or more operations are defined relative to the one or more parent nodes in the metadata model and are applied to data represented by the one or more child nodes in the metadata model, wherein the one or more parent nodes specify logical metadata, and wherein the one or more child nodes specify physical metadata representing the one or more data items, the method comprising: identifying a given data item and physical metadata of the given data item; accessing the metadata model, wherein edges connect the nodes, wherein an edge specifies a relationship between two nodes; identifying, in the metadata model, a child node representing the physical metadata of the given data item; traversing one or more edges in the metadata model to identify one or more parent nodes of the child node; determining, from logical metadata associated with the identified one or more parent nodes, one or more operations to be performed on the given data item; applying the one or more operations to the given data item to transform the given data item; as well as The transformed data items are stored in a memory.

2. The method of claim 1, wherein identifying the child node representing the physical metadata of the given data item comprises matching the physical metadata of the given data item with the physical metadata represented by the child node in the metadata model.

3. The method of claim 1 , further comprising accessing one or more metadata transformations to determine the one or more operations to be performed on the given data item, each metadata transformation specifying at least one operation to be performed on the data and at least one condition for performing the at least one operation.

4. The method of claim 3, wherein determining the one or more operations to be performed on the given data item comprises: selecting a metadata transform from the one or more metadata transforms; determining whether the logical metadata associated with the one or more parent nodes satisfies the at least one condition of the selected metadata transformation; as well as Responsive to determining that the logical metadata associated with the one or more parent nodes satisfies the at least one condition of the selected metadata transformation, determining that the one or more operations to be performed on the given data item include the at least one operation specified by the selected metadata transformation.

5. The method of claim 1 , wherein identifying the given data item comprises identifying the given data item to be accessed according to a processing specification, and wherein applying the one or more operations to the given data item comprises: inserting the one or more operations into the processing specification; as well as The processing specification is executed to apply the one or more operations to the given data item.

6. The method of claim 5, wherein the processing specification comprises a specification for a data flow graph, and wherein applying the one or more operations to the given data item comprises executing the data flow graph, wherein executing the data flow graph applies the one or more operations to the given data item.

7. The method of claim 1, further comprising updating the metadata model based on the transformed data items.

8. The method of claim 7, wherein updating the metadata model comprises adding one or more edges to an existing node in the metadata model to connect the existing node in the metadata model to the added node.

9. The method according to claim 7, further comprising: physical metadata identifying the transformed data item; identifying, in the updated metadata model, a child node representing the physical metadata of the transformed data item; traversing one or more edges in the updated metadata model to identify one or more parent nodes of the child nodes representing the physical metadata of the transformed data item; as well as It is determined from logical metadata associated with the one or more parent nodes identified by traversing the one or more edges in the updated metadata model that the one or more operations performed on the given data item are not to be performed on the transformed data item.

10. A method according to claim 9, wherein determining that the one or more operations performed on the given data item will not be performed on the transformed data item includes determining that the logical metadata associated with the one or more parent nodes identified by traversing the one or more edges in the updated metadata model is different from the logical metadata associated with the one or more parent nodes identified by traversing the one or more edges in the metadata model.

11. The method of claim 7, wherein updating the metadata model comprises adding one or more nodes to the metadata model, the added one or more nodes representing metadata of the transformed data item.

12. The method of claim 11, wherein some of the added nodes are copies of some of the nodes that already existed in the metadata model before the update.

13. The method of claim 12, further comprising outputting the updated metadata model for further processing of the data item.

14. The method according to claim 12, further comprising: physical metadata identifying the transformed data item; identifying, in the updated metadata model, a child node representing the physical metadata of the transformed data item; traversing one or more edges in the updated metadata model to identify one or more parent nodes of the child nodes representing the physical metadata of the transformed data item; determining, from logical metadata associated with the one or more parent nodes identified by traversing the one or more edges in the updated metadata model, one or more second operations to be performed on the transformed data item; applying, by the metadata management system, the one or more second operations to the transformed data item to further transform the transformed data item; and The further transformed data items are stored in memory.

15. The method of claim 1, wherein applying the one or more operations to the given data item comprises discarding data fields associated with the data item from further processing by a computer program.

16. The method of claim 1, wherein applying the one or more operations to the given data item comprises adding a data field to the data item or a data set associated with the data item.

17. The method of claim 1, wherein applying the one or more operations to the given data item comprises filtering data records associated with the data item.

18. A method according to claim 1, wherein the one or more operations to be performed on the given data item include a tokenization operation, and wherein applying the one or more operations to the given data item to transform the given data item includes applying the tokenization operation to the given data item to tokenize one or more fields of the given data item.

19. The method of claim 1, wherein the logical metadata associated with the one or more parent nodes comprises metadata received from a user through interaction with the metadata management system.

20. The method of claim 19, wherein the metadata received from the user is a personally identifiable information specification (PII specification) specified by the user via a graphical user interface, the graphical user interface including a visualization of the metadata model, and wherein the PII classification is specified by the user interacting with the visualization of the metadata model to select metadata items associated with one or more of the parent nodes for which the PII specification is to be specified in the model, and the method further comprising: updating the model after receiving the PII specification to associate the logical metadata with the one or more parent nodes, including indicating in the model that the one or more parent nodes of the parent nodes associated with the selected metadata item are associated with the PII specification, wherein the one or more operations applied to the given data item to transform the given data item include tokenizing one or more fields of the given data item.

21. The method according to claim 1, further comprising: accessing a plurality of data items stored in a hardware storage device; For each data item of the plurality of data items, identifying physical metadata and logical metadata corresponding to the data item; generating the metadata model based on the physical metadata and the logical metadata identified for each data item of the plurality of data items; providing access to the metadata model for the first application and the second application; The metadata model is accessed by at least one of the first application and the second application, wherein each of the first application and the second application is configured to: accessing the given data item to identify the physical metadata of the given data item; identifying, in the metadata model, the child node presenting the identified physical metadata for the given data item; traversing one or more edges in the metadata model to identify the one or more parent nodes of the identified child node; determining, from the logical metadata associated with the identified one or more parent nodes, at least one operation to be performed on the given data item; applying the at least one operation to the given data item to transform the given data item; as well as The transformed data items are stored.

22. A method according to claim 1, wherein the metadata model includes a first data structure corresponding to the parent node, the data structure including the logical metadata and at least a first pointer and a second pointer, wherein the first pointer points to a second data structure corresponding to the child node representing the physical metadata of the given data item, and the second pointer points to a third data structure corresponding to another child node representing the physical metadata of another data item different from the given data item.

23. A system, comprising: at least one processor; and a memory storing instructions executable by the at least one processor to perform operations including: identifying a given data item and physical metadata of the given data item; accessing a metadata model, the metadata model comprising one or more parent nodes specifying logical metadata and one or more child nodes specifying physical metadata, wherein edges connect the nodes, wherein the edge specifies a relationship between two nodes; identifying, in the metadata model, a child node representing the physical metadata of the given data item; traversing one or more edges in the metadata model to identify one or more parent nodes of the child node; determining, from logical metadata associated with the identified one or more parent nodes, one or more operations to be performed on the given data item; applying the one or more operations to the given data item to transform the given data item; as well as The transformed data items are stored in a memory.

24. A non-transitory computer readable medium storing instructions executable by at least one processor to perform operations comprising: identifying a given data item and physical metadata of the given data item; accessing a metadata model, the metadata model comprising one or more parent nodes specifying logical metadata and one or more child nodes specifying physical metadata, wherein edges connect the nodes, wherein the edge specifies a relationship between two nodes; identifying, in the metadata model, a child node representing the physical metadata of the given data item; traversing one or more edges in the metadata model to identify one or more parent nodes of the child node; determining, from logical metadata associated with the identified one or more parent nodes, one or more operations to be performed on the given data item; applying the one or more operations to the given data item to transform the given data item; as well as The transformed data items are stored in a memory.

Citation Information

Patent Citations

  • Managing record format information

    US20110153667A1

  • Discovering a semantic meaning of data fields from profile data of the data fields

    US20200380212A1

  • Executing computations expressed as graphs

    US5966072A

  • Managing data set objects

    US9977659B2

  • Analysis method and system of data communication protocol

    CN108183890A