Automatic Dependency Analyzer for Heterogeneously Programmed Data Processing Systems

JP7708828B2Active Publication Date: 2025-07-15AB INITIO TECHNOLOGY LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023170105
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-05-22
Filing Date
2023-09-29
Publication Date
2025-07-15
Estimated Expiration
2038-05-22

AI Technical Summary

Technical Problem

Existing data processing systems struggle with incomplete or inefficient dependency analysis across multiple programming languages, leading to inaccuracies and inefficiencies in managing data dependencies.

Method used

A dependency analyzer that processes programs written in multiple programming languages, constructing language-independent data structures to identify field-level lineage and dependencies, using a front-end module for language-specific parsing and a back-end module for unified dependency analysis.

Benefits of technology

Enhances the accuracy and efficiency of dependency analysis, providing comprehensive dependency information across heterogeneous programming environments, improving data processing system reliability and functionality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708828000025
    Figure 0007708828000025
  • Figure 0007708828000026
    Figure 0007708828000026
  • Figure 0007708828000027
    Figure 0007708828000027
Patent Text Reader

Abstract

To provide a method for calculating dependency information related to field-level lineage across programs in a plurality of computer programming languages.SOLUTION: In a computing environment 100, a dependency analyzer includes at least one computing device 103, 134 that generates dependency information between variables appearing in any of a plurality of programs written in different source languages for a data processing system 105. A data processing system analyzes each program, regardless of a language where a module is written, and records analysis information about each program in a first-type data structure to convert into a format that represents the dependency between the variables.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Often, large amounts of data are managed by a data processing system. The data processing system may be implemented on one or more computers coupled to one or more data sources. The computer can be configured to perform certain data processing operations by using software programs. Those software programs may be heterogeneous and written in multiple programming languages.

[0002] As an example, at a university, a data processing system that receives data as input from students requesting classes may process student registration. The data may include a student's name or identification number, a course title or number, an instructor, and other information related to the student or the course. The processing system can be programmed to process each request by accessing various data sources such as a data store containing information about each registered student. The data processing system can access this data store to determine whether a registration request for a class was made by a student who has paid tuition. Another data store may contain information about classes, including required classes. The data processing system can access this data in combination with data from another data store indicating classes that the student requesting the class has already completed, and can determine whether the student is eligible to register for the requested class.

[0003] By processing the data accessed according to the program, the data processing system can generate data indicating that the student is registered for the requested class and that the student is included in the class roster. When the data processing system completes processing a class registration request, the data processing system can access many data sources and can input or modify data in multiple data stores.

[0004] Similar patterns occur in many other important applications where a data processing system accesses multiple data sources and generates or modifies values stored in other data stores based on data accessed from the data sources or based on the values of other variables. Thus, many values of a variable can depend on the values of other variables, whether other values of the variable are accessed from a data source or exist within the data processing system.

[0005] In many cases, it is beneficial for a data processing system to provide dependency information about data elements that it generates or modifies. Ab Initio of Lexington, Massachusetts, USA, provides a "coordination system" that provides dependency information based on programs created to execute on a coordination system. The coordination system executes programs represented as data flow graphs, which are represented as operators and data flow between operators. Tools can analyze the graph and determine dependencies between variables by identifying operators where the value of a variable is set based on the values of one or more other variables. By tracing the flow back to an operator, the tool can identify other variables on which the variable input to the operator next depends. By tracing through the graph, the tool can identify all dependencies on any variable, directly or indirectly.

[0006] This dependency information can be input into a metadata store and used in many graph-related functions therefrom. For example, data quality rules may be applied to identify variables with patterns of unexpected values. These variables, and any variables in the graph that are dependent on them, may be flagged as suspect. As another example, when it is necessary to modify a part of the graph to provide a new function, correct an error, or exclude a function that is no longer desired, the variables generated or modified in that process can easily identify other parts of the dependent graph. Following the change, those parts of the graph can be inspected to ensure that the function is not impaired by changes made in other parts of the graph.

[0007] In some cases, the data processing system can use other programming languages instead of, or in addition to, the graph. For example, the data source can be a database programmed in SQL or a programming language defined by the database supplier, and the user can program the queries executed against that database. In a data processing system that can be implemented in large enterprises, there may be multiple programs written in multiple languages instead of, or in addition to, a graphical programming language. Dependency analysis tools operating on a program represented as a graph do not process these programs written in other languages. As a result, the dependency information for the entire data processing system may be incomplete or may require a great deal of effort to generate. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0008] Some embodiments provide a dependency analyzer for use in a data processing system configured to execute a program in any of a plurality of computer programming languages that access a plurality of records including fields. The dependency analyzer may be configured for dependency analysis of programs in a plurality of computer programming languages, and the dependency analysis may include calculation of dependency information regarding field-level lineage across computer programs in the plurality of computer languages. The dependency analyzer may comprise a front end configured to process a plurality of programs that control the data processing system. The plurality of programs may be written in a plurality of programming languages of the plurality of programming languages. The front end may comprise a plurality of front-end modules, each front-end module being configured to receive, as input, a computer program of a plurality of programs written in a programming language of the plurality of programming languages, and to output a language-independent data structure representative of dependencies to fields of the data processing system created within the program, the language-independent data structure being composed of one or more dependency components, each dependency component including an indication of at least one field and a dependency associated with at least one field. The data analyzer may also comprise a back end configured to receive the language-independent data structures from the plurality of front-end modules and to output dependency information of the data processing system representative of dependencies created within the plurality of programs. The dependency information may include field-level dependencies introduced into a plurality of the plurality of programs such that the dependency information includes field-level lineage across the plurality of programs.

[0009] Some embodiments provide a dependency analyzer for use in a data processing system comprising at least one computer hardware processor and at least one non-transitory computer-readable storage medium storing processor-executable instructions that cause the at least one computer hardware processor to perform an act that includes constructing, for each of a plurality of programs configured to be executed by the data processing system when executed by the at least one computer hardware processor, a first data structure that reflects operations within the program based on an analysis of the program. For each of the plurality of first data structures, operations reflected in the first data structure that affect dependencies on one or more data elements may be selected such that operations that do not affect dependencies are not selected. For each selected operation, an indication of the selected operation may be recorded in a second data structure by recording dependency components from a set of dependency components. The set of dependency components for each of the plurality of programs may be independent of the language in which the program is written. A plurality of second data structures may be processed. The processing may include identifying control flow or data flow dependencies on data elements, which dependencies may occur during execution of any of the plurality of programs. The identified dependencies within the dependency data structure may be processed, and the dependency data structure may include data indicating data elements within the plurality of programs configured to indicate dependencies between data elements.

[0010] Some embodiments provide a method of operating a dependency analyzer configured for use in a data processing system, using at least one computer hardware processor that executes within the data processing system and executes a plurality of programs configured to specify data processing operations in a programming language, where the plurality of programs use a plurality of programming languages, and constructing a respective first data structure of the plurality of programs by analyzing the programs. For each of the plurality of first data structures, the method can include selecting an operation reflected in the first data structure that affects a dependency on one or more data elements, and for each selected operation, recording an instruction of the selected operation in a second data structure, where the instruction is recorded in the second data structure using a set of dependency components common to all second data structures. The second data structure can be processed, and the processing can include identifying control flow or data flow dependencies on data elements that can occur during the execution of any of the plurality of programs. The identified dependencies can be recorded in a dependency data structure.

[0011] Some embodiments can provide an operation method of a dependency analyzer for use with a data processing system. The method includes using at least one computer hardware processor that executes a plurality of programs specifying data processing operations in a programming language within the data processing system, where the plurality of programs use a plurality of programming languages, and constructing a first data structure for each of the plurality of programs by analyzing the programs. For each of the plurality of first data structures, the method may include selecting an operation reflected in the first data structure that affects the control flow dependency to one or more data elements. For each selected operation, the instruction of the selected operation may be recorded in a second data structure, and the instruction is recorded in the second data structure using a set of dependency components common to all second data structures. The second data structure can be processed, and the processing includes identifying the dependency to the data element as a result of the processing in any of the plurality of programs. The identified dependency within the dependency data structure.

[0012] Yet another embodiment can provide at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform acts on a plurality of programs that specify data processing operations in respective programming languages, where the plurality of programs use a plurality of programming languages. The acts include constructing, by parsing the programs, respective first data structures of the plurality of programs. For each of the plurality of constructed first data structures, operations reflected in the first data structure that affect control flow or data flow dependencies to one or more data elements can be extracted from the first data structure, and the extracted operations can be recorded in a second data structure, where the instructions for the extracted operations are recorded in the second data structure using a set of dependency components common to all second data structures. The second data structure can be processed, and the processing includes identifying dependencies to data elements as a result of processing in any of the plurality of programs. The identified dependencies can be recorded in a dependency data structure.

[0013] Still other embodiments can provide a dependency analyzer for use in a data processing system, the dependency analyzer comprising means for constructing a first data structure representing each of a plurality of programs, the first data structure reflecting operations in each of the plurality of programs; means for recording in a second data structure an indication of an operation in a first data structure that affects control flow or data flow dependencies such that information regarding operations in each of the plurality of constructed first data structures that do not affect dependencies is omitted from the second data structure; means for processing the second data structure to identify control flow or data flow dependencies to data elements as a result of execution of operations in any of the plurality of programs; and means for recording the identified dependencies in a dependency data structure.

[0014] The foregoing is a non-limiting summary of the invention defined by the appended claims.

[0015] The following drawings are referred to in explaining various aspects and embodiments. It is to be understood that these drawings are not necessarily drawn to scale. Items that appear in multiple drawings are denoted by the same or similar reference numerals in all of the drawings in which they appear.

Brief Description of the Drawings

[0016]

Fig. 1A

Fig. 1B

Fig. 2

Fig. 3A

Fig. 3B

Fig. 3C

Fig. 3D

Fig. 4

Fig. 5

Fig. 6

Modes for Carrying Out the Invention

[0017] The inventors have recognized and understood that the accuracy, efficiency, and reliability of a data processing system can be improved by automatically converting source code representing a program for the data processing system into a combination of components independent of the language representing data dependencies. The components generated for the program can be combined with a part of a system including programs prepared in different languages and a data structure that potentially captures the dependencies of the entire system, whether represented in a graph or other form. Then, the data structure can be processed to generate dependency information of the system, including programs coded in multiple languages such as graphs and one or more other languages. Such an approach leads to an improved dependency analysis tool that generates more complete dependency information for the data managed by the data processing system.

[0018] In some embodiments, a set of components may be used to represent dependencies. The inventors have recognized and understood that useful information regarding dependencies to data elements within a data processing system can be represented by a finite set of dependency components. This set of structural elements can include, for example, components representing assignments, procedure definitions, procedure returns, procedure call sites, conditional control flows in the "if...else" format, and conditional control flows in the "while...do" format. In some embodiments, the set of dependency components consists of, consists essentially of, or includes only variations of these components, such as the assignment of constant values.

[0019] The inventors further recognized and understood that the automatic processing of source code representing a part of a program can be efficiently executed in multiple stages. In the first-stage processing, the source code can be parsed. The parsed source code can be represented in a first data structure that can be manipulated by a computer, such as a parse tree. In some embodiments, the first-stage processing can be executed using commercially available tools such as a parser generator. Examples of commercially available parser generators are Yacc and ANTLR. Both of these generate a parser from a grammar file that specifies the grammar of a programming language.

[0020] A parser is language-dependent and is often generated by a parser generator from the grammar rules of a language. Thus, a data processing system configured to identify dependencies of programs written in multiple languages can include multiple parsers and / or a parser generator that generates multiple parsers. Thus, the data processing system can access multiple parsers and select and, if necessary, generate a parser for each program based on the programming language in which the program is written.

[0021] In the second-stage processing, information regarding dependencies can be extracted from the parse tree or other data structure generated in the first stage. The dependency information can be expressed as a combination of components. The components may be selected from the above-described set or from other suitable sets of components. The result of the processing can be represented in a second computer data structure, which can also be formatted as a tree. However, it should be understood that the result of the second-stage processing can be stored in any data structure that holds information regarding the operations executed by a program that can affect dependencies based on data elements within the program or data elements accessed by the program.

[0022] The processing of the second stage is also language - dependent. Different languages can support different operations and can affect dependencies in different ways. Thus, the tools applied to each of the first data structures can be selected for the programming language of the program being processed. For example, a data processing system has multiple tools for the second - stage processing, and an appropriate tool can be selected for each program. The selection can be made dynamically by the data processing system to match the tool to the language in which the program being processed is written.

[0023] Each of the tools for the second - stage processing may be implemented in any suitable way. In some embodiments, each tool may be programmed to recognize operations in the first data structure that affect dependencies based on data elements. These dependencies often become one program variable, sometimes called a target variable, and one or more other program variables, sometimes called source variables. This type of data - flow dependency is used here as an example. However, other types of dependencies are possible, and it should be understood that they can be analyzed using the techniques described herein. For example, control - flow dependencies are also possible, where the execution of a code block depends on the values of one or more data elements.

[0024] In an example where the first data structure is an analysis tree, the tool can traverse the analysis tree to generate a second data structure. When detecting an operation that affects a dependency, the tool can create an entry in the second data structure using one or more components from a set of components. In this example, hierarchical information reflecting the order of operations can be carried over to the second data structure. The tool may not create an entry in the second data structure for operations reflected in the first data structure that do not affect dependencies. Thus, the second data structure is more compact than the first data structure.

[0025] Once the second data structure has been calculated for each program, the second data structure for the programs used together in the data processing system can be used together. The second data structure can identify which variables are dependent on other variables. This information can be represented as a transformation that defines each operation that introduces a dependency, the target variable affected by the operation, and the source variable from which the data used in that transformation is obtained.

[0026] The second data structure representing the dependencies of the system can be processed by a dependency analysis tool. Each target recorded in the second data structure can be identified and the source variables for that target can be listed. According to some embodiments, a program may have nested scopes such that functions, or other components that define internal blocks of code, may occur when defined within other components that define internal blocks of the program or external blocks of code. In some embodiments, the dependencies may only exist within the inner scope so as not to affect variables that are output or modified by the data processing system. In some embodiments, such purely local dependencies can be removed, leaving only the external dependency information.

[0027] In some embodiments, a first external variable may depend on a local variable, and the local variable may depend on one or more second external variables. In this scenario, an indirect dependency of the first external variable on the one or more second external variables may be identified in a second stage of processing. Such identification can be done automatically by propagating the local dependencies to the dependency analysis tool and recasting the dependency of the first external variable on the internal variable as a dependency on one or more second external variables on which the local variable depends.

[0028] Regardless of the specific content, this information regarding data dependencies can be written into a metadata structure such as a metadata repository used in a data processing system. The information can be written in combination with other appropriate information. In some embodiments, for example, information regarding the source code of a program may be captured together with information characterizing the dependencies of that program. As a specific example, a portion of the code including the statement where the dependency is created can be saved.

[0029] As an example of other types of information that can optionally be saved in some embodiments, information regarding control dependencies can be saved. Control dependencies can be identified by identifying blocks of a program that are conditionally executed depending on the value of one or more variables. This information can be captured similarly to data dependency information, but rather than recording variables as target entities that directly depend on other source variables, code blocks can be identified as depending on one or more flow control source variables. Any variable having a value that may depend on whether the code block is executed can be identified as a target variable and associated with those flow control source variables.

[0030] Dependency information can be used in any of a number of ways. For example, during the operation of a data processing system, dependency information can be used when evaluating rules that may enable or block the processing operation of specific data. As an example, the operation of receiving as input data that depends on a variable identified as broken or untrustworthy may be suppressed.

[0031] It is understood that the embodiments described herein may be implemented in any of a number of ways. For illustrative purposes only, specific examples of implementation are provided below. It is understood that these embodiments and the features / functions provided may be used individually, all together, or in any combination of two or more (the aspects of the technology described herein are not limited in this regard).

[0032] Figure 1A is a block diagram of a computing environment 100 for illustrative purposes in which some embodiments of the techniques described herein may operate. The computing environment 100 includes a data processing system 105 configured to operate on data stored in a data storage device 104.

[0033] In some embodiments, the data storage device 104 may include one or more storage devices that store data in one or more suitable formats of any appropriate type. For example, the storage device portion of the data storage device 104 may use one or more database tables, spreadsheet files, flat text files, and / or files in any other suitable format (e.g., mainframe native format) to store data. The storage device may be of any appropriate type and may include one or more servers, one or more database systems, one or more portable storage devices, one or more non-volatile storage devices, one or more volatile storage devices, and / or any other device configured to store data electronically. In some embodiments, the data storage device 104 may include one or more online data streams in addition to, or instead of, the storage device. Thus, in some embodiments, the data processing system 105 may have access to data provided on one or more data streams in any suitable format.

[0034] In embodiments where the data storage device 104 includes multiple storage devices, the storage devices may be located in the same location (e.g., within a building) in one physical location or may be distributed across multiple physical locations (e.g., within multiple buildings, in different cities, states, or countries). The storage devices may be configured to communicate with each other using one or more networks, such as the network 106 shown in FIG. 1A.

[0035] In some embodiments, the data stored by the storage device may include one or more data entities such as one or more files, tables, data in rows and / or columns of a table, spreadsheets, datasets, data records (e.g., credit card transaction records, call records, and bank transaction records), fields, variables, messages, and / or reports. The storage device may store thousands, millions, tens of millions, or hundreds of millions of data entities. Each data entity may include one or more data elements.

[0036] As described herein, a data element may be any data element that is stored and / or processed by a data processing system. For example, a data element may be a field in a data record, and the value of the data element may be the value stored in the field of the data record. As a specific non-limiting example, a data element may be a field that stores the name of the caller in a data record that stores information about a call (this data record may be part of a plurality of data records regarding calls made by a customer of a telecommunications company), and the value of the data element may be the value stored in the field. As another example, a data element may be a cell in a table (e.g., a cell existing in a specific row and column of a table), and the value of the data element may be the value of the cell in the table. As another example, a data element may be a variable (e.g., in a report), and the value of the physical element may be the value of the variable (e.g., in a specific instance of a report). As a specific non-limiting example, a data element may be a variable in a report regarding an applicant that represents the credit score of an applicant for a bank loan, and the value of the data element may be a numerical value of the credit score (e.g., a numerical value between 300 and 850). The value of the data element representing the applicant's credit score may vary depending on the data used to generate the report regarding the applicant for the bank loan.

[0037] In some embodiments, a data element may take on any suitable type of value. For example, a data element may take on a numerical value, an alphabetic value, a value from a discrete set of choices (e.g., a finite set of categories), or any other suitable type of value, and the aspects of the techniques described herein are not limited in this regard. These values can be assigned by processing on a computer accessing the data store 104.

[0038] The data processing system 105 may be programmed with a “coordination system” or other suitable program that enables the data processing system to execute a program, and may be described to configure the data processing system 105 for a particular enterprise based on the type of data processing the enterprise requires. Thus, the data processing system 105 may include one or more computer programs 109 configured to operate on the data of the data storage device 104. The computer programs 109 may be of any suitable type and may be written in any suitable programming language. For example, in some embodiments, the computer programs 109 are written at least in part using Structured Query Language (SQL) and may include one or more computer programs configured to access data in one or more database portions of the data storage device 104. As another example, in some embodiments, the data processing system 105 is configured to execute a program in the form of a graph, and the computer programs 109 may include one or more computer programs developed as a data flow graph. A data flow graph may include components called “nodes” or “vertices” that represent data processing operations performed on input data, and links between the components that represent the flow of data. Techniques for executing the computations encoded by a data flow graph are described in U.S. Patent No. 5,966,072, entitled “Execution of Computations Represented as Graphs,” which is hereby incorporated by reference in its entirety.

[0039] Furthermore, data processing system 105 may include tools and / or utilities that perform the dependency analysis functions described herein. In the illustrated embodiment of FIG. 1A, data processing system 105 further includes a development environment 108 that can be used by a person (e.g., a developer) to develop one or more computer programs 109 that operate on the data of data storage device 104. For example, in some embodiments, user 102 can use computing device 103 to specify a computer program, such as a data flow graph, and interact with the development environment to save the computer program as part of computer programs 109. An environment for developing a computer program as a data flow graph is described in U.S. Patent Application Publication No. 2007 / 0011668, entitled "Parameter Management for Graph-Based Applications," which is hereby incorporated by reference in its entirety. The dependency analysis tool can be accessed by a developer, such as user 102, in the development environment or in any other manner.

[0040] In some embodiments, one or more of computer programs 109 may be configured to perform any suitable operations on the data of data storage device 104. For example, one or more of computer programs 109 may be configured to access data from one or more sources, transform the accessed data (e.g., by changing a data value, filtering a data record, changing a data format, classifying data, combining data from multiple sources, splitting data into multiple parts, and / or in any other suitable manner), calculate one or more new values from the accessed data, and / or write the data to one or more destinations.

[0041] In some embodiments, one or more of the computer programs 109 may be configured to perform calculations on the data of the data storage device 109 and / or generate reports from the data of the data storage device 109. The calculations performed and / or the reports generated may be related to one or more quantities related to the business. For example, the computer program may be configured to access credit history data about a person and determine that person's credit score based on the credit history. As another example, the computer program may access the call logs of multiple customers of a telephone company and generate a report indicating which of those customers are using more data than is allowed by their data plan. As yet another example, the computer program may access data indicating the types of loans made by a bank and generate a report indicating the overall risk of the loans made by the bank. These examples are illustrative and non-limiting, as the computer program can be configured to generate any suitable information (e.g., for any suitable business purpose) from the data stored in the data storage device 104. The program 109 may be called by the user 130, for example, to obtain business information from data processing operations.

[0042] In the illustrated embodiment, data processing system 105 also includes a metadata repository 110 that supports the performance of various tasks that may occur during the operation of system 105, including tasks related to maintaining program and data governance. In the illustrated embodiment, metadata repository 110 includes a data dependency module 118, a control dependency module 120, and a code summary 122. Information regarding data dependencies, control dependencies, and code summaries can be obtained as described herein. This information can be stored in any suitable format. By way of non-limiting example, data dependency module 118 may comprise one or more tables indicating, for each of a plurality of variables of computer program 109, other variables of the computer program on which a dependency exists. Control dependency module 120 can similarly store information regarding variables having values that may potentially be affected by the execution of code blocks of program 109 that are conditionally executed depending on the value of one or more variables.

[0043] Code summary module 122 can store a portion of the code in program 109 where either a data or control dependency is introduced. The code summary can identify the code by storing a copy of the associated source code, or a pointer to a portion of the source code file that includes the associated code. The code can be identified in a suitable manner, such as by a copy of the source code or by a pointer to a portion of the source code file that includes the associated code.

[0044] Figure 1B shows a process that can be executed to generate dependency information. Figure 1B shows that multiple programs, shown as programs 150, 160, and 170 in Figure 1B, can be processed to identify dependency information. These programs can collectively control a data processing system to process data from one or more data sources. The programs may be written in different programming languages. In the example of Figure 1B, program 150 is shown written in language 1, program 160 is written in language 2, and program 170 may be written in language N. However, not all languages need to be different, and in some embodiments, some or all of the multiple programs may be written in the same language. As a specific example, a portion of a program may be written as a stored procedure call that can be executed to process data within a database. Others may be written in a graphical programming language, and others may be written in still other languages.

[0045] Regardless of the combination of languages in which the programs are written, each program can be processed using language-dependent processing. In Figure 1B, the language-dependent processing is executed in front-end modules 152, 162, and 172, and each of these modules is configured to process the program in its respective language.

[0046] The output of each front-end module may be a dependency data structure that reflects the dependencies introduced by processing using a program. As shown in FIG. 1B, each front-end can output its own dependency data structure, which can be the lower-level parse trees 154, 164, and 174. For example, a lower-level parse tree can be created by parsing a program to generate a parse tree and then making the parse tree more fine-grained by excluding information not related to dependencies. According to some embodiments, the lower-level parse tree may be represented using a finite set of dependency components. Each parse tree can be made more fine-grained by representing dependency information in a language-independent parse tree using language-independent dependency components, so the lower-level parse trees 154, 164, and 174 may be language-independent.

[0047] The lower-level parse trees are then processed together to enable the generation of dependency information for all of the programs 150, 160, 170 that can be executed by the data processing system. This processing can be performed at the back-end 180. The back-end processing may be language-independent, and the back-end 180 can be used in any data processing system regardless of the language in which the programs of the system are written.

[0048] The dependency information generated by the backend 180 can be used in any desired way. In the embodiment of FIG. 1A, the dependency information is stored in a data store that functions as a metadata repository 110. The metadata repository 110 can be implemented in any suitable way. Examples of such repositories can be found, for example, in U.S. Patent Application Publication No. 2010 / 0138431A1, which is incorporated herein by reference in its entirety. Such a metadata repository can hold dependency information introduced in any of the programs that collectively implement a data processing system. The dependency information can include data indicating data elements within a plurality of programs and can be configured to indicate dependencies between the data elements. For example, the dependency information may be structured as a series of lists, each list associated with a target data element and including entries that are source data elements on which the target may depend. Such information can alternatively or additionally be stored as a table where each row and column corresponds to a variable and the cells at the intersections of the rows and columns indicate whether there is a dependency between the row variable and the column variable. In some embodiments, the information in each cell may further indicate information about the dependency, such as whether it is a data flow dependency or a control flow dependency, or whether additional information about the dependency can be accessed.

[0049] The information in metadata repository 110 can be used in any suitable way. For example, it can be used when operating or maintaining computer program 109. As a specific example, programmer 102 may modify a part of one of the computer programs 109 that affects the value of the first variable. Programmer 102 can identify other parts of computer program 109 with variables that depend on the first variable based on the information in metadata repository 110. Programmer 102 can test these parts to ensure that their operation is not interrupted by the change. As another example, user 130 of data processing system 105 who interacts with the system via computing device 134 can analyze the data in the reports generated by data processing system 105. For example, user 130 can question the validity of the calculated values presented in the report and access the information in metadata repository 110 to identify the source of the data used to calculate the questionable values. User 130 can then investigate the quality of these data sources and thereby verify the quality of the data in the report. The information in the metadata repository can be accessed and processed automatically, for example, by calling a tool.

[0050] Regardless of how the metadata in repository 110 is used, even if there are programs written in multiple different programming languages, enabling data processing system 105 to generate the metadata can increase the reliability of data processing system 105 and allow data processing system 105 to execute new functions.

[0051] These features can be enabled, in whole or in part, by accessing data lineage. Data lineage indicates where in a data processing system the data used to generate the value of a particular variable is obtained or processed. FIG. 2 is a data lineage diagram 200 of an exemplary programmed data processing system. As described herein, the data lineage and its diagram can be generated by identifying dependencies. The data lineage diagram 200 includes nodes 202 representing data entities and nodes 204 representing transformations applied to the data entities.

[0052] A data entity can be a dataset such as, for example, a file, a database, or data formatted in any suitable manner. In some embodiments, a data entity may be a "field", which may be an element accessed from a particular data store. For example, in a data store formatted as a table where cells are organized into rows and columns, a field may be a cell. As another example, in a database composed of multiple records, a field may be a field within a record. Regardless of the particular configuration of the data store accessed by the data processing system, the program implementing the data processing system can read from and / or write to those fields such that those fields are referenced in the program implementing the data processing system. Thus, those programs create dependencies associated with those fields, and some fields have values that depend on the values of other fields. In some embodiments, the dependency analyzer described herein calculates and stores dependency information associated with fields, and that information can reflect dependencies introduced in any of the programs that can be affected by or affect the value of a particular field. This field-level dependency information enables the dependency analyzer to provide information regarding the field-level lineage of the entire program in any of a plurality of computer programming languages in which the data processing system can be programmed.

[0053] Figure 2 shows each node 202 having the same symbol, but not all nodes need to be of the same type. On the contrary, in modern enterprises, the nodes 202 can be of different types, such as those resulting from enterprises using components of ORACLE, SAP, or HADOOP. Components from ORACLE or SAP may implement a database known in the art. Components of HADOOP may implement a distributed file system. A common database on the HADOOP distributed file system is known as HIVE. Such components can be used for database implementation or for implementation of datasets stored in a distributed file system. Therefore, it should be understood that the technology described herein can operate with programs in many languages that perform operations on datasets in any of a number of formats.

[0054] Furthermore, each node does not need to correspond to a single hardware component. Rather, several data sources that may be represented as a single node in FIG. 2 can be implemented with distributed database technology. HADOOP is an example of distributed database technology that can implement one data source across multiple computer hardware devices.

[0055] The transformation represented by node 204 may be implemented by a computer program. In FIG. 4, all of the nodes 204 are shown with the same symbol, but not all of the programs need to be written in the same programming language. On the contrary, the languages used may depend on the type of node 202 from which the node 204 obtains data, or may vary from node to node for any of several reasons in an enterprise environment. The technology described herein can be applied to such heterogeneous environments.

[0056] The data lineage diagram 200 illustrates upstream lineage information regarding one or more data elements in the data entity 206. An arrow entering a node representing a transformation indicates which data entity is provided as input to the transformation. An arrow leaving a node representing a transformation of data indicates to which data entity the result of the transformation is provided. Examples of transformations include, but are not limited to, performing any suitable type of calculation, classifying data, filtering data to remove one or more portions of the data based on any suitable criteria (e.g., filtering data records to remove one or more data records), merging data (e.g., using a join operation or any other suitable manner), performing any suitable database operation or command, and / or any suitable combination of the above transformations. The transformation may be implemented using one or more computer programs of any suitable type, including, by way of example and not limitation, one or more computer programs implemented as a data flow graph.

[0057] Data lineage diagrams such as the diagram 200 shown in FIG. 2 can be useful for several reasons. For example, illustrating the relationship between data entities and transformations can help a user determine how a particular data element was obtained (e.g., how a particular value in a report was calculated). As another example, a data lineage diagram can be used to determine which transformations were applied to various data elements and / or data entities.

[0058] In some embodiments, data lineage can represent the relationships between data elements, the data entities that contain these data elements, and / or the transformations applied to the data elements. The relationships between data elements, data entities, and transformations can be used to determine relationships between other things such as, for example, systems (e.g., one or more computing devices, databases, data warehouses, etc.), and / or applications (e.g., one or more computer programs that access data managed by a data processing system). For example, if a portion of a data element in a certain table in a certain database stored in system "A" located at a certain physical location is shown to be derived from a portion of a different data element in a different table in a different database stored in system "B" within the data lineage, the relationship between system A and system B can be inferred. As another example, when an application program reads one or more data elements from a certain system, the relationship between this application program and this system can be inferred. As yet another example, when an application program accesses data elements on which operations were performed by another application program, the relationship between these application programs can be inferred. Any one or more of these relationships may be shown as part of a data lineage diagram.

[0059] It should be understood that FIG. 2 shows data lineage in a way that each target data element is shown in a state where the source data elements linked to the dependent party are directly related. Those source data elements may depend on other data elements. Since the source data elements of each data element that is the source of the target data element can be identified by tracing back dependencies through the data lineage as shown in FIG. 2, the expression of the data lineage shown in FIG. 2 makes it possible to identify these indirect dependencies. In the embodiments described herein, the data lineage can be stored in such a way that this traceback enables the identification of indirect dependencies. However, it should be understood that in some embodiments, the dependency information can be stored in any suitable way that may include identifying and storing indirect dependencies.

[0060] It should also be understood that a data processing system can manage a large number of data elements (e.g., millions, billions, or trillions of data elements). For example, a data processing system that manages data associated with credit card transactions can process billions of credit card transactions per year, and each transaction can include a plurality of data elements such as, for example, credit card numbers, dates, vendor IDs, purchase amounts, and the like. Thus, data lineage can represent relationships between a large number of data elements, the data entities that contain those data elements, and / or the transformations applied to the data elements. Since data lineage can contain a large amount of information, it is important to present that information in a way that can be digested by a viewer. Accordingly, in some embodiments, the information of the data lineage may be visualized at different levels of granularity. Some aspects of various techniques for visualizing the derived lineage information and techniques for generating and / or visualizing data lineage are described below. That is, (1) U.S. Patent Application No. 12 / 629,466, entitled "Visualizing Relationships Between Data Elements and Graphical Representations of Data Element Attributes," filed on December 2, 2009; (2) U.S. Patent Application No. 15 / 040,162, entitled "Filtering Data Lineage Diagrams," filed on February 10, 2016; (3) U.S. Patent Application No. 14 / 805,616, entitled "Data Lineage Summarization," filed on July 22, 2015; (4) U.S. Patent Application No. 14 / 803,374, entitled "Managing Parameter Sets," filed on July 20, 2015; and (5) U.S. Patent Application No. 14 / 803,396, entitled "Managing Lineage Information," filed on July 20, 2015, each of which is incorporated by reference in its entirety.

[0061] Regardless of how this information is used, lineage information can be derived according to the processes described herein by tools that implement front-end, programming language-dependent processing and back-end, programming language-independent processing. This approach reduces both the programming and processing requirements of the system. For example, back-end processing can be shared across programs implemented in many different programming languages, reducing the amount of memory and other computing resources required for system implementation. Front-end processing is language-specific but can leverage existing parser technologies developed for compilers, interpreters, or other existing tools.

[0062] FIG. 3A is a functional block diagram of front-end processing. The front-end processing shown in FIG. 3A converts source code in any of a plurality of programming languages into a data structure represented in a language-independent manner. FIG. 3A shows the processing of program 310. Program 310 may be one of computer programs 109 (FIG. 1A) and may be written in any suitable programming language.

[0063] Program 310 is provided as input to parser 312. Parser 312 may be a parser of a known structure configured to process program code in the language in which program 310 is written. The output of parser 312 may be parse tree 314. Parse tree 314 may be in a known format and may represent the operations specified by the program code within program 310. Parse tree 314 may be stored in computer memory as a first data structure representing program 310. However, it should be understood that the operations executed according to program 310 can be represented in any suitable manner.

[0064] Next, the parse tree 314 can be processed by an extraction module 316. The extraction module 316, like the parser 312, can be specific to the programming language in which the program 310 is written. When executed, the extraction module 316 can identify elements in the program 310 that represent operations that can create either a data flow or a control flow dependency by traversing the data structure representing the parse tree 314. Upon detecting such an operation, the extraction module 316 can create an entry in a second data structure, herein represented as the dependency data set 318.

[0065] The extraction module 316 can be implemented by applying the inventor's insight that a collection of operations of a program that can introduce dependencies can generally represent, for the purpose of identifying data and control flow dependencies, with a small number of constructs. Further, the number of operations supported by a programming language that can introduce dependencies is also finite and relatively small. Thus, the extraction module 316 may be configured to traverse the parse tree 314 using known techniques to identify each operation in the program 310 that affects the dependencies and input information representing the dependencies into a second data structure.

[0066] Once an operation that affects such a dependency is identified, the extraction module 316 can select one or more dependency components that represent the identified operation. Such a selection may be made using table lookups, heuristics, or other suitable techniques. The extraction module 316 can then parameterize the dependency components using the parameters of the operation identified in the parse tree 314.

[0067] As a specific example, when the extraction module 316 identifies an "add" operation in the parse tree 314, that operation may be mapped to a dependency component. In this example, the "add" operation includes three variables "a", "b", and "c", where "a" is the output of the operation and has a value that depends on the values of "b" and "c". These three variables can be used by the extraction module 316 when parameterizing the dependency component. According to an exemplary embodiment, the dependency component may be referred to as an "assignment". In this example, the dependency component may indicate that a value that depends on source variables "a" and "b" is assigned to target variable "a".

[0068] In some embodiments, it should be understood that whether there are dependencies between variables as a result of executing a program may depend on the values assigned to the variables of the program being executed. However, for simplicity, the extraction module 316 may be configured to classify any operation whose result may depend on the values of one or more other variables as an operation that affects the dependencies. Similarly, in some scenarios, part of the program may not be executed based on the values of one or more variables. Nevertheless, the extraction module 316 may be configured to process all operations of the parse tree 314 that can execute or effectuate the values of the variables and record those operations in the dependency dataset 318 under all circumstances.

[0069] FIG. 3A shows that the dependency dataset 318 includes a plurality of dependency components 320A, 320B... 320ZZ. Each dependency component may be in one form of a finite number of dependency components that characterize the program 310.

[0070] Examples of these structures are provided below. In some embodiments, the extraction module may be programmed to recognize a subset of these structures. In other embodiments, the following examples represent a finite set of configurations recognized by the extraction module. In other embodiments, the dependencies may be represented in other ways.

Example

[0071] Example 1: "Function" According to some embodiments, a part of the code representing a function may appear as follows.

Number

[0072] After this code is converted into an analysis tree and processed, it can be expressed as follows.

Number

[0073] Example 1 shows a single function represented as a procedure definition. From the perspective of dependencies, a procedure has a scope, input, output, or bidirectional parameter. A function is a procedure with a return parameter.

[0074] "3:1 4:1" are the start and end positions in the source file of the generated component. This information can function as a pointer to the code that generated the dependency and can be used to generate a code summary that can be stored in the metadata repository 110. Alternatively or additionally, the extraction module can capture information about the source code that caused the dependency by copying the code line. According to some embodiments, the extraction module may be programmed to show only the code summary where the dependency is created, in the form of a pointer, a copy of the source code, or some other method. The code can be summarized in any suitable way, such as by selecting the code line where the assignment is made, along with one or a finite number of lines before and / or after that line including the assignment. Other rules can be coded into the extraction module to reduce the storage capacity required for the code summary. For example, only the operators of the code line can be copied, or only the lines that explicitly reference the target variables within the code block can be copied.

[0075] Example 2: Assignment As a second example, the following dependencies affecting the operation may appear within the source code.

Number

[0076] Through the front-end processing, its operation may be expressed as a dependency component as follows.

Number

[0077] In this and other examples, angle brackets are used for dependency lists, parentheses are used for parameter and argument lists, and curly braces are used for ranges. For example, "y[x]" indicates a single dependency, where "y" depends on "x". In the language that constitutes the set of dependency components provided as examples in this specification, a dependency always has a single target element (in this case, "y") and a source element list. In this example, the source element list is "[x]", which is composed of one element "x". In this example, pointers to the source code for generating a source code summary at a later processing stage are also shown. Here, the relevant parts of the source code are shown as "21:1 23:1" and "22:5 22:11", which could be pointers to the lines within the source code file where the source code statements creating the corresponding dependencies can be found.

[0078] This example also shows data elements, namely "x" and "y". These elements represent basic / lowest-level / atomic fields. In this example, "x" appears in the source list of the elements and "y" is the (only) target element.

[0079] Knowledge of the directionality of parameters / arguments can be used in dependency analysis. Thus, dependency components may also represent directionality. There are three types, namely IN, OUT, and INOUT. In the language of this embodiment, if directionality is not mentioned, IN is implied.

[0080] It should be understood that this syntax makes the dependency components readable by humans. Since the processes described in this specification can be automated, it is not a requirement of the present invention that the dependency components be human-readable. Thus, any information representing the dependency components can be stored in the computer's memory and used to identify the dependency components.

[0081] Example 3: Return Since the value of the variable to be returned may depend on the processing within the function, including the variables used in the processing of that function, a function that returns the value of a variable may also create dependencies. Function

Number

Number

[0082] "_RETURN" is a special data element name used to specify an unnamed entity returned by a function. This example shows that data elements other than explicitly declared variables may have dependencies or may be created. In the processes described herein, data elements can be processed regardless of how they are created. "_RETURN" is an example of a local data element and has meaning only within a program. However, even local data elements can be used for assignment to "global variables", can control part of the execution of a program, or can be used in other ways to create data flow or control flow dependencies. If the value of "_RETURN" depends on other data elements, those dependencies can then affect the "global variables" or otherwise create recorded dependencies. Thus, "_RETURN" need not be included in the final metadata store created, but may be considered when identifying dependencies in the embodiments described herein.

[0083] In this example, the function returns only a single value. If the function returns multiple values, the "_RETURN" component may be modified to correspond to two or more values. Further, similar to the syntax used to represent other dependency components, "_RETURN" functions as a symbol and it should be understood that it is formatted here for human readability. When implemented in a data processing system, any suitable symbol can be used.

[0084] Example 4: Constant Assignment Dependencies within code containing constant assignments are as follows.

Number

Number

[0085] In this embodiment, "_declare" and "_CONST" are symbols of other special elements such as the "_RETURN" symbol. Note that the dependency components do not contain information regarding the values of the constant variables "u" and "v". Such information is not necessary and can be omitted for analyzing dependencies.

[0086] In some embodiments, the component "_CONST" may be omitted, but in some embodiments, it may be included to make it easier for a person to understand a collection of dependency components representing one or more programs and to simplify the computer processing of data structures including those components.

[0087] Example 5: Call Site Dependencies within the code including a call to a function are defined as follows.

Number

Number

[0088] This example includes a procedure or function call site. The arguments are provided to the procedure at the call site. Data entities functioning as arguments may affect the values of other data elements so that the processing within the called procedure may depend on these arguments.

[0089] In this example, constant values are assigned to the data entities "u" and "v". Therefore, the processing within the procedure based on those arguments produces a certain result and can be captured in the dependency components without indicating a dependency on "u" and "v". In some embodiments, dependency information is generated in a two-stage process, where the first stage may include language-dependent processing and the second stage may include language-independent processing.

[0090] In some exemplary embodiments described herein, dependencies on local variables can be identified and eliminated in the second-stage processing, such that the final dependency information is represented from the perspective of external variables rather than local variables that have a scope only within a function or other code blocks within a program. Thus, in this example, in the first-stage processing, the dependency of the variable "w", and thus the value returned by the function "f5()", can be represented as depending on "u" and "v". In the second-stage processing, "u" and "v", and the value returned by the function "f5()", can be shown to depend on constant values.

[0091] This process of propagating local dependencies can continue throughout the program being processed. When the function "f5()" is called in the program, its return value can be assigned to another variable. In the first-stage processing, that variable can be shown to depend on the value returned by the function "f5()". In the second-stage processing, the dependency on the local variable representing the value returned by the function "f5()" can be replaced with an indication of the global variables on which the value returned by the function "f5()" depends. In this example, the value returned by the function "f5()" can depend on constant values. If the value returned by the function "f5()" depends on external variables, the dependencies may be recorded as depending on those external variables instead of being based on constant values.

[0092] Example 6: Call Sites with Constant Arguments Dependencies within the code that include a call to a function are defined as follows.

Number

Number

[0093] An argument given a constant value can be treated as if it were a data element to which the constant value is assigned.

[0094] Example 7: Reading or Writing to the Global Table Operations on data entities other than elemental data elements can also be handled by the systems described herein. In some embodiments, the data entities may be hierarchical. Operations can be performed on the data entities at any level of the hierarchy. Thus, the dependency components can correspond to the identification of data entities at any applicable level of the hierarchy. One such data entity is, for example, a global table that is an example of a data set. The data set may contain other data entities within it. For example, a table contains rows and columns, which can be regarded as data entities.

[0095] Operations that affect or depend on the table itself can be performed, potentially affecting any of the data entities within the table. The table also has rows, columns, or other subdivisions, creating lower levels in the hierarchy. Operations may affect or depend only on data elements within subdivisions such as rows and columns. However, operations that depend on any subset of the data set are possible, such as operations on a single data element in a row or column of the table.

[0096] To specify a data entity in such a hierarchy, tuple notation can be used. The entity at the highest level of the hierarchy is specified first, followed by the entity at the next highest level in the hierarchy. As a specific example, a tuple can be specified in "dot" notation, where each symbol forming the tuple can be separated by a ". ". In the following example, the first value of the tuple (here "T1") identifies a table. The second value of the tuple (here "C1") identifies a column of that table. Instead, the second value could identify a row. In some embodiments, a tuple can include values that identify both a row and a column, specifying data elements at an even lower level of the hierarchy.

[0097] The dependencies of operations in such a table are as follows.

Number

Number

[0098] Example 8: An expression having a plurality of elements used as arguments at a call site The dependencies created by a function call are defined as follows.

Number

Number

[0099] This example shows the difference between a dependency source list enclosed in square brackets "[]" and arguments enclosed in parentheses "()".

[0100] Example 9: A conditional statement that creates a control flow dependency This embodiment shows how programming statements that create conditional control flows can be represented for dependency analysis. In this embodiment, the conditional statement is in the "if then... else" format. This embodiment is related to a function that includes a conditional statement as follows.

Number

[0101] Such "if then... else" statements are supported in many programming languages. How such a statement or any statement is represented within a program can vary from language to language. Since the processing of the extraction module 316 can be specific to one programming language, any variations in how the "if then... else" statement is represented for each language can be accounted for by configuring the extraction module to identify the components of the specific language in which that module is used. Such statements can be represented by the following dependency components.

Number

[0102] In this embodiment, for the purpose of representing the dependency components, the contents of the "then" and "else" branches from the "If" statement are merged into a single block. In the execution of the program, either one of the branches is executed, but since the determination of which branch to execute depends on the same data entity, it is possible for the dependencies of both the "then" block and the "else" block to be represented in the same way. In this embodiment, since which branch is taken is determined by the value of the variable "x", for the control flow, both depend on the variable "x". The control flow dependency of that block can be represented as "_condition[x]" as described above.

[0103] Code that shares the same control flow dependencies may be treated as one block. The dependency analysis according to some embodiments described herein is based on static code analysis, and such treatment is appropriate to identify dependencies that may occur during the execution of the code.

[0104] In this example, when each branch is executed, a value is assigned to the variable "y" based on a constant. Considering both branches together in this way further leads to the conclusion that within that code block, a value that depends on a constant is assigned to y. Therefore, the value returned by the function depends on a constant, but a particular constant depends on the value of x. Such dependencies between control and data flow can be expressed as "_condition[x,_CONST]".

[0105] One skilled in the art will recognize that for a programming construct that creates three or more possible control flows, there are three or more code blocks that are conditionally executed. The execution of each of these code blocks may potentially depend on different variables. In that case, the programming structure may be represented by different dependency components of each block or a subset of blocks.

[0106] Example 10: Control Flow Dependencies Created by a "While Do" Loop Control flow dependencies may also be created by a statement that implements a "while do" loop. In such a loop, a code block is executed while the mentioned expression evaluates to true. In this scenario, there is a control flow dependency between that block and the data entities that affect the value of that expression. An example of a function that uses such a statement is as follows.

Number

[0107] The dependencies of that function can be expressed as follows.

Number

[0108] In this embodiment, the code block inside the "while do" loop can be executed at least once. Therefore, the dependencies created by the statements inside that block are represented by the extracted dependency information. Additional dependencies can be created by variables that can change inside the loop such that their values depend on the number of times the loop is executed. In that scenario, any such variable depends on a data entity whose value can affect the number of times the loop is executed.

[0109] In this embodiment, the loop is executed several times depending on the values of "i" and "x". For this reason, a condition is mentioned, that is, "_condition[i,x]" represents that the values of "i" and "x" can affect the control flow. The indication that "y[y,i]" and "i[i,_CONST]" inside the parentheses under the condition reflect the variables "y" and "i" respectively that can be changed in that loop means that they will have values that depend on the number of times the loop is executed.

[0110] Dependencies created by other conditional statements such as "for", "while do", "do until", "switch", or "case" can be represented in a similar way by this conditional block structure. As a specific example, "for" and "while do" obtain the condition and the block, and since the code inside the block is guaranteed to be executed at least once, "do until" copies the content of the block and places it above the condition / block.

[0111] Example 11: Level Call Tree This example shows that dependencies can be represented by a limited number of components even when a function calls other functions. In this example, three functions are defined, and the second function "g()" calls the first function "h()". The third function "f()" then calls the second function as follows.

Mathematics

[0112] This multi-level call tree can be represented in such a way that each function that calls another function captures the calls to other functions for subsequent processing.

Mathematics

[0113] As can be seen from this embodiment, there is no need to introduce additional components to represent the multi-level call tree in the program with the extracted dependency information. Rather, the dependency of the first function is captured by the above-described components (here, "y[x]"). Other functions are recorded in the extracted dependency information in a way that indicates that they depend on other functions. Any dependencies that may be introduced by one function calling another function may be resolved in the back-end (second phase) processing where local dependencies are propagated.

[0114] Example 12: Recursion Similar to the multi-level call tree, no additional components are needed to represent the recursion of the program with the extracted dependency information. Rather, information indicating that the recursion can be saved as part of a lower-level parse tree, and the back-end processing can calculate any additional dependencies that may arise from that recursion.

[0115] Therefore, as follows, a function calls itself recursively.

Mathematics

Mathematics

[0116] Here, since this variable is initially set to a constant value of "0" and its value cannot be changed throughout, it is shown that the variable "accumulator" has a dependency on the constant value. Further, conditional flow is shown depending on the value of the variable "x". Such a condition results from the "if...else" component within the statement "if(x==0)". Within the block associated with that condition, the variable "y" is given a value based on the value of "accumulator", and the value of "accumulator" may depend on the function "F11()".

[0117] This information is sufficient to enable the identification of dependencies in the back - end processing, even if it is not sufficient to specify the values of these variables.

[0118] Returning to FIG. 3A, the front - end processing can be executed for each program. In this example, programs 310, 320, and 350 are shown, but any number of programs may exist in the data processing system. The results of such processing are, respectively, dependency data sets 318, 338, or 358 for each program. Each dependency data set has dependency components such as dependency components 320A, 320B, and 320ZZ as described above for dependency data set 318. As a result, the dependency data set includes information regarding the selected operations reflected in the parse tree, particularly the operations that affect dependencies. Those selected operations may be reflected by the same set of dependency components, regardless of the source language in which the program processed to generate the dependency data set is written.

[0119] The specific processing to reach those dependency components may vary for each programming language. Thus, the parser 332 can be configured using techniques known in the art to generate a parse tree 334 suitable for the language in which program 330 is written. Similarly, the parser 352 may be configured to generate a parse tree 354 suitable for the language in which program 350 is written.

[0120] Each of the extraction modules, such as 336 and 356, can also be configured based on a specific programming language. As described above, each is output as a dependency component of the dependency dataset. However, depending on the way the executable statements that can generate dependencies are represented in the associated programming language, the input to each may be in a different format.

[0121] Figures 3B, 3C, and 3D provide examples of such front-end processing. In this example, code segment 360 (Figure 3B) appears in the program being processed. In this example, code segment 360 includes the assignment statement "a = b + c". That statement is within a conditional block defined by the conditional statement "If(x = 1){...}". This conditional statement is in the "If then...else" format with a null "else". That statement is defined within a function that is part of the program. The function definition can define the scope of variables used only within that function.

[0122] Code segment 360 can be represented as part of the parse tree 368 (Figure 3C), and other segments of the code are processed in a similar way to represent the rest of the program in other parts of the parse tree (not shown). In this example, node 370, and the nodes below node 370, are within the scope of the program shown as PROGRAM1.

[0123] Node 372 indicates that FUNCTION1 is defined within PROGRAM1. Node 372, and the nodes below node 372, are within the scope of the function shown as FUNCTION1. In this example, nodes 373 and 374 are immediately below node 372, and other nodes are below those nodes.

[0124] In this example, node 374 shows a data flow dependency. Node 373 shows a control flow dependency. Node 374 represents one or more statements in the program being analyzed that perform an assignment. In this example, in relation to node 374, a code summary 375 of those statements can be recorded. As shown in the above example, this code summary can be one or more pointers to lines in the source code file. Alternatively or additionally, the source code summary 375 may include a copy of the line or part of the line of code. The code lines reflected in the summary are an executable representation of the rules, heuristics, or selection logic that identify the code statements creating the dependency, and in some embodiments, can be automatically selected by the execution of additional statements in the program that provide the context of that statement. For example, the statements providing the context can be a certain number of statements before and / or after the statement creating the dependency. Based on the nature of the statement creating the dependency, additional or different statements can be selected. For example, if the statement includes source variables, the summary can include the statement that defines the variable, or the most recent statement that assigns a value to the variable, even if those statements do not appear near the statement creating the summarized dependency in the program.

[0125] Node 376 represents the left side of the assignment represented by node 374. This example shows that the target variable, which is the variable to which a value is being assigned, includes only one variable, the variable "a" represented by node 378 in this example.

[0126] Node 380 represents the right side of the statement that generates the value assigned to the variable "a". The right side includes an operation (in this case addition), as shown by node 382. The addition operation has two inputs, the variables "b" and "c", shown by nodes 392 and 394 respectively.

[0127] Node 373 represents the control flow dependency of FUNCTION1. The control dependencies associated with FUNCTION1 indicate that any value that depends on an operation within the scope of FUNCTION1 also has a control flow dependency, because depending on the control flow, the statements within FUNCTION1 that create the dependency may (or may not) be executed in a particular case. As shown by node 379, its control flow dependency is based on the value of variable "X". In this example, depending on the value of X, an assignment to variable "a" represented by node 374 may (or may not) occur. This control flow dependency captures the conditional execution associated with the "If then…else" statement in Figure 3B.

[0128] Next, a portion of the parse tree shown in Figure 3C can be converted into a dependency component 396 (Figure 3D) that can be included in the dependency dataset. Here, the dependency component indicates that variable "a" is the target and its value depends on source variables "b" and "c". The nomenclature of this example, unlike that of Example 1 above, also represents an assignment. However, it should be understood that both the above example and Figure 3B show the dependency components in a human-readable form and that other forms can be used. Both show the same dependency components and can be encoded in the computer in any way programmed for the computer to recognize and process. As a specific example, the dependency information may be represented as a list of each target variable. Each list may include all of the source variables on which the target depends.

[0129] A portion of the analysis represented by dependency component 396 fits within the scope of FUNCTION1. That scope can be associated with the dependency information, as shown in Figure 3D by the scope identifier "FUNCTION1". Such scope information is formatted for human understanding. It should be understood that for computer processing, the scope information can be represented in other ways, such as associating each list of source variables with respect to a target with the scope.

[0130] The control flow dependencies represented in FIG. 3C can also be reflected in the dependencies of the target variable "a". In FIG. 3D, the control flow dependencies are shown by a source variable "X" with a symbol (in this example, "*") to indicate that the dependency is a control flow dependency. Similar to other examples, the notation is intended to be human-readable. The control flow dependencies can be captured in a computer-processable manner that may include associating symbols such as "*" with source variables in the list. Alternatively or additionally, the control flow dependencies can be shown by storing the source variables that give rise to the control flow dependencies separately or separating them from the source variables of the data flow dependencies in some other appropriate way. Such data storage techniques can be used in embodiments where processing is distinguished based on control flow versus data flow dependencies. In embodiments where processing is not distinguished based on the nature of the dependencies, the source variables of the control flow dependencies may be stored and processed in the same manner as the data flow dependencies. Or, if processing based on control flow dependencies is not performed, the control flow dependencies may be omitted.

[0131] FIG. 3C shows dependency information encoded within a second data structure for only a portion of one program. Similar processing may be performed on the entire program. As shown in FIG. 3A, when multiple programs are executed in a data processing system, similar processing may be performed for each program. In the examples provided here, the programs are shown as being processed independently. However, it should be understood that a program can call functions defined in other programs. In that case, the dependency information may represent that a variable in one program depends on a variable from another program by indicating the scope associated with the variable.

[0132] The process that results in the representation of dependencies shown in FIG. 3D can indicate dependencies on local variables and dependencies on local variables. Local variables exist only within a program and, by definition, are not inputs or outputs to the program. Thus, in some embodiments, information about those local variables may be omitted from the dependency information stored in a metadata repository 110 (FIG. 1A) or the like. Nevertheless, this local dependency information can first be captured and used in the process of deriving the stored dependency information.

[0133] FIG. 4 provides an example of such a process that is executed at the backend after language-dependent processing is complete. It is shown that the processing is being performed by one or more tools that are not language-dependent. Each of the sub-modules, such as 316, 336, or 356, can output the dependency datasets 318, 338, 358 in a common format so that subsequent processing cannot be language-dependent.

[0134] In the example of FIG. 4, the processing on the dependency datasets 318, 338, 358 can start with a dependency processing module 410 that can operate on both local and global dependencies. The processing may create a list or other suitable representation of the dependencies of all variables as described above in connection with FIG. 3D. In some embodiments, dependencies on and / or to local variables and external variables can be stored. However, in some embodiments, more efficient operation of the data processing system can be achieved by removing such local dependencies that are not ultimately used.

[0135] Accordingly, further processing can be performed by the module 420 that propagates local dependencies. The module 420 propagates and then removes local dependencies. Local dependencies can be identified as variables whose values are only used in processing within the internal scope of the program and whose values are not loaded from or saved to an external data source. As a specific example, in the above-described Example 11, the values of "s" and "t" in the function "g()" are not saved or accessed from an external data source. Rather, the values of s and t are derived only from "x". Once the processing is completed as described above, the target data element that depends on "s" or "t" only needs to indicate that it depends on "x". Accordingly, references to local variables such as "s" or "t" can be deleted without losing information. Similarly, since it is not used in some embodiments, the data representing the data dependencies of "s" or "t" may also be deleted. This information is only associated with local dependencies and can be deleted by the module 420.

[0136] The module 420 can identify local variables by traversing the data stored for the program to determine whether its variables exist outside the program or are only used in internal calculations. As described above in connection with FIG. 3C, information about the program can be stored in a way that records range information about code blocks. The range information can reveal nesting as a result of the code with the inner range being shown as defined within another code with an outer range.

[0137] Accordingly, the module 420 can start processing in each innermost range. The module 420 may determine whether to access each variable defined and / or used within that range from outside that range. If not, the variable can be flagged as a local variable. If a variable identified as a local variable depends on one or more variables, the values of those variables can be replaced with the local variable as the source in the dependency component.

[0138] In some cases, such processing may result in local variables being replaced by external variables as sources of dependencies. However, in some cases, local variables may depend on other local variables. To account for such a possibility, the processing by module 420 may be recursively executed such that, when a local variable depends on other local variables, those other local variables are then processed until all local variables within its scope are replaced by external variables in the dependency expressions.

[0139] This process of "propagating" local dependencies can be repeated for each outer scope, with nested scopes being processed in order from the innermost to the outermost. It should be understood that complex programs may have many parts that create non-dependent nested scopes. Each nesting can be processed in order until the entire program is processed, and dependencies on local variables are replaced by the external variables on which those local variables then depend.

[0140] Once this local dependency propagation processing is complete, information regarding the dependencies of local variables may no longer be needed, unless the dependencies of the local variables are retained. Thus, module 420 may either delete the dependency information regarding the local variables or save the retained dependency information in a way that the information regarding local dependencies is not processed thereafter.

[0141] Loader module 430 can save the retained dependency information in metadata repository 450. In the embodiment shown in FIG. 4, the metadata repository is shown as saving data dependencies 452 and control dependencies 454 separately. However, as noted above, such information can be saved in the same data structure, and flow control dependencies can be marked separately from data flow dependencies. Alternatively, in some embodiments, no distinction may be made between data flow dependencies and control flow dependencies.

[0142] The loader module 430 can convert the representation of the dependency information as described above in relation to FIG. 3C into the format of the metadata repository 450. As a specific example, the metadata repository 450 may store the dependency information as a collection of tables or lists indicating the source variables on which each target variable depends. The information may be stored such that the data processing system can generate the data lineage diagram 200 (FIG. 2).

[0143] Furthermore, in some embodiments, the loader module 430 may store the code summary 456 in the metadata repository. As described above, when the source code statements that generate dependencies are identified, they can be stored in relation to the dependency components in the dependency dataset, either alone or in combination with statements that can reveal information about their dependencies. This code information may be stored in whole or in part as the code summary 456. For example, the program may include multiple instances where values are assigned to variables. In some embodiments, the loader module 430 may store the code associated with each of these instances. Alternatively or additionally, the loader module 430 may selectively store the code associated with only some of these instances. The loader module 430 may be programmed, for example, to select the code associated with the first instance of such an assignment, or the first instance until the code summary reaches a predetermined size, or apply rules or coded heuristics to select the instance that is most likely to be accessed by the developer. However, such selection may be performed by the loader module 430 in any suitable way to limit the size of each code summary. Alternatively or additionally, such selection may be performed by the sub-module described in relation to FIG. 3A, or other suitable components.

[0144] Regardless of how the selection is made, by saving identification information of code that may cause dependencies on data elements or specific code blocks, a data processing system can enable additional functionality to be provided to developers such as user 102 (FIG. 1A). For example, code summary 456 can be used by a programmer to modify or debug a program and attempt to understand how the value was assigned to a variable with an unexpected value.

[0145] FIG. 5 is a flowchart of an exemplary process 500 for generating dependency information according to some embodiments of the techniques described herein. Process 500 can be executed by any suitable system and / or computing device, for example, by data processing system 105 described with reference to FIG. 1A and programmed with the modules described above with respect to FIGS. 3A and 4.

[0146] Process 500 begins with front-end processing 510. Front-end processing 510 may be executed for each program used to configure the data processing system.

[0147] Within front-end processing 510, the process begins at act 502 and a parser for the program being processed is selected. The parser can be selected based on the programming language of the particular program being processed. According to some embodiments, the data processing system may store parsers for each programming language supported by the data processing system. In other embodiments, the data processing system can include a parser generator that can generate a parser for any such programming language, and the data processing system can be configured to process a program described in a desired language by providing a grammar file or other information regarding that programming language that the parser generator uses.

[0148] In act 504, the selected parser can be used to analyze the program. The result of act 504 can be a parse tree for the program.

[0149] In act 506, the parse tree can be flattened. As described above, flattening the parse tree can be performed by a lower module configured for the programming language of the program being processed. The result of the flattening act can be a data structure that stores dependency components for data elements and code blocks within the program.

[0150] In act 508, the dependency data sets generated for all programs (program 109, FIG. 1A, etc.) can be optionally combined as needed. In some embodiments, a program written in one language may implement an overall data processing function by calling functions or otherwise incorporating code written in another programming language. Thus, the parse trees generated by processing each program independently can be combined to represent such behavior. The combined dependency data set can be the front-end output that is transmitted to the back-end processing 520. Alternatively or additionally, the parse trees prepared for different programs can be combined during back-end processing by identifying, for each target variable, the source variables on which that variable can depend, regardless of which program contains the statements that create those dependencies.

[0151] Within back - end processing 520, the process begins at act 522. At act 522, local dependencies are identified. In the case of a program with nested scopes, this process can be first executed for each part of the program representing the innermost scope. The identification of local variables can be performed in an appropriate way, such as by accessing a table that can be constructed during the parsing operation for each variable. That table can indicate where a variable is accessed or modified so that variables that are only accessed or modified within the scope being processed can be identified as local variables for that scope.

[0152] At act 524, local dependencies can be propagated. Propagating local dependencies can include using dependencies on external variables that local variables depend on instead of dependencies based on local variables.

[0153] The process proceeds to decision block 530 and branches depending on whether the scope being processed is an external scope. If not, the process branches back to act 522 where the acts of identifying and propagating local variables are repeated. In this way, a part of the program starting from the innermost scope and proceeding to the outermost scope is processed.

[0154] Once all of the outermost scope is reached, the process branches to act 532. At act 532, dependency information regarding local variables can be removed.

[0155] At act 534, the remaining information may be stored in a metadata repository or in other appropriate locations for later use.

[0156] The process 500 shown in FIG. 5 can be executed in any suitable computing system environment. FIG. 6 illustrates an example of a suitable computing system environment 700 in which the techniques described herein can be implemented. The computing system environment 700 is only an example of a suitable computing environment and is not intended to suggest any limitation as to the scope of use or functionality of the techniques described herein. The computing environment 700 should not be interpreted as having any dependency or requirement related to any one or combination of the components illustrated in the exemplary operating environment 700.

[0157] The techniques described herein can be used with a number of other general purpose or special purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that may be suitable for use with the techniques described herein include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, and the like.

[0158] A computing environment can execute computer-executable instructions such as program modules. In general, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The techniques described herein may be executed in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including memory storage devices.

[0159] Referring to FIG. 6, an exemplary system for implementing the techniques described herein includes a general-purpose computing device in the form of a computer 710. The components of computer 710 may include, but are not limited to, a processing device 720, a system memory 730, and a system bus 721 that couples various system components including the system memory to the processing device 720. The system bus 721 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Extended ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus, also known as a Mezzanine bus.

[0160] Computer 710 generally includes various computer-readable media. Computer-readable media can be any available media that can be accessed by computer 710, and includes both volatile and non-volatile media, removable and non-removable media. By way of example, and not limitation, computer-readable media may include computer storage media and communication media. Computer storage media is implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data, and includes volatile and non-volatile, removable and non-removable media. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or other media that can be used to store desired information and can be accessed by computer 710. Communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed so as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the above is also to be included within the scope of computer-readable media.

[0161] The system memory 730 includes computer storage media in the form of volatile and / or non-volatile memory such as read-only memory (ROM) 731 and random access memory (RAM) 732. The basic input / output system 733 (BIOS), which contains basic routines that help transfer information between elements within the computer 710 during startup and the like, is generally stored in the ROM 731. The RAM 732 generally contains data and / or program modules that are immediately available and / or currently being operated on by the processing unit 720. By way of example and not limitation, FIG. 7 illustrates an operating system 734, application programs 735, other program modules 736, and program data 737.

[0162] The computer 710 can also include other removable / non-removable, volatile / non-volatile computer storage media. By way of example only, FIG. 7 illustrates a hard disk drive 741 that reads from or writes to non-removable, non-volatile magnetic media, a flash drive 751 that reads from or writes to removable, non-volatile memory such as a flash memory 752, and an optical disk drive 755 that reads from or writes to removable, non-volatile optical disks such as a CD-ROM or other optical media. Other removable / non-removable, volatile / non-volatile computer storage media that can be used in the exemplary operating environment include, but are not limited to, magnetic tape cassettes, flash memory cards, digital versatile disks, digital video tapes, solid state RAM, solid state ROM, and the like. The hard disk drive 741 is generally connected to the system bus 721 through a non-removable memory interface such as the interface 740, and the magnetic disk drive 751 and the optical disk drive 755 are generally connected to the system bus 721 by a removable memory interface such as the interface 750.

[0163] The drives described above and illustrated in FIG. 7, and the computer storage media associated therewith, provide storage of computer-readable instructions, data structures, program modules, and other data of computer 710. In FIG. 7, for example, hard disk drive 741 is illustrated as storing operating system 744, application programs 745, other program modules 746, and program data 747. These components may be the same as, or different from, operating system 734, application programs 735, other program modules 736, and program data 737. Note that operating system 744, application programs 745, other program modules 746, and program data 747 are given different numbers here only to illustrate that they are at least different copies. A user can input commands and information into computer 710 via input devices such as keyboard 762 and a pointing device 761 generally called a mouse, trackball, or touchpad. Other input devices (not shown) may include a microphone, joystick, game pad, satellite dish, scanner, and the like. These and other input devices are often connected to processing device 720 by user input interface 760 coupled to the system bus, but may also be connected by other interfaces and bus structures such as a parallel port, game port, or universal serial bus (USB). Monitor 791 or other type of display device is also connected to system bus 721 via an interface such as video interface 790. In addition to the monitor, the computer may also include other peripheral output devices such as speakers 797 and printer 796 that can be connected through output peripheral interface 795.

[0164] Computer 710 can operate in a networked environment using logical connections to one or more remote computers, such as remote computer 780. Remote computer 780 can be a personal computer, a server, a router, a network PC, a peer device, or other common network node, and generally, only memory storage device 781 is illustrated in FIG. 7, but includes many or all of the elements described above in relation to computer 710. The logical connections depicted in FIG. 7 include local area network (LAN) 771 and wide area network (WAN) 773, but may also include other networks. Such networking environments are commonplace in offices, enterprise-scale computer networks, intranets, and the Internet.

[0165] When used in a LAN networking environment, computer 710 is connected to LAN 771 through network interface or adapter 770. When used in a WAN networking environment, computer 710 generally includes a modem 772, or other means for establishing communications over a WAN 773 such as the Internet. Modem 772, which can be internal or external, may be connected to system bus 721 via user input interface 760 or other appropriate mechanism. In a networked environment, program modules depicted in relation to computer 710, or portions thereof, may be stored in a remote memory storage device. By way of example and not limitation, FIG. 7 illustrates remote application program 785 as residing in memory device 781. It will be understood that the network connections shown are exemplary and that other means of establishing a communications link between computers may be used.

[0166] Although some aspects of at least one embodiment of the present invention have been described as above, it is to be understood that various changes, modifications, and improvements will readily occur to those skilled in the art.

[0167] Such changes, modifications, and improvements are intended to be part of this disclosure and are intended to be within the spirit and scope of the invention. Further, while the advantages of the invention are presented, it is to be understood that not all embodiments of the technology described herein will encompass all of the described advantages. Some embodiments may not implement any of the features described herein as being advantageous, and in some cases, one or more of the described features may be implemented to obtain further embodiments. Accordingly, the above description and drawings are merely examples.

[0168] The above-described embodiments of the technology described herein may be implemented in any of a number of ways. For example, these embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code may be provided on a single computer or distributed among multiple computers and may be executed on any suitable processor or group of processors. Such processors may be implemented as integrated circuits and include integrated circuit components having one or more processors known in the industry by names such as CPU chips, GPU chips, microprocessors, microcontrollers, or coprocessors. Alternatively, the processor may be implemented in a custom circuit such as an ASIC or in a semi-custom circuit resulting from the configuration of a programmable logic device. As a further alternative, the processor may be part of a larger circuit or semiconductor device, whether commercially available, semi-custom, or custom. As a specific example, some commercially available microprocessors have multiple cores, such that one or a subset of the multiple cores can constitute the processor. However, the processor can be implemented using any suitable format of circuitry.

[0169] Furthermore, it is understood that the computer may be embodied in any of a number of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, the computer may be incorporated into a device that is generally not considered a computer, such as a personal digital assistant (PDA), a smartphone, or any other suitable portable or fixed electronic device, provided that the device has suitable processing capabilities.

[0170] Also, the computer may have one or more input devices and output devices. These devices can be used, in particular, to present a user interface. Examples of output devices that can be used to provide a user interface include a printer or display screen for visual representation of output, and a speaker or other sound generation device for audible representation of output. Examples of input devices that can be used for a user interface include a keyboard, as well as pointing devices such as a mouse, a touchpad, and a digitizer tablet. As another example, the computer may receive input information by voice recognition or in other audible formats.

[0171] Such a computer can be interconnected by one or more networks in any suitable form, including a local area network or a wide area network such as a corporate network or the Internet. Such networks may be based on any suitable technology, and may operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.

[0172] Also, the various methods or processes outlined herein may be encoded as software executable on one or more processors using any one of a variety of operating systems or platforms. Additionally, such software may be written using any of a number of suitable programming languages and / or programming or scripting tools, and may be compiled as executable machine code or intermediate code to be executed on a framework or virtual machine.

[0173] In this regard, the present invention may be embodied as a computer-readable storage medium (or multiple computer-readable media) (e.g., computer memory, one or more floppy disks, compact disc (CD), optical disc, digital video disc (DVD), magnetic tape, flash memory, circuit configurations in a field programmable gate array or other semiconductor device, or other tangible computer storage media) encoded with one or more programs that, when executed on one or more computers or other processors, perform a method of implementing the various embodiments of the present invention described above. As is apparent from the above examples, a computer-readable storage medium can hold information for a sufficient time to provide computer-executable instructions in a non-transitory form. Such one or more computer-readable storage media may be portable so that the one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement the various aspects of the present invention as described above. As used herein, the term "computer-readable storage medium" encompasses only non-transitory computer-readable media that can be regarded as a product (i.e., an article of manufacture) or a machine. Alternatively or additionally, the present invention may be embodied as a computer-readable medium other than a computer-readable storage medium, such as a propagated signal.

[0174] The terms "program" or "software" are used generically herein to refer to any type of computer code or set of computer-executable instructions that can be used to program a computer or other processor to implement the various aspects of the present invention as described above. Additionally, according to certain aspects of the present embodiment, one or more computer programs that, when executed, perform the methods of the present invention need not reside on a single computer or processor and may be distributed in a modular fashion among a number of different computers or processors to implement the various aspects of the present invention.

[0175] Computer-executable instructions can be in many forms, such as program modules, executed by one or more computers or other devices. In general, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Generally, the functionality of program modules may be combined as desired or distributed in various embodiments.

[0176] Also, data structures may be stored on a computer-readable medium in any suitable form. For simplicity of illustration, data structures may be shown as having fields related by location within the data structure. Such relationships can also be achieved by allocating locations within the computer-readable medium that convey the relationships between fields to the storage of the fields. However, any suitable mechanism may be used to establish the relationships between the information in the fields of the data structure, including by using pointers, tags, or other mechanisms that establish relationships between data elements.

[0177] Various aspects of the present invention may be used alone, in combination, or in various arrangements not specifically described in the embodiments described above. Thus, in its application, it is not limited to the details and arrangements of the components described in the above description or illustrated in the drawings. For example, aspects described in one embodiment can be combined with aspects described in other embodiments in any manner.

[0178] Also, the present invention may be embodied as a method by way of example. The acts performed as part of this method may be ordered in any suitable manner. Thus, embodiments may be constructed in which the acts are performed in an order different from that shown (which may include performing some acts simultaneously even if shown as sequential acts in the exemplary embodiments for purposes of illustration).

[0179] Furthermore, some acts are described as being performed by a "user". It is to be understood that the "user" need not be a single individual, and in some embodiments, the acts attributed to the "user" may be performed by a team of multiple individuals and / or an individual in combination with computer-aided tools or other agencies.

[0180] The use of ordinal terms such as "first", "second", "third", etc. in the claims that modify claim elements does not, by itself, imply priority, precedence, or order of one claim element over another claim element, or the temporal order in which acts of a method are performed, but rather these claim elements are distinguished by using them as mere labels to distinguish one claim element having a certain name from another element having the same name (except for the use of ordinal terms).

[0181] Also, the expressions and terms used in this specification are for illustrative purposes and are not to be regarded as limiting. The use of "including", "comprising", "having", "containing", "involving", and variations thereof in this specification is meant to cover the items listed thereafter and their equivalents, as well as additional items.

Claims

1. A dependency analyzer for use in a data processing system configured to execute a program in any of a plurality of computer programming languages that accesses a plurality of records including fields, the dependency analyzer being configured for dependency analysis of programs in the plurality of computer programming languages, the dependency analysis including calculating dependency information related to field-level data lineage across programs in the plurality of computer programming languages, the dependency analyzer being A front end configured to process a plurality of programs that control the data processing system, the plurality of programs being written in a plurality of programming languages of the plurality of programming languages, the front end including a plurality of front end modules, each front end module being configured to receive as input a computer program of the plurality of programs written in a programming language of the plurality of programming languages and to output a language-independent data structure representing dependencies to fields of the plurality of records created within the program, the language-independent data structure being composed of one or more dependency components, each dependency component including an indication of at least one field and a dependency associated with the at least one field. A back end configured to receive the language-independent data structures from the plurality of front end modules and to output dependency information of the data processing system representing dependencies created within the plurality of programs, the dependency information including field-level data lineage across the plurality of programs, the dependency information including field-level dependencies introduced into a plurality of the plurality of programs.

2. The front-end module generates an analysis tree by analyzing an input program, and subordinates the analysis tree by representing operations in the analysis tree that create dependencies using components from a set of dependency components, thereby generating a language-independent data structure, the dependency analyzer according to claim 1, configured to do so.

3. The dependency information represents data flow dependencies, the dependency analyzer according to claim 2.

4. The dependency information represents control flow dependencies, the dependency analyzer according to claim 2.

5. A first tool configured to load the dependency information into a metadata repository, A second tool configured to receive, as input, an instruction for a field used by the data processing system, access the metadata repository, and output an instruction for a program among the plurality of programs that introduces a dependency affecting the indicated field, the dependency analyzer according to claim 1, further comprising.

6. The data processing system further includes a plurality of data sources, The plurality of programs are configured to perform conversions on data within the plurality of data sources, The plurality of data sources are heterogeneous, the dependency analyzer according to claim 1.

7. The plurality of data sources include at least one data source that is an ORACLE database or a SAP database, or a dataset stored in a HADOOP distributed file system, the dependency analyzer according to claim 6.

8. A method performed by a dependency analyzer used in a data processing system configured to execute a program in any of a plurality of computer programming languages that access a plurality of records including fields, the dependency analyzer being configured for dependency analysis of programs in the plurality of computer programming languages, the dependency analysis including calculating dependency information related to field-level data lineage across programs in the plurality of computer programming languages, the method comprising: A plurality of programs that control the data processing system are processed by a front end of the dependency analyzer, wherein the plurality of programs are described in a plurality of programming languages among the plurality of programming languages, and the front end includes each front end module receiving, as an input, a computer program of the plurality of programs described in a programming language among the plurality of programming languages, and outputting a language-independent data structure representing dependencies to fields of the plurality of records created in the program, the plurality of front end modules being configured such that the language-independent data structure is composed of one or more dependency components, each dependency component including an indication of at least one field and a dependency associated with the at least one field, and receiving, by the back end, a language-independent data structure from the plurality of front end modules and outputting dependency information of the data processing system representing dependencies created in the plurality of programs, wherein the dependency information includes field-level data lineage across the plurality of programs, such that the dependency information includes field-level dependencies introduced into a plurality of the plurality of programs, and A method including this.

9. The method according to claim 8, further comprising generating a language-independent data structure by analyzing an input program by the front end module to generate an analysis tree and subordinating the analysis tree by representing operations in the analysis tree that create dependencies using components from a set of dependency components.

10. The method according to claim 9, wherein the dependency information represents data flow dependencies.

11. The method according to claim 9, wherein the dependency information represents control flow dependencies.

12. loading the dependency information into a metadata repository by a first tool The method according to claim 8, further comprising receiving, by a second tool, an indication of a field used by the data processing system as input, accessing the metadata repository, and outputting an indication of a program of the plurality of programs that introduces a dependency affecting the indicated field.

13. The data processing system further includes a plurality of data sources, The method further includes performing a transformation on data within the plurality of data sources, The plurality of data sources are heterogeneous, the method according to claim 8.

14. The method according to claim 13, wherein the plurality of data sources comprise at least one data source that is at least one of an ORACLE database or a SAP database, or a dataset stored in a HADOOP distributed file system.

15. At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to execute a method of using a dependency analyzer for use in a data processing system configured to execute a program in any of a plurality of computer programming languages to access a plurality of records including fields, the dependency analyzer being configured for dependency analysis of programs in the plurality of computer programming languages, the dependency analysis including calculating dependency information related to field-level data lineage across programs in the plurality of computer programming languages, and using the dependency analyzer comprises Controlling a plurality of programs that control the data processing system by the front end of the dependency analyzer, wherein the plurality of programs are described in a plurality of programming languages among the plurality of programming languages, and the front end, each front end module receives as input a computer program of the plurality of programs described in the programming language among the plurality of programming languages, and outputs a language-independent data structure representing the dependencies to the fields of the plurality of records created within the program, the plurality of front end modules being configured to: the language-independent data structure is composed of one or more dependency components, each dependency component including an indication of at least one field and a dependency associated with the at least one field; Receiving, by a back end, a language-independent data structure from the plurality of front end modules, and outputting dependency information of the data processing system representing dependencies created within the plurality of programs, wherein the dependency information includes field-level data lineage across the plurality of programs, and the dependency information includes field-level dependencies introduced into a plurality of the plurality of programs; Including at least one non-transitory computer-readable storage medium.

16. The processor-executable instructions cause the at least one computer hardware processor to generate a parse tree by parsing an input program by the front-end module, and to subordinate the parse tree by representing operations in the parse tree that create dependencies using components from a set of dependency components, thereby generating a language-independent data structure, the at least one non-transitory computer-readable storage medium of claim 15.

17. The dependency information represents data flow dependencies, the at least one non-transitory computer-readable storage medium of claim 16.

18. The dependency information represents control flow dependencies, the at least one non-transitory computer-readable storage medium of claim 16.

19. The processor-executable instructions cause the at least one computer hardware processor to cause the dependency information to be loaded into a metadata repository by a first tool, cause a second tool to receive, as an input, an indication of a field used by the data processing system, access the metadata repository, and output an indication of a program of the plurality of programs that introduce a dependency affecting the indicated field, the at least one non-transitory computer-readable storage medium of claim 15. **Claim 20** The data processing system further includes a plurality of data sources, the processor-executable instructions cause the at least one computer hardware processor to perform a conversion on data within the plurality of data sources, the plurality of data sources being heterogeneous, the at least one non-transitory computer-readable storage medium of claim 15. **Claim 21** The plurality of data sources comprises at least one data source that is at least one of an ORACLE database or a SAP database, or a dataset stored in a HADOOP distributed file system, the at least one non-transitory computer-readable storage medium of claim 20.

Citation Information

Patent Citations

  • Extraction device for information from parsing tree

    JP1995239787A

  • Program structure recovery using multiple languages

    US20110302563A1