A method and system for generating FlinkSQL field lineage
By custom parsing the DDL and DML process of Flink SQL, Flink SQL field blood ties are generated, which solves the problem of not being able to generate field blood ties in the existing technology, and realizes dependency analysis between fields and desensitization of sensitive data.
Patent Information
- Application Number
- CN202111603842.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-12-24
AI Technical Summary
Existing metadata management methods cannot generate Flink SQL fields, resulting in the inability to clearly understand the dependence between the source table field and the final target table field.
By customizing the DDL and DML parsing process for Flink SQL, parsing SQL statements to obtain library names, table names and field names, using linked list arrays and filtering to select field arrays, combining upward recursive merging algorithms and reverse breadth-first methods to generate blood relationships between fields.
It realizes the influencing range analysis of the table, conveniently performs automatic pull-up of task dependence and field-level desensitization of sensitive data, and clearly describes the blood relationship between fields.
Smart Images

Figure CN114238416B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method and system for generating FlinkSQL field lineage. Background Art
[0002] In big data technology, real-time data warehouse technology has been widely applied in enterprises, and major enterprises and institutions have established or explored real-time data warehouses that conform to their own business scenarios. With the increasing requirements for real-time, Apache Flink has increasingly become the first choice for many enterprises to build real-time data warehouses; among them, the use of Flink SQL is increasing. Therefore, the metadata management based on Flink SQL becomes more and more important, especially the field-level lineage relationship, clearly knowing the corresponding final target table fields of the source table fields. Through the field-level lineage relationship, it is possible to conveniently analyze the influence range of the table, automatically pull up task dependencies, and implement field-level desensitization for sensitive data.
[0003] Currently, the management method for metadata generated based on Flink SQL cannot generate Flink SQL field lineage.
[0004] In view of this, this application is specifically proposed. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that currently, the management method for metadata generated based on Flink SQL cannot generate Flink SQL field lineage. The purpose is to provide a method and system for generating FlinkSQL field lineage. By customizing the DDL and DML parsing processes for Flink SQL, parsing the SQL statement to obtain the database name, table name, and field name, generating an array containing all the tables in the SQL, then sequentially traversing the array to obtain the relationship between fields, and finally obtaining the Flink SQL field lineage in a way that shows the relationship between fields.
[0006] The present invention is realized through the following technical solutions:
[0007] On the one hand, the present invention provides a method for generating FlinkSQL field lineage, including the following steps:
[0008] S1: Define a linked list array and a filtered selection field array;
[0009] S2: Parse the SQL statement to obtain multiple column names, multiple filtered condition column names, and multiple table names;
[0010] S3: For each column name, store the column name and the mapping relationship between the column name and the table name into a linked list array, and multiple column names correspond to multiple linked list arrays;
[0011] S4: For each filtered condition column name, store the filtered condition column name and the mapping relationship between the filtered condition column name and the table name into a filtered selection field array. Multiple filtered condition column names correspond to multiple filtered selection field arrays.
[0012] S5: Define a global linked list array and a global filtered selection field array.
[0013] S6: Solve for the mapping relationship between the column name and the table name for the multiple linked list data of each linked list array, and update the global linked list array.
[0014] S7: Process each filtered selection field array to obtain the mapping relationship between the filtered condition column name and the table name, and update the global filtered selection field array.
[0015] S8: Merge the updated global linked list array and the global filtered selection field array and then find the difference.
[0016] S9: Display the data obtained after finding the difference, and obtain the FlinkSQL field lineage based on the display result.
[0017] In view of the specific data structure and storage method of column names and table names in Flink SQL, the present invention customizes the DDL and DML parsing processes for Flink SQL, parses the SQL statement to obtain the database name, table name, and field name, and can easily obtain the field-level lineage relationship through the custom data structure, so as to realize the dependency relationship between the underlying source table fields and the final target table fields. Secondly, the upward recursive merging algorithm and the reverse breadth-first method are respectively adopted to process the column fields and the filtered selection fields to obtain the final field-level lineage result. Finally, for the hierarchical influence relationship between fields, the filtered fields do not directly affect the final fields, which is more conducive to describing the lineage relationship between fields.
[0018] By using the method and system for generating FlinkSQL field lineage provided by the present invention, it is possible to conveniently perform an impact scope analysis on the table, realize the automatic pulling of task dependencies, and desensitize at the field level for sensitive data.
[0019] As a further description of the present invention, the specific content of S3 is: for each column name, perform the following steps to obtain multiple linked list arrays:
[0020] S3.1: Write the mapping relationship between the column name and the table name as the key value into a temporary Map array in the form of "column name % table name".
[0021] S3.2: If the column name has an alias, write the alias as the value into the said Map array; if the column name has no alias, write the column name as the value into the said Map array.
[0022] S3.3: Add the Map array storing the key value and value to a linked list array.
[0023] As a further description of the present invention, the said S4 is specifically: for each filtering condition column name, perform the following steps to obtain multiple filtering selection field arrays:
[0024] S4.1: Add the mapping relationship between the filtering condition column name and the table name in the form of "filtering condition column name % table name" as the key value to the filtering selection field array.
[0025] S4.2: If in the SQL statement, the right side of the "=" in the filtering condition statement is a SQL subquery statement, assign the value of the filtering selection field array to the FlinkTable object; if in the SQL statement, the right side of the "=" in the filtering condition statement is not a SQL subquery statement, assign the value of the filtering selection field array to Null.
[0026] As a further description of the present invention, after the said S4, perform the following steps:
[0027] For each table name, merge the linked list array corresponding to the table name and the filtering selection field array into a FlinkTable object, and multiple table names correspond to multiple FlinkTable objects.
[0028] Write the said multiple FlinkTable objects into the ListTable array.
[0029] As a further description of the present invention, after the said S5, perform the following steps:
[0030] Traverse the ListTable array and take out the said multiple FlinkTable objects.
[0031] Judge whether the linked list arrays in each FlinkTable object are empty; if the linked list arrays in all FlinkTable objects are empty, do not execute the said S6 to S8 anymore; otherwise, execute the said S6.
[0032] As a further description of the present invention, the said S6 is specifically: solve the multiple linked list arrays by using the upward recursive merging algorithm, including:
[0033] S6.1: Take out the multiple linked list arrays corresponding to the multiple FlinkTable objects.
[0034] S6.2: Define a local linked list array;
[0035] S6.3: Iteratively traverse multiple linked list arrays, and take out the linked list arrays; if the local linked list array is empty, assign the taken-out linked list array to the local linked list array; if the local linked list array is not empty, execute S6.4;
[0036] S6.4: Traverse the local linked list array in S6.3, and obtain multiple key values and multiple value values in the linked list array; separate two adjacent key values among the multiple key values with "%" to form a ListString array;
[0037] S6.5: Define a local ListString array;
[0038] S6.6: Traverse the ListString array in S6.4 in sequence, and determine whether each key value in the ListString array is included in the local linked list array; if so, add the key value to the local ListString array, otherwise return to S6.4;
[0039] S6.7: Determine whether the local ListString array is empty; if it is empty, update the local linked list array with the multiple key values in S6.4 as keys and Null as values; otherwise, update the local linked list array with the multiple key values in S6.4 as keys and the local ListString array as values;
[0040] S6.8: Return to S6.3 until the key values in each linked list array are obtained.
[0041] As a further description of the present invention, the specific content of S7 is as follows: Obtain the mapping relationship between the filtered condition column names and table names for the multiple filtered selection field arrays by using the reverse breadth-first method, including:
[0042] S7.1: Take out the multiple linked list arrays corresponding to multiple FlinkTable objects, and reverse the order of the linked list arrays to obtain a reverse-ordered array;
[0043] S7.2: Iteratively traverse the reverse-ordered array, and take out the linked list arrays in the reverse-ordered array;
[0044] S7.3: Layer by layer traverse the filtered selection field arrays, and take out the key values and value values in the filtered selection field arrays; if the value value in the filtered selection field array is empty, perform the next layer of traversal; otherwise, execute S7.4;
[0045] S7.4: Determine whether the value retrieved in S7.3 exists in the filtered selection field array described in S7.2; if so, store the value retrieved in S7.3 in a temporary array;
[0046] S7.5: Retrieve the temporary array and determine whether the temporary array is empty; if not, update the filtered selection field array in the FlinkTable object with the value in the temporary array;
[0047] S7.6: Obtain the value in each filtered selection field array according to the method in S7.2 to S7.6.
[0048] As a further description of the present invention, before S8, filter the fields in the global linked list array with empty values.
[0049] As a further description of the present invention, S9 is specifically: represent the data obtained after the difference as the influence relationship of the filtered selection field array on the Key value in the linked list array, and the influence relationship of the Key value in the linked list array on the value, to obtain the FlinkSQL field blood relationship.
[0050] On the other hand, the present invention provides a system for generating FlinkSQL field blood relationship, including:
[0051] A custom module for defining a linked list array, a filtered selection field array, a global linked list array, and a global filtered selection field array;
[0052] An SQL statement parsing module for parsing an SQL statement to obtain multiple column names, multiple filtered condition column names, and multiple table names;
[0053] A data adding module for storing the column name and the mapping relationship between the column name and the table name in a linked list array, and storing the filtered condition column name and the mapping relationship between the filtered condition column name and the table name in a filtered selection field array;
[0054] A data processing module for solving multiple linked list data in each linked list array to obtain the mapping relationship between the column name and the table name, and updating the global linked list array; for processing each filtered selection field array to obtain the mapping relationship between the filtered condition column name and the table name, and updating the global filtered selection field array;
[0055] A merging and difference module for merging and taking the difference between the updated global linked list array and the global filtered selection field array;
[0056] A display module for displaying the data obtained after the difference, and obtaining the FlinkSQL field blood relationship according to the display result.
[0057] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0058] 1. A method and system for generating FlinkSQL field lineage provided by an embodiment of the present invention implement the generation of field lineage relationships in Flink SQL through a custom data structure and algorithm, thereby realizing the dependency relationship between the underlying source table fields and the final target table fields;
[0059] 2. A method and system for generating FlinkSQL field lineage provided by an embodiment of the present invention can clearly describe the relationships among the filtered selection fields, source table fields, and target fields through hierarchical display of fields;
[0060] 3. A method and system for generating FlinkSQL field lineage provided by an embodiment of the present invention can conveniently perform an impact scope analysis on tables, realize the automatic startup of task dependencies, and perform field-level desensitization for sensitive data. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0062] Figure 1 It is a schematic diagram of the SQL statement parsing process provided by an embodiment of the present invention;
[0063] Figure 2 It is a schematic diagram of the method flow for obtaining the corresponding relationship of column names and fields by using an upward recursive merging algorithm provided by an embodiment of the present invention;
[0064] Figure 3 It is a schematic diagram of the method flow for obtaining the corresponding relationship of filtered fields in a reverse breadth-first manner provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and do not limit the present invention.
[0066] Embodiment
[0067] Currently, the management method for metadata generated based on Flink SQL cannot generate Flink SQL field lineage. In response to this, this embodiment provides a method for generating Flink SQL field lineage, including the following steps:
[0068] Step 1: Obtain the complete SQL statement, parse the SQL engine through Apache Calcite, customize the DDL and DML parsing processes for Flink SQL, and obtain the database name, table name, and field name.
[0069] Step 1.1: Define a FlinkColumn object, which is used to represent the column name parsed in Flink SQL and contains the column name (name), column data type (type), and column comment (comment); define a FlinkTable type, which is used to represent the tables used in the SQL statement and contains two data types: a linked list array List<Map<String,Set <string>>>, which represents the mapping list of the field column names of the table and the filtered selection field array Map<String, FlinkTable>, represents the filtered selection fields of the table.
[0070] Both the linked list array and the filtered selection field array described below are each represented by List<Map<String, Set <string>>> and Map<String, FlinkTable>.
[0071] Step 1.2: Parse the complete SQL statement obtained in Step 1.1 to obtain column names; write the mapping relationship between the column name and the table name into a temporary Map array as the key value in the form of "column name % table name"; if the column name has an alias, write the alias as the value into the Map array; if the column name does not have an alias, write the column name as the value into the Map array; add the Map array storing the key value and the value to a List<Map<String, Set <string>>>.
[0072] Step 1.3: Parse the complete SQL statement obtained in Step 1.1 to obtain the filtered condition column names; add the mapping relationship between the filtered condition column names and the table names to the Map<String, FlinkTable> in the form of "filtered condition column name % table name"; if the right side of the "=" in the filtered condition statement in the SQL statement is a SQL subquery statement, assign the value of the Map<String, FlinkTable> to a FlinkTable object; if the right side of the "=" in the filtered condition statement in the SQL statement is not a SQL subquery statement, assign the value of the Map<String, FlinkTable> to Null.
[0073] The method flow for parsing the SQL statement is as Figure 1 shown.
[0074] Step 1.4: The List<Map<String, Set corresponding to the table name <string>>> Merge with Map<String, FlinkTable> into a FlinkTable object, where multiple table names correspond to multiple FlinkTable objects; write the multiple FlinkTable objects into a ListTable array.
[0075] Step 2: Based on the ListTable array containing all the tables in the SQL obtained in Step 1 above, traverse the array sequentially to obtain the blood relationship between fields.
[0076] Step 2.1: Define a global linked list array Map<String, Set <string>> object, used to store the field correspondence results of all tables in step 1; define the global filtered selection field array Set <string>The object is used to store the filtering selection field results of all tables; define a List<Map<String,Set <string>>> object, used to assist in screening out the corresponding relationships of filtering fields.
[0077] The global linked list array and the global filtered selection field array described below are respectively represented by Map<String,Set <string>> and Set <string>representation
[0078] Step 2.2: Traverse the ListTable array obtained in Step 1 in sequence, and take out the FlinkTable object therein. Determine whether the field column name mapping list array is empty. If it is empty, jump out of the loop; otherwise, go to Step 2.3.
[0079] Step 2.3: Take out the field column name mapping list List<Map<String,Set <string>>>, traverse in order and take out the Map<String,Set <string>> Object. Then, an upward recursive merging algorithm is used to solve the problem, and the operation steps are as Figure 2 shown, including:
[0080] Step 2.31: Define a local type of Map<String, Set <string>> Object, used to store the final result of one loop and return it.
[0081] Step 2.32: Iteratively traverse the Map<String,Set retrieved in Step 2.2 <string>> Object; Merge it with the object defined in step 2.31. If the object defined in step 2.31 is empty, directly assign the traversed obtained value to the value in step 2.31; if the defined object in step 2.31 is not empty, execute step 2.33.
[0082] Step 2.33: Iteratively traverse the Map object taken out in step 2.32 to obtain its key value and value. Separate the obtained Key values with the "%" sign and convert them into a List<String> array.
[0083] Step 2.34: Define a local List<String> array to store the temporary results. Sequentially traverse the List<String> array in step 2.33, and judge whether its value is included in the Map object defined in step 2.31. If it exists, add it to the List<String> array defined in step 2.32; if it does not exist, return to step 2.33.
[0084] Step 2.35: Judge whether the content of the List<String> array defined in step 2.34 is empty. If it is empty, use the key obtained in step 2.33 as the key and null as the value, and update it in the Map object in step 2.31; if it is not empty, also use the key obtained in step 2.33 as the key and the List<String> array in step 2.34 as the value, and update the Map object in step 2.31.
[0085] Step 2.36: Continue to jump to step 2.34 for execution. If the loop is completed, return to the local object defined in step 2.31.
[0086] Step 2.4: Obtain the data result returned in step 2.3 and add or update it to the Map<String, Set> in step 2.1 <string>> into the object.
[0087] Step 2.5: Obtain the filtered field Map set in the FlinkTable object in Step 2.2; obtain the corresponding relationship of the filtered fields in a reverse breadth-first manner. The operation steps are as Figure 3 shown, including:
[0088] Step 2.51: Obtain the List<Map<String,Set defined in Step 2.1 <string>>> Object, and then reverse the array.
[0089] Step 2.52: Iteratively traverse the reversed array in Step 2.51 and take out the Map<String, Set <string>> Object.
[0090] Step 2.53: Traverse the filtered field Map set layer by layer, and take out the key value and value in the Map set. If the value is empty, jump out and proceed to the next loop; otherwise, go to Step 2.54;
[0091] Step 2.54: Take out the Map<String, Set in the array described in Step 2.51 <string>> Determine whether the value in the type exists in the value in Step 2.53; if it exists, store it in a temporary array.
[0092] Step 2.55: Take out the temporary array in Step 2.54, and if it is not empty, update the filtered field Map set in the FllinkTable object obtained in Step 2.2.
[0093] Step 2.56: Determine whether the loop is completed and jump out to return the final result; otherwise, go to Step 2.54.
[0094] Step 2.6: Obtain the result in Step 2.5 and add or update it to the Set in Step 2.1 <string>In the object.
[0095] Step 2.7: Filter out the Map<String, Set in Step 2.1 <string>Among the objects of >, the fields with empty value values.
[0096] Step 2.8: Merge the Map<String,Set in Step 2.7 <string>> Object and Set <string>Object; retrieve Map<String, Set <string>The key value set of the object > and Set <string>The difference of the object sets, and the result obtained will be only in Set <string>Some data inside.
[0097] Step 2.9: After the combined operation in Step 2.7, a field set Map<String, Set will be obtained from the source table to the target table <string>>. Among them, the key value is composed of the column names and table names of the source table combined according to "%"; the value is a set representing the field set of the corresponding target table; a set Set containing all filtered and selected fields <string>, where the value is composed of the source table filtering fields and the table name combined according to "%".
[0098] Step 2.10: Display of the mutual relationship between the final fields. Represent it as a set of filtering fields Set <string>Affect Map<String, Set <string>> The key values in the set, and then the key values affect the value values. In this way, the lineage relationship at the entire field level is obtained.
[0099] Correspondingly, this embodiment also provides a system for generating FlinkSQL field lineage, which is used to execute the method for generating FlinkSQL field lineage described above. The system includes:
[0100] A custom module for defining a linked list array, a filtered selection field array, a global linked list array, and a global filtered selection field array;
[0101] An SQL statement parsing module for parsing the SQL statement to obtain multiple column names, multiple filtered condition column names, and multiple table names;
[0102] A data addition module for storing the column names and the mapping relationship between the column names and the table names into a linked list array, and storing the filtered condition column names and the mapping relationship between the filtered condition column names and the table names into a filtered selection field array;
[0103] A data processing module for solving the multiple linked list data of each linked list array to obtain the mapping relationship between the column names and the table names, and updating the global linked list array; for processing each filtered selection field array to obtain the mapping relationship between the filtered condition column names and the table names, and updating the global filtered selection field array;
[0104] A merging and difference-finding module for merging and then finding the difference between the updated global linked list array and the global filtered selection field array;
[0105] A display module for displaying the data obtained after finding the difference, and obtaining the FlinkSQL field lineage according to the display result.
[0106] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.< / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string> < / string>
Claims
1. A method for generating FlinkSQL field lineage, characterized in that, It includes the following steps: S1: Define an array of linked lists and an array of filtered selection fields; S2: Parse the SQL statement to obtain multiple column names, multiple filtered condition column names, and multiple table names; S3: For each column name, store the column name and the mapping relationship between the column name and the table name in an array of linked lists, and multiple column names correspond to multiple arrays of linked lists; S4: For each filtered condition column name, store the filtered condition column name and the mapping relationship between the filtered condition column name and the table name in an array of filtered selection fields, and multiple filtered condition column names correspond to multiple arrays of filtered selection fields; S5: Define a global array of linked lists and a global array of filtered selection fields; S6: Use the upward recursive merging algorithm to solve multiple arrays of linked lists to obtain the mapping relationship between column names and table names, and update the global array of linked lists; S7: Use the reverse breadth-first method for the multiple arrays of filtered selection fields to obtain the mapping relationship between the filtered condition column names and the table names, and update the global array of filtered selection fields; S8: Merge and then find the difference between the updated global array of linked lists and the global array of filtered selection fields; S9: Display the data obtained after finding the difference, and obtain the FlinkSQL field lineage based on the display result.
2. The method for generating FlinkSQL field lineage according to claim 1, characterized in that, The specific content of S3 is: For each column name, execute the following steps to obtain multiple arrays of linked lists: S3.1: Write the mapping relationship between the column name and the table name in the form of "column name % table name" as the key value into a temporary Map array; S3.2: If the column name has an alias, write the alias as the value into the Map array; if the column name does not have an alias, write the column name as the value into the Map array; S3.3: Add the Map array storing the key value and the value to an array of linked lists.
3. The method for generating FlinkSQL field lineage according to claim 1, characterized in that, The specific content of S4 is: For each filtered condition column name, execute the following steps to obtain multiple arrays of filtered selection fields: S4.1: Add the mapping relationship between the filtered condition column name and the table name in the form of "filtered condition column name % table name" as the key value to the array of filtered selection fields; S4.2: If the right side of the "=" in the filtered condition statement in the SQL statement is a SQL subquery statement, assign the value of the array of filtered selection fields to the FlinkTable object; if the right side of the "=" in the filtered condition statement in the SQL statement is not a SQL subquery statement, assign the value of the array of filtered selection fields to Null.
4. The method for generating FlinkSQL field lineage according to claim 1, characterized in that, After S4, execute the following steps: For each table name, merge the array of linked lists and the array of filtered selection fields corresponding to the table name into a FlinkTable object, and multiple table names correspond to multiple FlinkTable objects; write the multiple FlinkTable objects into the ListTable array.
5. The method for generating FlinkSQL field lineage according to claim 4, characterized in that, After S5, execute the following steps: Traverse the ListTable array and take out the multiple FlinkTable objects; Determine whether the linked list arrays in each FlinkTable object are empty; if the linked list arrays in all FlinkTable objects are empty, then do not execute the steps from S6 to S8; otherwise, execute S6.
6. The method for generating FlinkSQL field lineage according to claim 1, 4 or 5, characterized in that, The method of using the upward recursive merging algorithm to solve multiple linked list arrays to obtain the mapping relationship between column names and table names and update the global linked list array includes: S6.1: Take out the multiple linked list arrays corresponding to multiple FlinkTable objects; S6.2: Define a local linked list array; S6.3: Iteratively traverse the multiple linked list arrays and take out a linked list array; if the local linked list array is empty, assign the taken-out linked list array to the local linked list array; if the local linked list array is not empty, then execute S6.4; S6.4: Traverse the local linked list array in S6.3 to obtain multiple key values and multiple value values in the linked list array; separate two adjacent key values among the multiple key values with "%", forming a ListString array; S6.5: Define a local ListString array; S6.6: Sequentially traverse the ListString array in S6.4 and determine whether each key value in the ListString array is included in the local linked list array; if so, add the key value to the local ListString array, otherwise return to S6.4; S6.7: Determine whether the local ListString array is empty; if it is empty, update the local linked list array with the multiple key values in S6.4 as keys and Null as values; otherwise, update the local linked list array with the multiple key values in S6.4 as keys and the local ListString array as values; S6.8: Return to S6.3 until the key values in each linked list array are obtained.
7. A method for generating FlinkSQL field lineage according to claim 1, 4 or 5, characterized in that The method of using the reverse breadth-first method for the multiple filtered selection field arrays to obtain the mapping relationship between the filtered condition column names and table names and update the global filtered selection field array includes: S7.1: Take out the multiple linked list arrays corresponding to multiple FlinkTable objects and reverse the order of the linked list arrays to obtain a reverse array; S7.2: Iteratively traverse the reverse array and take out the linked list arrays in the reverse array; S7.3: Layer-by-layer traverse the filtered selection field array and take out the key values and value values in the filtered selection field array; if the value value in the filtered selection field array is empty, then perform the next layer of traversal; otherwise, execute S7.4; S7.4: Determine whether the value value taken out in S7.3 exists in the filtered selection field array described in S7.2; if so, store the value value taken out in S7.3 in a temporary array; S7.5: Take out the temporary array and determine whether the temporary array is empty; if it is not empty, update the filtered selection field array in the FlinkTable object with the value values in the temporary array; S7.6: Obtain the value values in each filtered selection field array according to the method from S7.2 to S7.
6.
8. A method for generating FlinkSQL field lineage according to claim 1, characterized in that Before S8, fields with empty value in the global linked list array are filtered out.
9. A method for generating FlinkSQL field lineage according to claim 1, characterized in that Specifically, S9 is as follows: Represent the data obtained after the difference operation as the influence relationship of the filtered selection field array on the Key values in the linked list array, and the influence relationship of the Key values in the linked list array on the value values, to obtain the FlinkSQL field lineage.
10. A system for generating FlinkSQL field lineage, characterized in that It includes: A custom module for defining a linked list array, a filtered selection field array, a global linked list array, and a global filtered selection field array; An SQL statement parsing module for parsing SQL statements to obtain multiple column names, multiple filtered condition column names, and multiple table names; A data addition module for storing the column names and the mapping relationship between the column names and the table names into a linked list array, and storing the filtered condition column names and the mapping relationship between the filtered condition column names and the table names into a filtered selection field array; A data processing module for solving multiple linked list arrays using an upward recursive merging algorithm to obtain the mapping relationship between the column names and the table names, and updating the global linked list array; for obtaining the mapping relationship between the filtered condition column names and the table names from the multiple filtered selection field arrays using a reverse breadth-first method, and updating the global filtered selection field array; A merging and difference module for merging and taking the difference between the updated global linked list array and the global filtered selection field array; A display module for displaying the data obtained after the difference operation, and obtaining the FlinkSQL field lineage based on the display result.
Citation Information
Patent Citations
Data traceability tool construction method, data processing method, device and equipment
CN113434533A
Blood relationship representation method based on Elastic Search
CN113590610A