Data platform metadata automatic generation method based on large model

By standardizing and statically parsing SQL code, and combining it with deep semantic understanding of large models, the problem of low metadata generation efficiency in complex data platforms by traditional methods is solved, and efficient and accurate data lineage generation is achieved.

CN120873206APending Publication Date: 2025-10-31ZHEJIANG NON-LINEAR DIGITAL TECH CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202511018594.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional methods of metadata management and data lineage generation are inefficient, time-consuming, and difficult to guarantee the real-time performance and accuracy of data when dealing with large-scale and complex enterprise-level data platforms, especially when dealing with complex SQL code, where it is difficult to accurately parse the data lineage.

Method used

We adopt an automatic metadata generation method for data platforms based on large models. By standardizing and statically parsing the original SQL code, we identify the code type, perform preliminary lineage analysis, use a large language model for deep semantic understanding, integrate simple and complex lineage relationships, and generate complete data platform metadata.

Benefits of technology

It significantly improves the accuracy and completeness of metadata generation, makes up for the shortcomings of traditional parsing tools in understanding complex semantics, and improves the efficiency and accuracy of metadata generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873206A_ABST
    Figure CN120873206A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of metadata generation, and discloses a data platform metadata automatic generation method based on a large model, which comprises the following steps of: firstly, performing unified standardization and preliminary grammatical analysis on an original SQL (Structured Query Language) code to obtain a structured intermediate representation; on the basis, preliminary blood relationship analysis based on rules is carried out, and simple column references are quickly identified and processed. For complex expressions which are difficult to accurately analyze by a traditional method, code snippets and context information of the complex expressions are accurately extracted and submitted to a large language model for deep semantic understanding and complex blood relationship analysis. And finally, integrating the complex consanguinity analyzed by the large model with the initial consanguinity list to form a comprehensive and accurate field-level consanguinity, and further generating complete data platform metadata. In this way, the defect that a traditional analysis tool understands complex semantics is effectively overcome, and the accuracy and integrity of metadata generation are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of metadata generation technology, and more specifically, to a method for automatically generating metadata for a data platform based on a large model. Background Technology

[0002] As the cornerstone for carrying massive amounts of data and supporting various data applications, the management level of a data platform's metadata directly affects the maturity of data governance, the credibility of data assets, and the efficiency of data applications. Metadata, as "data on data," details the entire lifecycle of data from its generation and transformation to its consumption, including data structure, meaning, and lineage. Among these, data lineage is crucial for understanding data flow, tracing data sources, assessing data quality, conducting impact analysis, and meeting compliance requirements.

[0003] However, traditional methods of metadata management and data lineage generation face numerous challenges and limitations. Generally, early data lineage tracing relied heavily on manual analysis or simple rule matching at the table or field level. This approach proves inefficient, time-consuming, and error-prone when dealing with massive enterprise-level data platforms with complex and ever-changing data chains, making it difficult to guarantee data real-time performance and accuracy. With the evolution of architectures such as data warehouses, data lakes, and even data platforms, the complexity of SQL code in data processing flows has increased exponentially. Numerous ETL (Extract, Transform, Load) scripts, data model building statements, view definitions, and stored procedures constitute a complex network of data transformations. Summary of the Invention

[0004] To address the aforementioned technical problems, this application is proposed. Embodiments of this application propose a method for automatically generating metadata for a data platform based on a large model.

[0005] According to one aspect of this application, a method for automatically generating metadata for a data platform based on a large model is provided, comprising: obtaining an original SQL code file; standardizing the original SQL code file to obtain standardized code text and code type identifiers; based on the code type identifiers, calling a corresponding static SQL parser to perform syntax analysis on the standardized code text to obtain a structured intermediate representation of the code; performing rule-based preliminary lineage analysis and fragment extraction on the structured intermediate representation of the code to obtain a preliminary lineage list and a list of complex expression IR nodes to be analyzed and their corresponding code fragments; performing large model semantic understanding and complex lineage parsing on the list of complex expression IR nodes to be analyzed and their corresponding code fragments to obtain a complex lineage list; integrating the preliminary lineage list and the complex lineage list to obtain a structured field-level lineage relationship list; and associating the lineage relationships in the structured field-level lineage relationship list with the corresponding data tables and field source data to obtain data platform metadata.

[0006] In one possible implementation, the original SQL code file is standardized to obtain standardized code text and code type identifier, including: identifying the database type or processing framework to which the original SQL code file belongs to obtain the code type identifier; and performing case-consistent processing, removing unnecessary comments, and formatting on the original SQL code file to obtain the standardized code text.

[0007] In one possible implementation, based on the code type identifier, the corresponding static SQL parser is invoked to perform syntactic analysis on the standardized code text to obtain a structured intermediate representation of the code, including: performing lexical analysis on the standardized code text to obtain a sequence of lexical units; the static SQL parser performs syntactic analysis on the sequence of lexical units based on the grammatical rules of the target SQL dialect to obtain a code abstract syntax tree as a structured intermediate representation of the code.

[0008] In one possible implementation, the structured intermediate representation of the code undergoes rule-based preliminary lineage analysis and fragment extraction to obtain a preliminary lineage list and a list of complex expression IR nodes to be analyzed, along with their corresponding code fragments. This includes: traversing the structured intermediate representation of the code to obtain a list of code nodes; extracting the expressions of each code node in the list of code nodes to obtain a list of code node expressions; performing type judgment on each code node expression in the list of code node expressions to obtain a type judgment result; if the type judgment result is a simple column reference, generating a preliminary lineage record based on the corresponding code node expression to obtain the preliminary lineage list; if the type judgment result is not a simple column reference, marking the corresponding code node as a complex expression IR node to obtain the list of complex expression IR nodes to be analyzed and their corresponding code fragments.

[0009] Compared with existing technologies, the automatic generation method for data platform metadata based on a large model provided in this application first performs unified standardization and preliminary syntactic analysis on the original SQL code to obtain a structured intermediate representation. Based on this, rule-based preliminary lineage analysis is conducted to quickly identify and process simple column references. For complex expressions that are difficult to parse accurately using traditional methods, their code fragments and contextual information are precisely extracted and then passed to a large language model for deep semantic understanding and complex lineage parsing. Finally, the complex lineage parsed by the large model is integrated with the preliminary lineage list to form a comprehensive and accurate field-level lineage relationship, thereby generating complete data platform metadata. This effectively compensates for the shortcomings of traditional parsing tools in understanding complex semantics, significantly improving the accuracy and completeness of metadata generation. Attached Figure Description

[0010] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0011] Figure 1 The illustration shows a schematic flowchart of a method for automatically generating metadata for a large model-based data platform according to an embodiment of this application.

[0012] Figure 2 The illustration shows a schematic flowchart of step S2 in the method for automatically generating metadata for a large model-based data platform according to an embodiment of this application.

[0013] Figure 3The illustration shows a schematic flowchart of step S3 in the method for automatically generating metadata for a large model-based data platform according to an embodiment of this application.

[0014] Figure 4 The illustration shows a schematic flowchart of step S4 in the method for automatically generating metadata for a large model-based data platform according to an embodiment of this application.

[0015] Figure 5 The illustration shows a schematic flowchart of step S5 in the method for automatically generating metadata for a large model-based data platform according to an embodiment of this application.

[0016] Figure 6 The illustration shows a schematic flowchart of step S53 in the method for automatically generating metadata for a large model-based data platform according to an embodiment of this application. Detailed Implementation

[0017] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.

[0018] Figure 1 The illustration shows a schematic flowchart of a method for automatically generating metadata for a large-model data platform according to an embodiment of this application. Figure 1 As shown, this application provides a method for automatically generating metadata for a data platform based on a large model, including: S1: obtaining the original SQL code file; S2: standardizing the original SQL code file to obtain standardized code text and code type identifiers; S3: based on the code type identifiers, calling the corresponding static SQL parser to perform syntax analysis on the standardized code text to obtain a structured intermediate representation of the code; S4: performing rule-based preliminary lineage analysis and fragment extraction on the structured intermediate representation of the code to obtain a preliminary lineage list and a list of complex expression IR nodes to be analyzed and their corresponding code fragments; S5: performing large model semantic understanding and complex lineage parsing on the list of complex expression IR nodes to be analyzed and their corresponding code fragments to obtain a complex lineage list; S6: integrating the preliminary lineage list and the complex lineage list to obtain a structured field-level lineage relationship list; S7: associating the lineage relationships in the structured field-level lineage relationship list with the corresponding data tables and field source data to obtain data platform metadata.

[0019] For example, in step S1, the raw SQL code files are obtained. It should be understood that in any mature data platform, data processing, cleaning, aggregation, and model building all rely on the support of the SQL language. Whether it's the transformation logic defined in ETL (Extract, Transform, Load) tools, the creation of views and stored procedures in a data warehouse, or the data model building statements used in data lake and data platform projects, their core carriers are various SQL scripts. These SQL code files represent the precise definition and execution path of data from its source to its final destination, undergoing layers of evolution and computation. They record in detail the source of fields, calculation rules, transformation logic, and relationships between data; these are key elements in constructing accurate data lineage and metadata. Without these raw SQL codes as input, subsequent standardization processing, syntax parsing, lineage analysis, and even large-scale model semantic understanding will lose their data foundation, and the entire metadata generation method cannot be effectively implemented. Therefore, the first step is to obtain these raw SQL code files containing rich data transformation information.

[0020] Specifically, the methods for obtaining the raw SQL code files include: First, they can be obtained from the version control system of the data platform. Most standard data development processes incorporate SQL scripts into version control systems such as Git and SVN. By integrating with the APIs of these systems, the latest SQL code files can be retrieved periodically or on demand, ensuring that the most authoritative and complete codebase in the current production environment is obtained. Second, they can be extracted directly from the metadata storage of the database, data warehouse, or data lake. For example, view definitions, stored procedures, functions, etc., stored in the database exist in the form of SQL code. By connecting to the system tables or metadata views of the database, these SQL definition texts can be queried and exported. Third, they can be exported from ETL tools or data integration platforms. Many enterprises use dedicated ETL tools or data integration platforms to orchestrate data flows. These tools often provide the ability to export their internally defined JOBs or transformation scripts, which contain a large amount of SQL code. By parsing these exported files, the embedded SQL logic can be extracted. Fourth, they can be obtained through file system scanning or specific directory monitoring. In some scenarios, developers may store SQL script files in specific file servers or shared directories. At this point, a script can be written to traverse and identify the specified file directory, filtering out all SQL code files with suffixes such as .sql and .ddl for collection.

[0021] For example, in step S2, the original SQL code file is standardized to obtain standardized code text and code type identifier. It should be understood that the original SQL code file may come from different database systems, each of which may have differences in SQL dialects, resulting in slight differences in the syntax rules for the same operation across different databases. Furthermore, even within the same database system, developers may have various habits when writing SQL: for example, using different capitalization conventions; including numerous comments for explaining business logic or debugging; having non-standard formatting; and even mixing in embedded SQL from other languages, such as JDBC statements in Java code or SQL strings in Python scripts. Directly inputting these raw, non-standardized SQL code files into the subsequent parser can lead to the following problems: First, compatibility issues: different SQL dialects require different parsing rules. If the code type is not clearly defined, the parser will be unable to select the correct set of syntax rules, leading to parsing failure or errors. Second, robustness issues: comments, formatting differences, mixed capitalization, and other "noise" information are redundant or even interfering with syntax parsing. These factors significantly increase parsing complexity, reduce parser robustness, and make it more prone to crashing or producing inaccurate results when encountering non-standard formats. Third, there are efficiency issues: parsing raw text containing a large amount of redundant information slows down processing and increases unnecessary computational overhead. Fourth, there are large model understanding biases: while large language models have powerful semantic understanding capabilities, their core is based on probability and pattern learning. If the input data format is highly inconsistent, the model may introduce unnecessary complexity or even biases when understanding semantics, affecting the accuracy of lineage inference. Therefore, in this application's technical solution, the original SQL code file is standardized to obtain standardized code text and code type identifiers.

[0022] In one embodiment, such as Figure 2 As shown, the original SQL code file is standardized to obtain standardized code text and code type identifier, including: S21: identifying the database type or processing framework to which the original SQL code file belongs to obtain the code type identifier; S22: standardizing the case of the original SQL code file, removing unnecessary comments, and formatting to obtain the standardized code text.

[0023] Specifically, the original SQL code file is first analyzed to determine its SQL dialect or specific SQL processing engine type, thereby accurately identifying the code type and obtaining a code type identifier. In one specific embodiment, heuristic rules and pattern matching are used to identify the database type or processing framework to which the original SQL code file belongs. Specifically, a database dialect feature library is constructed, which contains SQL keywords, built-in functions, system table names, or syntax structures unique to various mainstream databases or processing frameworks. When a raw SQL file is received, the system scans its content and counts the frequency of different dialect features. For example, if the NVL() function, DUAL table, and CONNECT BY clause frequently appear in the file, it is likely to point to an Oracle database; if LIMIT and OFFSET, AUTO_INCREMENT keywords, or backticks are used for identifier references, it is highly likely to be MySQL or PostgreSQL; while the TOP keyword, NOLOCK hint, and GETDATE() function are commonly used in SQL Server; for SQL in a data lake, it may contain syntax elements unique to HiveQL or Spark SQL such as EXPLODE, LATERALVIEW, and ARRAY_CONTAINS. Furthermore, this application can also pre-define a feature set for each dialect and assign a weight to each matched feature. For example, discovering a strong Oracle-specific function might have a higher weight than discovering a general keyword. By accumulating the weights of the matched features, a total score is calculated for each possible dialect. Finally, the dialect with the highest score that exceeds a certain confidence threshold is selected as the recognition result, where the confidence threshold can be set according to the actual situation. If the scores of all dialects are below the threshold or the scores are close, they may be marked as general SQL or an unknown dialect and submitted to a more general parser for testing, or the user may be prompted for manual confirmation.

[0024] After identifying the code type identifier, specifically, the original SQL code file undergoes case unification, unnecessary comments are removed, and formatting is performed to obtain the standardized code text. In one specific embodiment, lexical analysis preprocessing is first performed, using a lightweight lexical analyzer to decompose the SQL text into lexical units. During this process, all comments can be directly identified and removed, as comments have no impact on SQL execution logic and lineage. Simultaneously, strings and numeric literals are processed to ensure they are not mistakenly affected by subsequent standardization steps. Next, case unification is performed. While most SQL databases are case-insensitive when processing keywords, case unification by the parser improves consistency. Typically, keywords and built-in functions are converted to a unified case, while table names and column names retain their original case or are converted according to the database's actual case sensitivity settings. Then, redundant whitespace characters and line breaks are removed. Multiple spaces, tabs, and line breaks in an SQL statement are often syntactically equivalent to a single space. Standardization compresses these into a single space, thereby reducing text length and simplifying subsequent lexical and syntax analysis. Finally, formatting and consistency processing are performed. It's understandable that while standardization aims to remove redundancy, light formatting can increase consistency. For example, ensuring operators have leading and trailing spaces, or standardizing the use of statement terminators (such as semicolons). This benefits the robustness of subsequent static parsers.

[0025] For example, in step S3, based on the code type identifier, the corresponding static SQL parser is invoked to perform syntactic analysis on the standardized code text to obtain a structured intermediate representation of the code. It should be understood that the original SQL code, even after standardization, is essentially still a linear sequence of characters. Although it conforms to certain human-defined grammatical rules, directly processing such linear text is extremely inefficient and prone to errors for computer programs to deeply understand and automate their analysis. This raw text lacks an inherent hierarchical structure and semantic relationships, making it difficult to directly reveal the relationships between the various components of an SQL statement, such as which columns are selected, which tables are joined, which conditions are filtered, and which expressions are calculated fields. Therefore, further, the corresponding static SQL parser is invoked to perform syntactic analysis on the standardized code text, thereby transforming this flat character sequence into a more efficient and expressive structured intermediate representation according to specific SQL dialect grammatical rules.

[0026] In one embodiment, such as Figure 3As shown, based on the code type identifier, the corresponding static SQL parser is invoked to perform syntactic analysis on the standardized code text to obtain a structured intermediate representation of the code, including: S31: performing lexical analysis on the standardized code text to obtain a lexical unit sequence; S32: the static SQL parser performs syntactic analysis on the lexical unit sequence based on the grammatical rules of the target SQL dialect to obtain a code abstract syntax tree as a structured intermediate representation of the code.

[0027] First, in the lexical analysis phase, the standardized code text is read and grouped character by character, decomposing it into a series of smallest units with independent meaning, namely, a sequence of lexical units. This process is similar to breaking down a natural language article into individual words and punctuation marks. The lexical analyzer identifies various elements in the SQL language, such as keywords, identifiers, operators, numeric literals, string literals, delimiters, and parentheses. While identifying these lexical units, the lexical analyzer also assigns a type to each unit and records its corresponding text value.

[0028] Next comes the syntax analysis phase, the core of the static SQL parser. In this phase, the static SQL parser processes the sequence of lexical units generated by the lexical analyzer based on the grammatical rules of the target SQL dialect, and constructs a code abstract syntax tree (CAST) as a structured intermediate representation of the code. The code type identifier identified in the previous stages plays a decisive role here, indicating which specific dialect's grammatical rules the system should load and apply. For example, if the code type identifier is MySQL, the parser will load the MySQL SQL syntax specification; if it's Oracle, it will load the Oracle syntax specification. This process can be visualized as assembling lexical units into meaningful sentence structures according to grammatical rules.

[0029] Specifically, the parser maintains a state machine internally and identifies the validity of lexical unit sequences based on predefined context-free grammar rules. For example, a simple SELECT statement might have the following syntax: select_statement = SELECT selected_elements FROM table_expression. When the parser sees the SELECT keyword, it expects to see a series of selected_elements, then the FROM keyword, followed by table_expression. If the lexical unit sequence strictly conforms to these rules, the parser constructs a code abstraction syntax tree in memory accordingly and uses it as a structured intermediate representation of the code.

[0030] For example, in step S4, a rule-based preliminary lineage analysis and fragment extraction are performed on the structured intermediate representation of the code to obtain a preliminary lineage list and a list of complex expression IR nodes to be analyzed, along with their corresponding code fragments. It should be understood that SQL code is massive and structurally diverse. A significant portion of the lineage relationships are direct and explicit, such as simply selecting a field in a table or directly assigning a value from one field to another. The syntactic structures corresponding to these lineage relationships are highly standardized and can be quickly and accurately identified and extracted using preset syntactic rules and pattern matching on the already constructed abstract syntax tree, i.e., the structured intermediate representation of the code. Through this rule-based preliminary lineage analysis, a preliminary lineage list can be quickly generated. This lineage is highly accurate, and the processing cost is significantly lower than calling a large language model. This strategy avoids the necessity of submitting all SQL expressions, regardless of their complexity, to a computationally intensive large model for processing, thereby significantly reducing overall computational overhead and response time.

[0031] In one embodiment, such as Figure 4 As shown, the structured intermediate representation of the code undergoes rule-based preliminary lineage analysis and fragment extraction to obtain a preliminary lineage list and a list of complex expression IR nodes to be analyzed and their corresponding code fragments. This includes: S41: Traversing the structured intermediate representation of the code to obtain a list of code nodes; S42: Extracting the expressions of each code node in the code node list to obtain a list of code node expressions; S43: Performing type judgment on each code node expression in the code node expression list to obtain a type judgment result; S44: If the type judgment result is a simple column reference, generating a preliminary lineage record based on the corresponding code node expression to obtain the preliminary lineage list; S45: If the type judgment result is not a simple column reference, marking the corresponding code node as a complex expression IR node to obtain the list of complex expression IR nodes to be analyzed and their corresponding code fragments.

[0032] Specifically, the process begins by traversing the structured intermediate representation of the code to obtain a list of code nodes. After obtaining the abstract syntax tree (AST), the structured intermediate representation of the SQL statement, a traversal algorithm is initiated, typically using depth-first search (DFS) or breadth-first search (BFT). The purpose of the traversal is to systematically access every node in the abstract syntax tree, as each node represents a syntactic construct or expression in the SQL statement. During the traversal, all nodes considered potentially containing lineage information, especially those related to various expressions, are collected to form an ordered list of code nodes.

[0033] Next, the expressions of each code node in the code node list are extracted to obtain a code node expression list. This step involves extracting the logical structure or textual content of the specific expression represented by each node in the code node list obtained in the previous step from its code abstract syntax tree. For example, if a node in the structured intermediate representation of the code represents the direct name of a field, a simple addition operation, or a function call, this information will be extracted and encapsulated. Importantly, what is extracted here is the semantic carrier of the expression, not just the string text, to facilitate subsequent structured type determination. These extracted expression entities together form the code node expression list, serving as the direct input for subsequent type determination.

[0034] Then, type determination is performed on each code node expression in the code node expression list to obtain the type determination result. This step is crucial for preliminary lineage analysis and complex expression identification. It categorizes the complexity of each expression according to preset, strictly defined rules. The implementation depth of type determination relies on the accurate identification of SQL language syntax structure and code abstract syntax tree node types. Specifically, the core criterion for determining whether an expression is a simple column reference is whether it directly points to a column located in the query source (e.g., a table or view in the FROM clause) without any complex transformations such as arithmetic calculations, logical operations, function calls, or conditional judgments during the reference process.

[0035] In a specific implementation, firstly, type checking is performed on the code abstraction syntax tree (CAST) nodes. In typical CAST designs, nodes representing column references have a specific type marker, such as a column reference node. When such nodes are encountered, a preliminary judgment is made. Secondly, a source tracing check is performed. For an expression identified as a column reference node, the parser traces its parent and ancestor nodes in the CAST to confirm whether it directly originates from a field in a recognizable table or view. This means that the column reference cannot be an operand of any arithmetic operation or function call, but rather an element directly selected from the data source. For example, directly selecting a product name field in a query, which directly comes from the product table, is a simple column reference.

[0036] A more crucial criterion is the absence of calculations or nested functions. Even if a node is a column reference, if it is included in any arithmetic expression (e.g., adding the column's value to another number), logical expression (e.g., the column's value participating in a Boolean comparison), or passed as an argument to a function (e.g., string truncation or date formatting of the column's value), then the expression containing that column reference is considered complex. Therefore, a simple column reference means that the column is directly selected or assigned a value without any semantic transformation at the expression level. For cases using aliases, such as renaming a total amount field to transaction total in a query, as long as its underlying source is a single and direct field, it is still considered a simple column reference because the alias itself does not introduce complex semantic calculations.

[0037] Any expression that does not conform to the above criteria for simple column references is considered a complex expression. This judgment is based on its internal structure and semantic characteristics, typically including one or more of the following features: 1. Function calls: If the expression contains any built-in SQL functions (e.g., functions for date and time processing, string manipulation, mathematical calculations) or user-defined functions (UDFs). For example, concatenating two string fields or summing a numeric field. 2. Arithmetic or logical operations: If the expression involves two or more operands combined using arithmetic operators (such as addition, subtraction, multiplication, and division) or logical operators (such as AND, OR, and NOT). For example, calculating the difference between two numeric fields or determining whether two conditions are simultaneously met. 3. Conditional expressions: If the expression contains explicit conditional branching structures, such as determining different output results or text descriptions based on the value of a field. 4. Subqueries or nested queries: If the evaluation result of the expression depends on the return of a nested subquery. 5. Aggregate functions: If the expression is an aggregate function that performs summary calculations on a set of data, even if its parameters are simple column references, aggregation itself is a complex transformation. For example, calculating the average or maximum value of a set of numeric fields. VI. Type Conversion: If the expression contains explicit or implicit type conversion operations. VII. Bitwise Operators or Other Complex Operators: Other less common, indirect syntax operators.

[0038] During implementation, each expression node in the code node expression list is traversed and recursively checked. For example, for a code abstract syntax tree node representing a complex statement, the system checks its child nodes. If a node's type explicitly belongs to a function call node, binary operation node, conditional expression node, or subquery node, it is directly marked as a complex expression. At a deeper level, even if a node representing a column reference seems simple, if its parent or ancestor node explicitly indicates that it is included in any of the aforementioned complex operations, then the entire expression using that column reference as an operand or parameter is considered a complex expression. For example, performing addition on a field; the field itself is a simple column, but the whole expression "the field plus another value" is a complex expression.

[0039] Depending on the type determination result, different actions are taken: If the type determination result is a simple column reference, a preliminary lineage record is generated based on the corresponding code node expression to obtain the preliminary lineage list; for expressions identified as simple column references, the system directly extracts their source code table and column names from the code abstract syntax tree (or maps them to the original table columns via aliases) and records them. For example, selecting a field in a query and giving it an alias will result in a lineage record containing the target alias and its source original table and column. These records are aggregated into a preliminary lineage list, which is a deterministic and efficient source of lineage.

[0040] If the type determination result is not a simple column reference, the corresponding code node is marked as a complex expression IR node to obtain a list of complex expression IR nodes to be analyzed and their corresponding code fragments: for nodes identified as complex expressions, they themselves (i.e., the corresponding node object in the code abstract syntax tree as their internal representation) are added to the list of complex expression IR nodes to be analyzed. Simultaneously, a crucial operation is to accurately extract the corresponding code fragment from the original standardized SQL code text for this complex expression. This extraction must be clear in boundaries and accurate in content, ensuring the completeness of the fragment. For example, if a complex conditional judgment runs throughout the entire statement, the entire text of the conditional judgment will be extracted. Furthermore, the system will collect fragment context information related to this fragment, such as the fragment's position in the original SQL, the target output name in its SELECT list (if any aliases exist), the mapping relationship between table aliases and actual table names involved in the query, and detailed metadata information (such as the original table, column name, data type, etc.) of any source fields referenced internally by the complex expression.

[0041] For example, in step S5, the list of IR nodes of the complex expression to be analyzed and its corresponding code fragments are subjected to large-model semantic understanding and complex lineage parsing to obtain a complex lineage list. It should be understood that although static SQL parsers can accurately identify the syntax structure of SQL and establish lineage relationships for simple column selection, filtering, or join operations, their capabilities are insufficient when faced with expressions involving deep business logic, semantic transformation, or black-box nature. For example, when multiple nested function calls appear in the SQL code, and the time information of one field undergoes multiple transformations, such as numerical adjustment, formatting, and finally date truncation, the semantic chain of this complex transformation is beyond the understanding of traditional parsers; they cannot discern the true flow of the original lineage of the data after layers of function calls. Another example is complex conditional logic, where the value of one output field depends on complex Boolean judgments between multiple input fields. For instance, a customer's level depends on multiple conditions such as their spending amount, activity level, and participation in specific activities. Traditional tools can only identify the syntax of the conditional statements but cannot intelligently infer the specific depth of association between the final result and these conditional fields. Alternatively, aggregate functions may contain complex expressions, calculating the sum of sub-expressions after conditional filtering and transformation. Tracing their lineage requires a deep understanding of the semantics of internal conditional statements and the filtering logic of their internal fields. Even user-defined function or stored procedure calls, such as those calculating customer risk scores, are black boxes of business logic. Traditional parsers cannot penetrate their internal workings to understand their transformation relationships with input fields, nor can they infer the source of the final score's fields. Consequently, the lineage relationships generated by traditional methods are often incomplete, inaccurate, or even completely missing in complex scenarios. Therefore, this paper further performs large-scale model semantic understanding and complex lineage parsing on the IR node list of the complex expressions to be analyzed and their corresponding code snippets to obtain a complex lineage list.

[0042] In one embodiment, such as Figure 5 As shown, the method involves performing large-scale model semantic understanding and complex lineage parsing on the list of IR nodes of the complex expression to be analyzed and its corresponding code fragments to obtain a complex lineage list, including: S51: extracting a first code fragment to be analyzed from the list of IR nodes of the complex expression to be analyzed and its corresponding code fragments; S52: extracting the fragment context information of the first code fragment to be analyzed; S53: generating a complex lineage relationship question Propmt based on the first code fragment to be analyzed and the fragment context information; S54: inputting the complex lineage relationship question Propmt into the large language model to output a complex lineage record corresponding to the first code fragment to be analyzed.

[0043] Specifically, the first step is to extract the first code fragment to be analyzed from the list of complex expressions with IR nodes and their corresponding code fragments. This means the system will traverse all expressions marked as complex in the initial lineage analysis phase and process them one by one. For each node representing a complex operation in the code abstract syntax tree, its corresponding complete code fragment in the original SQL code text will be precisely extracted. For example, if the list to be analyzed contains a complex expression representing a summation calculation and a classification expression based on multiple conditional judgments, these complete text fragments will be used as input respectively to ensure that the large language model can perform independent and focused semantic analysis on them.

[0044] Next, the fragment context information of the first code fragment to be analyzed is extracted. This is a crucial step in ensuring that the large model can perform accurate semantic understanding, because a simple code fragment is often insufficient to provide enough semantic clues. In one embodiment, the fragment context information includes the location of the first code fragment to be analyzed in the original SQL code file, the target field alias, the table alias mapping, and the source fields referenced internally.

[0045] The location of the first code snippet to be analyzed within the original SQL code file refers to its starting line number and column number. While this location information doesn't directly participate in the core logic of lineage inference, it is helpful for subsequent metadata tagging, issue tracking, code auditing, or providing a user-friendly visualization interface. For example, during the debugging phase, precise location information can help developers quickly locate the problematic expression.

[0046] A target field alias refers to the field name that the complex expression ultimately represents in the data query results. For example, when calculating the average, this average will be assigned a specific name. Providing this alias helps the large model understand the expected final semantic output of the expression and produce the final lineage record using this alias, thereby ensuring naming consistency with the data models processed by downstream systems and facilitating subsequent integration and invocation.

[0047] Table alias mapping refers to the practice of using short aliases for tables in data queries to simplify expressions. For example, a table named "Customer Information" might be abbreviated to "C". The system builds a comprehensive mapping relationship, associating all table aliases used in SQL statements with their corresponding actual table names. For complex join queries, the join type and conditions must also be provided, which helps the larger model understand where the data comes from and how it was merged.

[0048] Internally referenced source fields refer to fields that are identified when a complex expression uses other fields (e.g., an expression calculating profit depends on the sales and cost fields). Detailed information includes the original table name, column name, original data type, and may even include a brief business description. This information is the cornerstone of accurate lineage reasoning in the large model, allowing it to understand the business meaning and data type of each field, thus more accurately determining how they participate in the final calculation. For example, if an expression calculating total contribution involves two sales-related numeric fields, the system will provide detailed information about these two original numeric fields and their source tables.

[0049] Next, as Figure 6 As shown, a complex bloodline relationship question Prompt is generated based on the first code segment to be analyzed and the segment context information. In one embodiment, generating a complex bloodline relationship question Prompt based on the first code segment to be analyzed and the segment context information includes: S531: embedding the first code segment to be analyzed and the segment context information into a preset Prompt template to obtain an initial complex bloodline relationship question Prompt; S532: semantically encoding the initial complex bloodline relationship question Prompt to obtain an initial complex bloodline relationship question Prompt semantic encoding vector; S533: semantically decoding the initial complex bloodline relationship question Prompt semantic encoding vector to obtain the complex bloodline relationship question Prompt.

[0050] Specifically, a pre-defined Prompt template is used first. This template aims to clearly articulate the task objective, the structure of the input data, and the expected output format to the large language model. It typically includes defining the model's role (e.g., instructing the model to act as an AI expert proficient in data lineage analysis), explicit task instructions (e.g., requiring the model to analyze a fragment of SQL expression and derive the precise field-level lineage relationship between its output and input fields), specific requirements for the expected output format (e.g., specifying that results should be returned in a structured text format, including key fields such as target column, source column, and source table), and guidance for situations where the source cannot be determined. Next, the system precisely fills the specified locations in this template with the text content of the previously extracted code snippets and all structured contextual information, thereby generating an initial complex lineage query Prompt.

[0051] Furthermore, the generated initial complex kinship question Prompt is fed into semantic encoding. This encoding process transforms discrete textual information into a semantically encoded vector of the initial complex kinship question Prompt that the model can process internally. This is accomplished through the embedding layer in the front of the large language model and the encoder part in its architecture, designed to capture the semantic information of the text. Then, using the decoder parameter mapping space weight matrix obtained through training, the semantically encoded vector of the initial complex kinship question Prompt is semantically decoded to obtain the complex kinship question Prompt.

[0052] As described above, since the type of the complex expression IR nodes in the list of complex expression IR nodes to be analyzed does not belong to simple column references, the code fragments and their context information corresponding to each complex expression IR node will have different semantic gradient responses after semantic encoding. Furthermore, since the preset Propmt template cannot dynamically adjust for specific encoded semantic gradient responses in order to improve its balanced adaptability, it is necessary to perform gradient response structure distribution remapping on the semantic encoding vector of the initial complex lineage question Propmt and the decoding parameter association domain of the semantic decoding to improve the same functionalized decoding parameter space mapping of the complex lineage question Propmt and improve the accuracy of the complex lineage question Propmt.

[0053] Therefore, in another embodiment, semantically decoding the initial complex kinship question Propmt semantic encoding vector to obtain the complex kinship question Propmt includes: extracting a decoding parameter mapping space weight matrix for semantic decoding; calculating the gradient partial derivative of the eigenvalues ​​of the decoding parameter mapping space weight matrix relative to each position in the initial complex kinship question Propmt semantic encoding vector to obtain a decoding weight tensor composed of n fine-grained decoding weight matrices; performing dilation correlation on the initial complex kinship question Propmt semantic encoding vector to obtain a Propmt semantic encoding dilation matrix; determining a semantic decoding fine-grained weight vector based on the vector standard deviation of the Propmt semantic encoding dilation matrix and the initial complex kinship question Propmt semantic encoding vector; and semantically decoding the initial complex kinship question Propmt semantic encoding vector based on the decoding weight tensor composed of n fine-grained decoding weight matrices and the semantic decoding fine-grained weight vector to obtain the complex kinship question Propmt.

[0054] Specifically, let V be the semantic encoding vector of the initial complex kinship question Propmt, then the decoding parameter mapping space of the semantic decoding can be represented as follows: Where M is the weight matrix of the decoding parameter mapping space. Here, in order to improve the mapping effect of the same functional decoding parameter space, the weight matrix of the decoding parameter mapping space is expanded to M1, M2, ..., M n This maps the decoding parameters. Transform into:

[0055]

[0056] Where V represents the initial complex kinship question Propmt semantic encoding vector, and ⊙ represents element-wise multiplication, that is, adding the corresponding value from the weight vector to the value at each position of the weight matrix. Represents matrix multiplication. This represents matrix addition, M1, M2, ..., M n This represents the fine-grained decoding weight matrices in the decoding weight tensor, ω1,ω2,...,ω n This represents each element in the semantic decoding fine-grained weight vector.

[0057] That is, the decoding parameter mapping space weight matrix is ​​expanded into a decoding weight tensor (M1, M2, ..., Mn) consisting of n fine-grained decoding weight matrices. n )∈R n×n×n and semantic decoding fine-grained weight vector (ω1,ω2,...,ω n )∈R 1×n R represents the set of real numbers.

[0058] Furthermore, in order to maintain the gradient response structure distribution of the decoding weight tensor with respect to the initial complex kinship question Propmt semantic encoding vector V, the weight matrix M of the decoding parameter mapping space is used to apply the following to each feature value v of the initial complex kinship question Propmt semantic encoding vector V. i The gradient partial derivative is used to obtain the corresponding fine-grained decoding weight matrix M. i ,Right now:

[0059]

[0060] Among them, v i M represents each feature value of the Propmt semantic encoding vector for the initial complex kinship question. i Indicates v i The corresponding fine-grained decoding weight matrix.

[0061] In the specific calculation process, it can be simplified to a fine-grained decoding weight matrix M. i Each eigenvalue m i-j,k The difference gradient response, i.e.:

[0062]

[0063] Where, m i-j,k M represents the fine-grained decoding weight matrix i Each eigenvalue m i-j,k .

[0064] The semantic decoding fine-grained weight vector (ω1,ω2,...,ω) n It is also expected that the feature distribution (v1, v2, ..., v) of the Propmt semantic encoding vector V, which is related to the initial complex blood relationship, will be used to ask questions. n Therefore, based on the weight space dilation, the initial complex blood relationship question Propmt semantic encoding vector V is first dilated and associated accordingly. Among them, M V Let T represent the Propmt semantic encoding inflation matrix, and let T represent the transpose of the vector. Then, based on the inflation correlation, a weight vector is obtained through predetermined coefficients, such as the standard deviation of the vector.

[0065]

[0066] Where (ω1,ω2,...,ω n ) represents the fine-grained weight vector for semantic decoding, and σ represents a predetermined coefficient, such as the standard deviation of the vector.

[0067] In this way, by remapping the gradient response structure distribution of the semantic decoding parameter association domain relative to the initial complex lineage question Propmt semantic encoding vector, the same functionalized decoding parameter space mapping of the complex lineage question Propmt is improved, thereby improving the accuracy of the complex lineage question Propmt.

[0068] Finally, the complex lineage relationship is input into the large language model via the Prompt, which outputs a complex lineage record corresponding to the first code segment to be analyzed. It should be understood that large language models typically employ a Transformer-based architecture, with a self-attention mechanism at its core, capable of efficiently processing sequence data and capturing long-distance dependencies. In one embodiment, the Prompt input to the large language model includes an encoder-decoder mode or a decoder-only mode. In the encoder-decoder architecture, the encoder is responsible for processing the input Prompt. It transforms each text unit in the Prompt into a high-dimensional, semantically rich context vector representation through a multi-layered self-attention mechanism and a feedforward neural network. In this process, the encoder can deeply capture the dependencies between text units within the Prompt, comprehensively understanding the semantics of the instructions, code segments, and all the information provided by the context. The decoder then generates new text units step by step based on the context vector output by the encoder and the previously generated text units, until the entire lineage record is output, following the output format specified in the Prompt.

[0069] To better perform lineage resolution tasks, large language models undergo extensive pre-training on massive amounts of text data, learning general language patterns, world knowledge, and the syntax and semantics of code languages. However, since SQL lineage resolution is a domain-specific task, further training is typically conducted using instruction fine-tuning or domain adaptation. This allows the model to learn to precisely understand SQL syntax and data processing logic, and to extract and infer field-level lineage relationships from complex SQL expressions. This fine-tuning often utilizes domain-specific datasets containing a large amount of SQL code, its corresponding lineage annotations, and even error examples and fix suggestions. In this way, the model not only masters general language capabilities but also possesses powerful specialized capabilities in SQL semantic understanding and lineage reasoning.

[0070] When a complex lineage relationship is input into the Prompt model, the large language model performs a series of complex internal calculations: it delves into the function calls, operators, conditional logic, and aggregation behaviors within the code snippet, while simultaneously understanding how these elements interact with the input data by combining the provided source field metadata and table alias mappings. Through complex transformations, calculations, and logical judgments, it ultimately produces output data. This process is a comprehensive application of syntactic rules, function semantics, and data type rules; it is not limited to the syntactic level but delves into the semantic level to understand the meaning of data transformations. Finally, the model generates a complex lineage record in text form, strictly adhering to the structured format required by the Prompt (e.g., a clear text structure containing information such as the target column, source column, its table, and transformation type).

[0071] To enable large language models to accurately and efficiently perform the aforementioned complex lineage resolution tasks, specialized training and domain knowledge enhancement are required. In one embodiment of this application, the knowledge sources and training process of the large language model are crucial to ensuring its powerful SQL semantic understanding and lineage reasoning capabilities.

[0072] First, regarding the knowledge sources of the large language model, it is built upon a general foundational model. This foundational model has been extensively pre-trained on massive and diverse data, including internet text, general code libraries, and academic literature, thereby mastering general language patterns, logical reasoning capabilities, and understanding of the basic syntax and structure of various programming languages, including SQL. However, to be competent in the highly specialized task of generating metadata for data platforms, it also needs to be infused with domain-specific knowledge. This domain knowledge primarily comes from a carefully constructed professional knowledge base oriented towards data lineage analysis. The construction of this knowledge base covers data from multiple dimensions: First, a large number of SQL code samples and their corresponding field-level lineage relationships, manually annotated or verified by experts. These samples cover various database dialects (such as Oracle, MySQL, SQL Server, HiveQL, Spark SQL, etc.) and various complex scenarios ranging from simple queries to those involving complex function calls, multi-level nested subqueries, conditional logic, and aggregation operations. Second, authoritative database technical documents, SQL syntax manuals, and best practice guides, which provide the model with precise definitions regarding function semantics, operator behavior, data type conversion rules, and other aspects. Thirdly, there are anonymous data model definitions, ETL scripts, and data dictionaries derived from real-world business scenarios. These resources help the model understand the business meaning of data fields and common transformation logic within specific industries or enterprises. By digitizing and structuring this professional knowledge, a rich knowledge source is formed, providing a solid foundation for subsequent model training.

[0073] Secondly, regarding the training process of the large language model, the main technical approach adopted is instruction fine-tuning or domain adaptation to deeply optimize the basic large model. The core of the training process is to enable the model to accurately output structured field-level lineage relationships when it receives a Prompt containing SQL code snippets and contextual information. The specific training dataset consists of paired "input-output" samples. The input is a complex lineage relationship question Prompt simulating real-world application scenarios, while the output is the standard-format lineage relationship label corresponding to that Prompt. During training, the model generates a predicted lineage relationship based on the input Prompt. The system then compares this prediction with the preset correct answers in the dataset, calculating the difference or loss. Subsequently, through backpropagation, the model's internal parameters (such as the weights in the Transformer layer) are adjusted based on this loss value, aiming to minimize the gap between the prediction and the true label. This process is iterative, with the model repeatedly learning and optimizing on tens of thousands or even millions of specialized samples, gradually mastering the professional ability to accurately infer field-level lineage relationships from complex SQL expressions. In this way, the model not only solidifies its understanding of SQL syntax, but more importantly, it learns to delve into the semantic level, understand the internal logic of data transformation, and ultimately can accurately parse complex data flow paths like an experienced data engineer.

[0074] To further illustrate the technical solution of this application, an embodiment in a specific application field will be described below. It should be noted that this embodiment is a specific application of the aforementioned general method, and its core idea remains consistent with the aforementioned embodiments, but it differs in the source of knowledge and the focus of model fine-tuning.

[0075] In one specific embodiment, the method of this application can be applied to the automatic generation of metadata for data platforms in the water conservancy industry. In this scenario, the knowledge sources and training process of the large language model will be more focused on the professional characteristics of the water conservancy industry.

[0076] In terms of knowledge sources, in addition to the aforementioned general SQL code libraries and technical documents, a key focus will be on building a professional knowledge base specifically for the water conservancy industry. This knowledge base will specifically collect and organize massive amounts of data related to the water conservancy field. This includes, but is not limited to, official documents such as water conservancy data standards, hydrological and water resources terminology specifications, and guidelines for water conservancy engineering design and management. These documents provide authoritative evidence for the model to understand water conservancy professional terms (such as "runoff", "reservoir capacity", "soil moisture", "water quality grade") and their inherent relationships. Simultaneously, a large amount of anonymous database design documents, data models, ETL scripts, and data report templates from real-world water conservancy information projects (such as reservoir scheduling systems, irrigation district management systems, flood control and drought relief command systems, etc.) will be collected. These real-world samples enable the model to learn the data processing logic and lineage patterns unique to the water conservancy industry. For example, how rainfall data is converted into runoff data through a watershed confluence model, or the complex constraints and calculation relationships between reservoir water level, inflow, and outflow.

[0077] In terms of model training, the same instruction-based fine-tuning approach is used, but the training dataset is centered on professional data from the water conservancy industry. The construction of training samples simulates common scenarios in water conservancy business analysis. For example, the input Prompt of a training sample might contain a SQL code snippet for calculating the multi-year average runoff at a specific cross-section, with contextual information including metadata such as relevant station information tables and flow data tables. The correct output from the model is a lineage record that precisely indicates the target field, "multi-year average runoff," originates from the "instantaneous flow" field in the flow data table and has undergone aggregation calculations over a specific time period. By fine-tuning on this highly specialized and scenario-based dataset, the large language model can deeply learn and internalize the business rules and data transformation logic of the water conservancy industry. This enables the model to not only perform general syntax parsing when processing SQL code in the water conservancy industry, but also to combine its background knowledge to accurately understand the specific calculations and dependencies of business terms such as "upstream water inflow" and "downstream flood discharge" at the data level. This significantly improves the accuracy and depth of metadata lineage generation in the water conservancy data platform, providing solid support for the refined management and intelligent application of water conservancy data assets.

[0078] For example, in step S6, the preliminary lineage list and the complex lineage list are integrated to obtain a structured field-level lineage relationship list. It should be understood that preliminary lineage analysis focuses on handling high-frequency, deterministic, simple scenarios to improve efficiency, while complex lineage parsing leverages the powerful capabilities of large models to overcome semantic understanding barriers that traditional methods struggle to overcome, achieving broader scenario coverage and higher accuracy. However, for downstream core metadata consumers such as data governance, data tracing, and impact analysis, what they need is a unified, seamless field-level lineage panorama, not isolated fragments. In actual execution, the output of an SQL statement may simultaneously depend on the direct transmission of simple columns and may also contain new fields calculated from complex expressions. If these two parts of lineage information are not effectively integrated, the generated metadata will be fragmented and incomplete, failing to truly and comprehensively reflect the complete flow and transformation path of data within the entire SQL statement, significantly diminishing its value in guiding business decisions and technical maintenance. Therefore, by integrating the preliminary bloodline list and the complex bloodline list to obtain a structured field-level bloodline list, it is ensured that regardless of the complexity of the bloodline relationship, it can ultimately be provided to the outside world in a consistent and standardized format, thereby laying a solid foundation for subsequent metadata management and application.

[0079] Specifically, first, ensure that every record in both the initial lineage list and every record generated from the large model in the complex lineage list strictly adheres to a predefined, standardized data format. Typically, such a field-level lineage record contains at least several key elements: for example, the name of the output field being analyzed and its corresponding logical table or alias, indicating the direction of the lineage; next, the names of one or more input source fields that the output field depends on, and their respective original tables or aliases, describing the origin of the lineage; furthermore, it may include the type of lineage relationship, such as direct pass-through, transformation operation, aggregation calculation, etc., to provide richer semantic information. By enforcing this unified data structure, the two independently generated lineage lists naturally possess format compatibility, removing obstacles for subsequent merging operations.

[0080] With a unified data structure, the integration process transforms into a logical merging operation. This involves logically appending or merging all lineage records collected in the initial lineage list with all lineage records output from the large model in the complex lineage list. This is similar to the union operation in set theory. A crucial consideration during this merging process is the potential for duplication or conflict. Theoretically, since the initial lineage analysis and the large model's complex lineage parsing process different types of expressions in SQL statements, the resulting lineage records are largely complementary, with a low probability of direct conflict. However, to ensure the purity and accuracy of the final generated list, the system typically incorporates a deduplication mechanism. This can be achieved by validating key identifiers of lineage records, such as unique combinations of target field-source field-source table. If identical lineage records are found (i.e., all key attributes are consistent), only one is retained. Alternatively, if subtle differences exist due to different parsing paths, priority rules may need to be set (e.g., complex lineage records from the large model have higher priority) to ensure the authority of the final list.

[0081] Through this unified data structure and rigorous merging logic, the originally separate lineage information is organically combined, forming a unified, coherent, and structured list of field-level lineage relationships. This final list is not merely a simple collection of information; it constructs a semantic graph that clearly depicts all field-level transformations and dependencies of data from source to target, from input to output, and from raw data to derived products.

[0082] For example, in step S7, the lineage relationships in the structured field-level lineage relationship list are associated with the corresponding data tables and field source data to obtain data platform metadata. It should be understood that although the preceding steps have successfully extracted and integrated a structured field-level lineage relationship list from complex SQL code, this lineage information itself is still relatively sparse. It merely indicates how data flows and is transformed, but lacks broader contextual information and business attributes. For example, the lineage of a field might indicate that it originates from the amount field of the order table, but if it is unknown that the order table is a core table of the financial system with the highest data quality level, or that the business meaning of the amount field is the total transaction amount including tax, then the value of this lineage information will be greatly reduced.

[0083] Therefore, further, pure lineage relationships are associated with the corresponding data tables and fields using richer and more comprehensive source data (i.e., their existing metadata information, such as data type, length, constraints, description, hierarchy, owner, data quality metrics, business terminology mappings, etc.). This association operation can weave scattered lineage clues into an interconnected, semantically rich metadata network, thereby truly supporting the various complex data governance and management needs in the data platform.

[0084] Specifically, first, a metadata baseline repository is built or accessed. The data platform maintains a metadata repository or database containing detailed information about all data tables and fields. This information may come from various sources, such as: Database system catalogs: automatically collecting physical metadata such as table names, column names, data types, primary keys, foreign keys, and indexes; ETL or data integration tool metadata: recording configuration information during data loading and transformation; Business metadata management tools: manually or automatically entered business definitions, data dictionaries, business terms, data owners, business responsible persons, etc.; Data quality or security tools: recording data quality rules, sensitive data tags, access permissions, etc. This metadata baseline repository is the foundation for association operations, providing rich context for lineage information. Second, information is associated through key-value matching. Each lineage record in the structured field-level lineage relationship list contains the name of the source field, the name of the table to which the source field belongs, the name of the target field, and the name of the table to which the target field belongs. These names are the keys for association. The system iterates through each lineage record in the lineage list. For each source table and source field mentioned in the lineage record, the system searches the pre-built metadata benchmark library, matches the table name and field name, and retrieves the corresponding detailed metadata. Similarly, the same matching operation is performed on the target table and target field mentioned in the lineage record to obtain all available metadata information.

[0085] Furthermore, association operations also include the integration of lineage relationships and the enrichment of metadata attributes. The original field-level lineage relationship list only represents the dependencies between fields. Through association, these lineage relationships themselves are also endowed with more metadata attributes. For example, based on information such as the data type and sensitivity of the source and target fields, the potential data quality risks and data security propagation paths involved in the lineage relationship can be automatically inferred. If the source field is sensitive data, then the target field derived through this lineage relationship should also be marked as sensitive. At the same time, operation metadata can be added to the lineage relationship by analyzing the SQL statement type (e.g., whether it is a DML operation or a DDL operation). Finally, a comprehensive data platform metadata is formed. The information after association is stored in the data platform's metadata store, forming a multi-dimensional, interconnected metadata graph.

[0086] In summary, the automatic generation method for data platform metadata based on a large-scale model provided in this application first performs unified standardization and preliminary syntactic analysis on the original SQL code to obtain a structured intermediate representation. Based on this, a rule-based preliminary lineage analysis is conducted to quickly identify and process simple column references. For complex expressions that are difficult to parse accurately using traditional methods, their code snippets and contextual information are precisely extracted and then passed to a large-scale language model for deep semantic understanding and complex lineage parsing. Finally, the complex lineage parsed by the large-scale model is integrated with the preliminary lineage list to form a comprehensive and accurate field-level lineage relationship, thereby generating complete data platform metadata. This effectively compensates for the shortcomings of traditional parsing tools in understanding complex semantics, significantly improving the accuracy and completeness of metadata generation.

[0087] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.

[0088] The flowcharts of the methods involved in this application are merely illustrative examples and are not intended to require or imply that connections, arrangements, or configurations must be made in the manner shown in the flowcharts. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0089] It should also be noted that the steps in the method of this application can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions of this application.

[0090] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0091] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A method for automatically generating metadata for a data platform based on a large model, characterized in that, include: Obtain the original SQL code file; The original SQL code file is standardized to obtain standardized code text and code type identifier; Based on the code type identifier, the corresponding static SQL parser is invoked to perform syntactic analysis on the standardized code text to obtain a structured intermediate representation of the code; The structured intermediate representation of the code is subjected to rule-based preliminary lineage analysis and fragment extraction to obtain a preliminary lineage list and a list of complex expression IR nodes to be analyzed and their corresponding code fragments; Large-scale semantic understanding and complex lineage parsing are performed on the list of IR nodes of the complex expression to be analyzed and its corresponding code snippets to obtain a complex lineage list; The preliminary bloodline list and the complex bloodline list are integrated to obtain a structured field-level bloodline list; The lineage relationships in the structured field-level lineage relationship list are associated with the corresponding data tables and field source data to obtain data platform metadata.

2. The method for automatically generating metadata for a data platform based on a large model according to claim 1, characterized in that, The original SQL code file is standardized to obtain standardized code text and code type identifiers, including: Identify the database type or processing framework to which the original SQL code file belongs to obtain the code type identifier; The original SQL code file is case-sensitive, unnecessary comments are removed, and formatting is performed to obtain the standardized code text.

3. The method for automatically generating metadata for a data platform based on a large model according to claim 2, characterized in that, Based on the code type identifier, the corresponding static SQL parser is invoked to perform syntax analysis on the standardized code text to obtain a structured intermediate representation of the code, including: Lexical analysis is performed on the standardized code text to obtain a sequence of lexical units; The static SQL parser performs syntactic analysis on the lexical unit sequence based on the grammatical rules of the target SQL dialect to obtain a code abstract syntax tree as a structured intermediate representation of the code.

4. The method for automatically generating metadata for a data platform based on a large model according to claim 3, characterized in that, The structured intermediate representation of the code undergoes rule-based preliminary lineage analysis and fragment extraction to obtain a preliminary lineage list and a list of complex expression IR nodes to be analyzed, along with their corresponding code fragments, including: The structured intermediate representation of the code is traversed to obtain a list of code nodes; Extract the expressions of each code node in the code node list to obtain a code node expression list; Type determination is performed on each code node expression in the code node expression list to obtain the type determination result; If the type determination result is a simple column reference, a preliminary lineage record is generated based on the corresponding code node expression to obtain the preliminary lineage list; If the type determination result is not a simple column reference, the corresponding code node is marked as a complex expression IR node to obtain the list of complex expression IR nodes to be analyzed and their corresponding code snippets.

5. The method for automatically generating metadata for a data platform based on a large model according to claim 1, characterized in that, Large-scale model semantic understanding and complex lineage parsing are performed on the list of IR nodes of the complex expression to be analyzed and its corresponding code snippets to obtain a complex lineage list, including: Extract the first code segment to be analyzed from the list of IR nodes of the complex expression to be analyzed and its corresponding code segments; Extract the fragment context information of the first code fragment to be analyzed; Based on the first code segment to be analyzed and the context information of the segment, a complex blood relationship question Propmt is generated; The complex bloodline question Propmt is input into the large language model to output a complex bloodline record corresponding to the first code segment to be analyzed.

6. The method for automatically generating metadata for a data platform based on a large model according to claim 5, characterized in that, The fragment context information includes the location of the first code fragment to be analyzed in the original SQL code file, the target field alias, the table alias mapping, and the source field referenced internally.

7. The method for automatically generating metadata for a data platform based on a large model according to claim 5, characterized in that, Based on the first code snippet to be analyzed and the context information of the snippet, a complex blood relation question Propmt is generated, including: The first code segment to be analyzed and the context information of the segment are embedded into a preset Propmt template to obtain an initial complex blood relationship question Propmt. The initial complex kinship question Propmt is semantically encoded to obtain the initial complex kinship question Propmt semantic encoding vector; The initial complex kinship question Propmt semantic encoding vector is semantically decoded to obtain the complex kinship question Propmt.

8. The method for automatically generating metadata for a data platform based on a large model according to claim 7, characterized in that, Semantic decoding of the initial complex kinship question Propmt semantic encoding vector to obtain the complex kinship question Propmt includes: Extract the weight matrix of the decoding parameter mapping space used for semantic decoding; Calculate the gradient partial derivative of the weight matrix of the decoding parameter mapping space with respect to the feature values ​​at each position in the initial complex bloodline question Propmt semantic encoding vector to obtain a decoding weight tensor composed of n fine-grained decoding weight matrices; The initial complex bloodline question Propmt semantic encoding vector is expanded and correlated to obtain the Propmt semantic encoding expansion matrix; Based on the Propmt semantic coding inflation matrix and the vector standard deviation of the Propmt semantic coding vector obtained from the initial complex lineage relationship, the fine-grained weight vector for semantic decoding is determined. Based on the decoding weight tensor composed of n fine-grained decoding weight matrices and the semantic decoding fine-grained weight vector, the initial complex kinship question Propmt semantic encoding vector is semantically decoded to obtain the complex kinship question Propmt.

Citation Information

Cited By

  • Data consanguinity extraction and completion method and system for development script, equipment and medium

    CN121234341A

  • Data bloodline extraction completion method and system, device and medium for developing scripts

    CN121234341B

  • Secret-related document classification and exchange permission management and control method based on natural language processing

    CN121389190A

  • Operator-level data blood relationship automatic generation method based on large model

    CN121390083A

  • SQL (Structured Query Language) blood relationship analysis method, system, equipment and medium

    CN121597772A