Method and device for constructing a blood relationship knowledge graph of civil affairs data based on a large model
By building a basic entity network and a large language model to parse complex SQL statements, identify field-level lineage associations, eliminate virtual table redundancy, and generate an optimized lineage relationship map, the complex SQL parsing and redundancy problems in traditional methods are solved, and efficient data lineage management and decision support are achieved.
Patent Information
- Application Number
- CN202510522095.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Traditional data lineage analysis methods have difficulty accurately identifying complex SQL statements, handling virtual table redundancy and data node duplication, and lack adaptability, especially when processing nested SQL logic.
A hierarchical and progressive approach based on a large language model is adopted to build a basic entity network, parse complex structured query language statements, identify field-level lineage relationships, eliminate virtual table redundancy and data node duplication, and use breadth-first search and data alignment algorithms to generate an optimized lineage relationship map.
It has achieved the construction of a high-precision and automated bloodline knowledge graph for civil affairs data, improved data management transparency and processing efficiency, simplified the data relationship network, and supported more effective data governance and decision-making.
Smart Images

Figure CN120032724B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management and knowledge graph technology, and in particular to a method and device for constructing a civil affairs data bloodline knowledge graph based on a large model. Background Art
[0002] As civil affairs operations become increasingly digital, the scale and complexity of civil affairs data continues to grow. Data lineage, the cornerstone of data governance, refers to the complete record of data from its source to its final application. This is crucial for ensuring data quality, enhancing data value, and achieving data compliance. Tracing and managing data lineage not only helps civil affairs departments understand the origins and development of data but also allows them to quickly identify the causes of data issues, effectively preventing data errors and misuse. Therefore, the construction of data lineage knowledge graphs has become a research hotspot in the fields of data management and knowledge graph technology.
[0003] Traditional data lineage analysis relies primarily on manual annotation or rule-based system log parsing. While these methods achieve a certain degree of data lineage tracking, their limitations gradually become apparent when faced with complex and ever-changing data processing scenarios. For example, traditional rule-based methods struggle to accurately identify dynamic relationships across heterogeneous data sources, especially when dealing with complex processing logic such as nested SQL statements. Furthermore, entity relationship extraction often relies on a fixed rule base, making it difficult for the system to adapt to dynamic changes in data processing logic and lacking adaptability.
[0004] With the rapid development of natural language processing (NLP) technology, large language models (LLMs), a breakthrough technology in the field, have demonstrated powerful semantic understanding and generation capabilities. By learning from massive amounts of text data, LLMs can understand and generate semantically rich linguistic information, providing new insights for the automatic extraction of data lineage. However, applying LLMs to data lineage extraction still faces many challenges. Summary of the Invention
[0005] The present invention provides a method and device for constructing a lineage knowledge graph of civil affairs data based on a large model, which solves the defects of the existing technology in accurately parsing complex SQL statements, processing virtual table redundancy and data node duplication, and realizes high-precision and automated construction of a lineage knowledge graph of civil affairs data.
[0006] The present invention provides a method for constructing a blood relationship knowledge graph of civil affairs data based on a large model, comprising the following steps:
[0007] Extract metadata features of civil affairs data to build a basic entity network, the basic entity network including data tables and fields;
[0008] Based on the basic entity network, the data table structure is parsed to generate a data table structure relationship model;
[0009] Based on the data table structure relationship model, the complex structured query language statements are parsed through the large model to identify field-level blood relationship and generate a preliminary blood relationship map;
[0010] Based on the preliminary blood relationship map, all physical source tables are traversed to eliminate virtual table redundancy and data node duplication, thereby generating an optimized blood relationship map;
[0011] Based on the optimized blood relationship map, common data nodes are identified and field-level data blood relationships are merged to generate the final civil affairs data blood relationship map.
[0012] According to the present invention, a method for constructing a blood relationship knowledge graph of civil affairs data based on a big model is provided. The method parses complex structured query language statements through the big model, identifies field-level blood relationship associations, and generates a preliminary blood relationship graph. The method specifically includes: parsing the structured query language statement into an enhanced abstract syntax tree; generating a comprehensive complexity score of the structured query language query based on the enhanced abstract syntax tree; if the complexity score exceeds a preset threshold, decomposing the complex structured query language statement into multiple atomic structured query language query units through the big model; extracting data table generation paths and field-level mapping rules from the decomposed structured query language query units to generate a preliminary blood relationship graph.
[0013] According to a method for constructing a large-model-based bloodline knowledge graph of civil affairs data provided by the present invention, the method of parsing a structured query language statement into an enhanced abstract syntax tree specifically includes: preprocessing a structured query language query string; performing preliminary parsing of the preprocessed structured query language query string to generate a basic abstract syntax tree; traversing each child node of the basic abstract syntax tree and performing different processing according to the type of the child node; and returning the enhanced abstract syntax tree after traversing and processing all child nodes.
[0014] According to a method for constructing a large-model-based bloodline knowledge graph of civil affairs data provided by the present invention, the comprehensive complexity score of the structured query language query generated according to the enhanced abstract syntax tree specifically includes: extracting complexity features in the enhanced abstract syntax tree, normalizing each extracted complexity feature, and calculating the normalized value of each complexity feature; assigning a corresponding weight to each normalized complexity feature; and calculating the comprehensive complexity score according to the normalized value of each complexity feature and the corresponding weight.
[0015] According to a method for constructing a large-model-based lineage knowledge graph of civil affairs data provided by the present invention, the method extracts data table generation paths and field-level mapping rules from the disassembled structured query language query units to generate a preliminary lineage relationship graph, which specifically includes: inputting the disassembled multiple atomic structured query language query units into the large model; parsing each structured query language query unit to extract table-level lineage relationships and field-level lineage relationships; and integrating the extracted table-level lineage relationships and field-level lineage relationships to generate a preliminary lineage relationship graph.
[0016] According to a method for constructing a large-scale model-based blood relationship knowledge graph of civil affairs data, the method is based on the preliminary blood relationship graph, traverses all physical source tables, eliminates virtual table redundancy and data node duplication, and generates an optimized blood relationship graph, specifically including: constructing an adjacency list of tables and fields according to the preliminary blood relationship graph; marking virtual tables from the adjacency list and filtering physical source tables; initializing the queue, visited node set and merged mapping list required for breadth-first search, and adding the physical source table to the queue; traversing the adjacency list through breadth-first search, merging field mapping relationships, and updating the merged mapping list and visited node set; deduplicating the merged mapping list to remove redundant mapping relationships; and converting the deduplicated mapping list into an optimized blood relationship graph.
[0017] According to a method for constructing a large-scale model-based civil affairs data lineage knowledge graph provided by the present invention, based on the optimized lineage relationship graph, public data nodes are identified and field-level data lineages are merged to generate a final civil affairs data lineage relationship graph, specifically including: traversing all table nodes in the optimized lineage relationship graph to identify public data nodes; for each identified public data node, triggering the field-level alignment process to perform field node matching; for successfully matched field nodes, merging their field-level data lineages and integrating the merged field-level data lineages into the optimized lineage relationship graph to generate a final civil affairs data lineage relationship graph.
[0018] The present invention also provides a device for constructing a blood relationship knowledge graph of civil affairs data based on a large model, comprising the following modules:
[0019] A basic entity network construction module is used to extract metadata features of civil affairs data to construct a basic entity network, the basic entity network including data tables and fields;
[0020] A data table structure relationship model generation module is used to parse the data table structure based on the basic entity network and generate a data table structure relationship model;
[0021] A preliminary blood relationship map generation module is used to parse complex structured query language statements through a large model based on the data table structure relationship model, identify field-level blood relationship associations, and generate a preliminary blood relationship map;
[0022] An optimized blood relationship map generation module is used to traverse all physical source tables based on the preliminary blood relationship map, eliminate virtual table redundancy and data node duplication, and generate an optimized blood relationship map;
[0023] The final blood relationship map generation module is used to identify common data nodes and merge field-level data blood relationships based on the optimized blood relationship map to generate the final civil affairs data blood relationship map.
[0024] The present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements a method for constructing a large-model-based bloodline knowledge graph of civil affairs data as described in any one of the above-mentioned methods.
[0025] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the method for constructing a large-model-based civil affairs data bloodline knowledge graph as described in any of the above.
[0026] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the method for constructing a large-model-based civil affairs data bloodline knowledge graph as described in any of the above.
[0027] The present invention provides a method and device for constructing a large-scale model-based lineage knowledge graph of civil affairs data, which includes the following beneficial effects: By constructing a data lineage knowledge graph, the source and flow of data can be clearly traced, thereby improving the transparency of data management, helping organizations to better understand their data assets, and optimize the use and management of data. By parsing data table structures and SQL statements and identifying field-level lineage associations, organizations can discover dependencies between data, and then quickly locate and repair data problems when they occur, thereby improving data quality. By generating an optimized lineage relationship graph, virtual table redundancy and data node duplication can be eliminated, thereby reducing unnecessary data processing and improving data processing efficiency. By identifying common data nodes and merging field-level data lineage, the data relationship network can be further simplified, helping organizations to conduct data governance and decision-making more effectively. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0029] Figure 1 This is the overall architecture diagram of the SQL data lineage extraction model based on the large model provided by the present invention.
[0030] Figure 2 It is a schematic diagram of the fusion process of the data lineage of the virtual table and the field provided by the present invention.
[0031] Figure 3 It is a schematic diagram of the field-level data lineage fusion process provided by the present invention.
[0032] Figure 4 It is a flow chart of the method for constructing a large-scale model-based civil affairs data bloodline knowledge graph provided by the present invention.
[0033] Figure 5 This is an example diagram of SQL disassembly input and output based on Prompt provided by the present invention.
[0034] Figure 6 This is a flow chart of data lineage extraction based on LLM provided by the present invention.
[0035] Figure 7 It is a structural diagram of the device for constructing the large-scale model-based civil affairs data bloodline knowledge graph provided by the present invention.
[0036] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0037] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0038] As civil affairs operations become increasingly digital, the scale and complexity of civil affairs data continues to grow. Data lineage tracing and management have become core requirements for civil affairs data governance. This is particularly true in scenarios like cross-departmental data sharing, policy effectiveness analysis, and data audits. Accurate lineage information can help civil affairs departments clarify data sources, flows, and processing logic, ensuring data accuracy and credibility.
[0039] Traditional data lineage analysis typically relies on manual annotation or rule-based system log parsing. These methods are limited by their difficulty in accurately identifying dynamic relationships across heterogeneous data sources, particularly complex processing logic such as nested SQL statements. Furthermore, traditional lineage graph construction often relies on fixed rule bases for entity relationship extraction, making it difficult to adapt to dynamic changes in data processing logic. Therefore, adaptive, fine-grained, automated lineage modeling methods are urgently needed.
[0040] In response to the above-mentioned technical bottlenecks, the present invention proposes a hierarchical and progressive method for constructing a civil affairs data lineage knowledge graph. Through a rule-based structured data lineage parsing method, data lineage extraction of structured metadata and field-level parsing of the data definition language are realized. Based on the semantic understanding representation capability driven by a large language model and an independently designed lineage extraction module, more efficient and accurate SQL field lineage analysis is achieved. The proposed lineage fusion algorithm is used to achieve comprehensive and accurate acquisition of data lineage results, and ultimately realize the construction of an accurate and effective data lineage knowledge graph.
[0041] (1) Regular Expression
[0042] Regular expressions are tools for text pattern matching and play a vital role in field-level parsing of data definition languages (DDLs). By customizing regular patterns, models can extract important structural information from DDL statements, such as table names, field names, field data types, and field comments. Regular expressions can quickly and accurately parse structured DDL statements, helping to establish mappings between tables and fields. For example, regular expressions can be used to extract field definitions from table creation statements, identify field data types, comments, and other information, and handle partition field extraction, foreign key constraints, and index information extraction.
[0043] (2) SQL syntax analysis
[0044] SQL syntax parsing is the basis for understanding the structure of SQL statements. Its purpose is to convert SQL statements into an abstract syntax tree (AST) that is easy for computers to process. In this paper, the SQL syntax parsing module uses the open source sqlparse library, which can convert complex SQL query statements into abstract syntax trees and identify key information such as table names, fields, and query conditions in the query through further syntax processing. However, due to the diversity and complexity of SQL statements
[12]
[13] , the standard sqlparse library cannot handle all complex syntax structures, such as CTE (common table expressions), nested queries, and JOIN conditions. Therefore, on this basis, this paper further enhances the processing capabilities of complex SQL structures, and supports the parsing of CTE, nested queries, window functions, etc. through customized extension rules, so that the model can be more accurate in parsing complex queries.
[0045] (3) Large Language Model and Fine-tuning
[0046] Large Language Models (LLMs) are a breakthrough technology in the field of natural language processing in recent years. They can understand and generate semantically rich language information by learning from massive amounts of text data
[16] . In the task of data lineage extraction, large language models can be used to parse complex SQL statements and identify relationships between tables and mappings between fields. Through deep learning technology, LLMs can automatically extract semantically layered lineage relationships from complex SQL queries without the need for manual rule writing. Especially when faced with complex operations such as nested queries, conditional filtering, and aggregation, LLMs can accurately predict data flows by understanding the context of the statement.
[0047] Fine-tuning is a technique that uses a small amount of task-related data to perform targeted optimization on an existing pre-trained model. While pre-trained large language models typically demonstrate strong capabilities across a wide range of text understanding tasks, fine-tuning is necessary to better understand SQL syntax and the characteristics of the fine-tuning data in order to adapt to specific application scenarios such as data lineage extraction. Fine-tuning a small-scale, open-source large language model can significantly improve the model's performance in data lineage extraction, particularly when solving domain-specific problems encountered during SQL parsing. This can significantly improve extraction accuracy and efficiency. The fine-tuning process not only helps the model better adapt to task requirements but also reduces computing resource consumption and protects data privacy.
[0048] In summary, the present invention discloses a method for constructing a lineage knowledge graph of civil affairs data based on a large model, which belongs to the field of data management and knowledge graph technology. First, in response to the difficulty of SQL complex logic parsing in traditional data lineage analysis, this method constructs a hierarchical and progressive knowledge extraction framework: based on the metadata feature extraction of civil affairs data, a basic entity network is constructed, and the DDL parsing technology is used to realize data table structure relationship modeling; in response to the bottleneck problem of traditional SQL semantic understanding, a lineage relationship extraction model based on a large language model and comparative learning is proposed, which uses the powerful semantic understanding ability of the large language model to break through the limitations of traditional rule matching and realize accurate identification of field-level lineage associations. In order to solve the virtual table redundancy problem and data node duplication problem caused by complex SQL disassembly, a virtual table mapping fusion based on breadth-first search and a data alignment algorithm based on key nodes are proposed. Compared with traditional methods, the present invention has stronger semantic understanding ability and scenario adaptability in complex scenario data lineage identification, and provides a high-precision decision support framework for data governance.
[0049] The following combination Figures 1-8 The embodiments of the present invention are described in detail.
[0050] 1. SQL data lineage extraction model based on large model
[0051] This paper proposes a SQL Hybrid Parsing Model (SQLHPM), which integrates the dual capabilities of grammatical structure analysis and semantic content understanding, and realizes adaptive adjustment of parsing strategies through dynamic complexity evaluation. Figure 1 shown.
[0052] The SQL Hybrid Parsing Model (SQLHPM) achieves efficient processing of complex SQL statements through a multi-stage collaborative mechanism. It comprises a syntax parsing module, a complexity assessment module, an SQL decomposition module, and a data lineage extraction module. The syntax parsing module, implemented based on the sqlparse library, constructs an enhanced abstract syntax tree (AST) from raw SQL statements. It accurately identifies complex syntactic structures such as nested queries and subquery associations, and generates standardized syntactic representations. The complexity assessment module quantifies features such as the nesting level, number of JOIN operations, and subquery depth in the AST to generate a comprehensive complexity score and trigger a threshold judgment mechanism. For highly complex statements with complexity scores exceeding the threshold, the SQL decomposition module leverages the semantic understanding capabilities of the large language model (GLM4-9b-chat) to decompose complex logic into atomic SQL query units to reduce coupling. Finally, the data lineage extraction module extracts data table generation paths and field-level mapping rules from the decomposed logical units, outputting a standardized lineage relationship map to extract data lineage relationships.
[0053] 2. Virtual table mapping fusion algorithm based on breadth-first search
[0054] This paper proposes a virtual table mapping fusion algorithm based on breadth-first search, using breadth-first search (BFS) in a graph structure to perform schema-level lineage fusion. The algorithm aims to analyze the mapping relationships between source and target tables, merge corresponding tables and field mappings, and eliminate redundancy. Specifically, the algorithm first establishes adjacency relationships between tables and marks each table as a virtual table. Based on this, the algorithm further performs BFS, traversing all physical source tables and merging their mapping relationships to ensure uniqueness and consistency of the results. Figure 2 The data lineage fusion process including virtual tables and fields is described in detail.
[0055] 3. Data alignment based on key nodes
[0056] The key to data alignment is finding common nodes. Generally speaking, each data table name is exact and unique. That is, if two data tables have the same name, they are the same data node. For data fields, if the data tables have the same name, find the nodes with the same fields under them. If there are any, then the node is considered a common node. After a lineage fusion, the number of fields in a common node may not match, so all fields must be merged. Figure 3 The field-level data lineage fusion process is described in detail.
[0057] Figure 4 This is a flow chart of the method for constructing a large-scale model-based civil affairs data lineage knowledge graph provided by the present invention. Figure 4 As shown, the method includes the following steps:
[0058] S410. Extract metadata features of civil affairs data and build a basic entity network, where the basic entity network includes data tables and fields.
[0059] Specifically, in the initial stages of data lineage analysis, traditional methods often find it difficult to accurately extract the relationships between data tables and fields from heterogeneous data sources, especially when faced with complex metadata structures, which can easily lead to information loss or association errors. In addition, the initial relationships between data tables and fields are usually scattered across different metadata sources, lacking a unified integration mechanism, making subsequent lineage analysis difficult to conduct efficiently. Building a basic entity network based on metadata feature extraction from civil affairs data can effectively solve these problems. By extracting key features of data tables and fields (such as table name, field name, data type, annotation, etc.) from metadata and constructing an initial entity network, structured basic data is provided for subsequent lineage analysis. This step ensures that the relationships between data tables and fields can be accurately captured and integrated, avoiding information loss and association errors.
[0060] By extracting metadata features and constructing a basic entity network, the initial relationships between data tables and fields are clearly structured, providing reliable foundational data for subsequent lineage analysis. The basic entity network integrates information scattered across different metadata sources into a unified network, reducing the complexity and time cost of data integration. The basic entity network can adapt to a variety of metadata structures and data sources, ensuring accurate extraction of relationships between data tables and fields in different scenarios. By accurately capturing the initial relationships between data tables and fields, the basic entity network provides high-quality data input for subsequent lineage analysis, significantly improving its accuracy.
[0061] S420: Based on the basic entity network, the data table structure is parsed to generate a data table structure relationship model.
[0062] By parsing the data table structure, the generated data table structure relationship model clearly shows the association relationship between data tables (such as primary and foreign key relationships, field mapping relationships, etc.), providing structured basic data for subsequent lineage analysis. Using DDL parsing technology, the structural information of the data table (such as table name, field name, data type, etc.) can be accurately extracted, avoiding errors caused by imperfect rules or manual intervention in traditional methods. The data table structure relationship model can integrate table structure information from different data sources to ensure that the relationship between data tables can be accurately parsed even in heterogeneous data environments. Through automated data table structure parsing, the need for manual intervention is reduced, the efficiency of data lineage analysis is significantly improved, and time and cost consumption are reduced. The generated data table structure relationship model provides a foundation for subsequent complex SQL parsing, ensuring that the relationship between tables and fields can be accurately identified when parsing complex SQL structures such as nested queries and JOIN operations.
[0063] S430. Based on the data table structure relationship model, the complex structured query language statements are parsed through the large model, the field-level blood relationship is identified, and a preliminary blood relationship map is generated.
[0064] According to a method for constructing a blood relationship knowledge graph of civil affairs data based on a big model provided by the present invention, complex structured query language statements are parsed through the big model, field-level blood relationship associations are identified, and a preliminary blood relationship graph is generated. The method specifically includes: parsing the structured query language statement into an enhanced abstract syntax tree; generating a comprehensive complexity score of the structured query language query based on the enhanced abstract syntax tree; if the complexity score exceeds a preset threshold, decomposing the complex structured query language statement into multiple atomic structured query language query units through the big model; extracting the data table generation path and field-level mapping rules from the decomposed structured query language query units to generate a preliminary blood relationship graph.
[0065] Specifically, the main components of the large-scale model-based SQL data lineage extraction model include: a syntax parsing module, a complexity assessment module, an SQL decomposition module, and a data lineage extraction module. The syntax parsing module, based on the sqlparse library, enhances SQL deep parsing capabilities and is responsible for parsing raw SQL statements into an enhanced abstract syntax tree (AST). The complexity assessment module generates a comprehensive complexity score for each SQL query based on the AST. If the complexity score exceeds a threshold, the SQL statement is input into the SQL decomposition module, which breaks the complex SQL statement into multiple smaller parts, reducing the difficulty of analysis and improving the efficiency of subsequent analysis and reasoning. The data lineage extraction module leverages the semantic understanding capabilities of the large-scale model to extract data lineage relationships.
[0066] Through the collaborative work of the syntax parsing module and the complexity assessment module, the present invention realizes the accurate parsing and complexity assessment of complex SQL statements, providing a solid foundation for subsequent data lineage extraction. This not only improves the accuracy of data lineage extraction, but also enhances the robustness and adaptability of the system. For SQL statements whose complexity scores exceed the preset threshold, the present invention decomposes them into multiple atomic SQL query units through the SQL disassembly module, which reduces the difficulty of analysis and improves the efficiency of subsequent analysis and reasoning. At the same time, the data table generation path and field-level mapping rules are extracted from the disassembled SQL query units to generate a preliminary lineage relationship map, which provides strong support for subsequent data governance and decision support. This design significantly improves the construction efficiency and accuracy of the data lineage knowledge map, and provides a strong guarantee for the effective management and utilization of data assets.
[0067] According to a method for constructing a large-model-based bloodline knowledge graph of civil affairs data provided by the present invention, a structured query language statement is parsed into an enhanced abstract syntax tree, specifically including: preprocessing the structured query language query string; performing preliminary parsing of the preprocessed structured query language query string to generate a basic abstract syntax tree; traversing each child node of the basic abstract syntax tree and performing different processing according to the type of the child node; and returning the enhanced abstract syntax tree after traversing and processing all child nodes.
[0068] Specifically, the present invention focuses on improving the sqlparse syntax parsing module design, including enhancing the refined processing of complex SQL structures, including common table expressions (CTEs), nested queries, and JOIN conditions.
[0069] For CTE, the main processing is to first perform enhanced WITH clause recognition: using a dynamic regular matching engine, design pattern: WITH\s+(RECURSIVE\s+)?([a-zA-Z_][a-zA-Z0-9_] )\s AS\s The .csh script, \(, supports parallel detection of recursive CTEs and multiple CTEs. A state machine tracks the nesting levels of parentheses, accurately locating the boundaries of CTE definitions and resolving misjudgments caused by traditional regular expression matching in complex clause nesting scenarios. A multi-level AST is then constructed, including independent syntax parsing of the extracted CTE definitions and the establishment of a hierarchical AST structure. Using tree pointer association technology, the CTE AST nodes are mounted as independent subtrees to the WITH branch of the main query AST, and an alias mapping table is established to implement semantic association. Finally, context-aware CTE reference resolution is performed: During the main query parsing phase, a context-aware processor is integrated to implement dynamic scope management. When a CTE alias is detected, the following process is executed:
[0070] a) Backtracking query AST WITH branch;
[0071] b) Verify the validity of the CTE scope (handle nested CTEs);
[0072] c) Inject the CTE logical table into the current parsing context;
[0073] d) Perform column-level dependency analysis (especially handling chained references between CTEs).
[0074] Nested queries are processed by recursively traversing the syntax tree. Specifically, the SQL statement's syntax tree is first traversed to identify the nested query clauses. The parsing function is then recursively called for each nested query to generate the corresponding AST. Finally, the results of the nested queries are integrated into the parsed results of the main query.
[0075] For JOIN conditions, the ON clause is parsed to extract the JOIN condition. The specific implementation steps are to traverse the syntax tree, find the JOIN keyword, and then extract the conditional expression in the ON clause. Finally, the conditional expression is parsed to extract the table name and column name. The pseudo code for SQL syntax parsing is as follows:
[0076] Algorithm 1: SQL syntax parsing to generate AST
[0077] Input: SQL query string Q
[0078] Output: AST containing CTE, nested queries, JOIN conditions, table information, and field information
[0079] 1. Procedure PARSE_SQL(Q)
[0080] 2. List context
[0081] 3. String preprocessed_sql ← PREPROCESS(Q) / / Preprocess SQL query string
[0082] 4. AST base_ast ← sqlparse.parse(preprocessed_sql) / / Parse SQL to generate syntax tree
[0083] 5. context ← {cte_defs: {}, temp_tables: , fields: {}}
[0084] 6. for each child in base_ast.tokens do
[0085] 7. if child is WITH clause then
[0086] 8. for cte_name, subquery in EXTRACT_CTE(child) do
[0087] 9. context.cte_defs[cte_name] ← PARSE_SQL(subquery)
[0088] 10. / / Recursively resolve CTE subquery
[0089] 11. context.temp_tables.add(cte_name)
[0090] 12. if child is Parenthesis then
[0091] 13. AST sub_ast ← PARSE_SQL(UNWRAP(child)) / / Recursively parse nested queries
[0092] 14. child.set_attr('nested', sub_ast)
[0093] 15. if child is JOIN keyword then
[0094] 16. String on_clause ← FIND_ON_CLAUSE(child) / / Find JOIN conditions
[0095] 17. child.set_attr('join_cond', PARSE_EXPR(on_clause)
[0096] 18. if IS_COLUMN(child) then
[0097] 19. String field_name ← GET_FIELD_NAME(child) / / Get the field name
[0098] 20. String table_name ← GET_TABLE_NAME(child) / / Get the table name
[0099] 21. context.fields.add({field_name, table_name, alias})
[0100] 22. end for
[0101] 23. return base_ast
[0102] By enhancing the refined processing of complex SQL structures, the present invention improves the accuracy and robustness of SQL parsing. In particular, when processing complex structures such as CTE, nested queries and JOIN conditions, it can accurately parse and retain contextual information, providing a reliable foundation for subsequent data lineage extraction. Through technical means such as recursive traversal of the syntax tree and dynamic scope management, the present invention achieves efficient parsing of complex SQL statements. This not only improves the efficiency of SQL parsing, but also reduces the consumption of computing resources, providing strong support for the construction of large-scale data lineage knowledge graphs. At the same time, the ability to accurately parse complex SQL statements also provides more accurate and comprehensive information for data governance and decision support.
[0103] According to a method for constructing a large-model-based bloodline knowledge graph of civil affairs data provided by the present invention, a comprehensive complexity score of a structured query language query is generated according to an enhanced abstract syntax tree, specifically including: extracting complexity features in the enhanced abstract syntax tree, normalizing each extracted complexity feature, and calculating the normalized value of each complexity feature; assigning a corresponding weight to each normalized complexity feature; and calculating a comprehensive complexity score based on the normalized value of each complexity feature and the corresponding weight.
[0104] Specifically, for the complexity evaluation module, the present invention calculates the comprehensive complexity score of each SQL query, evaluates its complexity, and thus determines whether the SQL statement needs to be input into the SQL disassembly module. The calculation formula of the complexity score is as follows:
[0105]
[0106] Where, Indicates the maximum and minimum normalization of feature i to ensure that the features are on the same scale, where Indicates the normalization calculation of the number of subqueries. Indicates normalization calculation of JOIN operands. Indicates normalized calculation of the number of CTEs Indicates normalized calculation of the maximum nesting depth. Represent the corresponding weight data respectively, and the complexity score is expressed as S.
[0107] By introducing a complexity assessment module and a specific complexity score calculation formula, this paper can accurately assess the complexity of SQL queries and provide a decision-making basis for subsequent SQL disassembly. This not only improves the efficiency and accuracy of processing complex SQL statements, but also helps optimize the construction process of the entire data lineage knowledge graph.
[0108] According to a method for constructing a large-model-based lineage knowledge graph of civil affairs data provided by the present invention, data table generation paths and field-level mapping rules are extracted from the disassembled structured query language query units to generate a preliminary lineage relationship graph, which specifically includes: inputting the disassembled multiple atomic structured query language query units into the large model; parsing each structured query language query unit to extract table-level lineage relationships and field-level lineage relationships; integrating the extracted table-level lineage relationships and field-level lineage relationships to generate a preliminary lineage relationship graph.
[0109] Specifically, for the SQL decomposition module, the present invention directly performs data lineage extraction on the SQL statement if the complexity score calculated by the complexity assessment module does not exceed the threshold. If the complexity score exceeds the threshold, the SQL statement is first decomposed. This decomposition primarily aims to simplify SQL syntax and improve the accuracy of data lineage extraction; it also reduces the length of subsequent input to the data lineage extraction module, avoiding errors caused by exceeding the input and output limits of large models. Subqueries, CTE definitions, JOIN conditions, and UNION components in complex SQL statements are written to create temporary tables, and the main query is then composed of these temporary tables.
[0110] The disassembly module is based on the open source version GLM-4-9B-Chat-1M in the pre-trained model GLM-4 series. This version supports a context length of 1M, approximately 2 million Chinese characters, and has strong semantic, mathematical, reasoning, and coding capabilities, meeting the requirements of general SQL data. The prompt used is as follows:
[0111] # Role: Senior SQL Development Engineer
[0112] ## Task Description
[0113] Please accurately parse complex SQL statements, carefully breaking down nested subqueries, CTEs, UNIONs, and JOINs, and converting them into multiple temporary table creation statements starting with 'CREATE TEMPORARY TABLE'. Please strictly adhere to the following requirements during the breakdown process:
[0114] (1) The output must consist entirely of SQL statements and must not contain any additional text, comments, or explanations.
[0115] (2) All generated SQL statements, whether repeated or not, must be listed one by one and must not be omitted.
[0116] (3) The disassembled SQL statements need to be arranged in the order of their original logical dependencies to ensure semantic integrity.
[0117] ## Please disassemble according to the following SQL statement. Only the SQL statement is output. No other text is allowed.
[0118] Figure 5 This example shows the input and output of SQL parsing using the prompt. To ensure the accuracy of the parsed SQL statements, the sqlparse library is used to perform preliminary SQL syntax verification.
[0119] The present invention adopts a method of fine-tuning a small-scale open source large language model to improve its ability in the field of data lineage extraction, but there is currently a lack of publicly labeled datasets for data lineage extraction tasks. Obtaining data lineage through manual labeling will consume a lot of manpower, financial resources and time costs. With the rapid iteration of large language models, their capabilities continue to improve, and they can complete more and more natural language processing tasks, including relationship extraction tasks and SQL generation tasks. The present invention uses a large language model with a large parameter scale in combination with an SQL structured template to generate SQL statements for a specific field, and extracts the data lineage in the generated SQL statements, and uses the extraction results as pseudo-data annotations to obtain sufficient "annotated" data. In the selection of large models, referring to the list of the classic Text2SQL evaluation dataset Spider, the top four are all GPT-4 enhanced with various black technologies. Therefore, the present invention uses GPT-4o as the data generator. In the selection of open source large models, the computing power and model evaluation results are comprehensively considered, and the GLM4-9b-chat model is selected for fine-tuning training. The specific implementation process is as follows Figure 6 Compared with large models, fine-tuning small models can not only significantly reduce computing resource requirements, but also provide certain advantages in ensuring data privacy and security.
[0120] To generate the dataset, this paper used GPT-4o for data generation and annotation. To ensure data diversity, 110 SQL language syntax combinations were pre-set, and SQL data was generated for each syntax combination. Furthermore, the large model outputs the confidence level of the data lineage analysis results, and only results with a confidence level greater than 0.6 are retained in the dataset. Through these steps, a SQL-DL dataset with both syntactic diversity and high confidence was constructed.
[0121] In the process of building the SQL-DL dataset, we first summarized 110 typical operation modes. We designed the following prompt words to generate SQL and extract the structure of data lineage:
[0122] # Role: Senior SQL Development Engineer
[0123] ## Task Description
[0124] Generate specific SQL statements based on the given SQL typical operation description. The business scenario is limited to the XX field, and the data lineage relationship is given in the following structure:
[0125] - The relationship between tables: expressed in the form of (source_table, target_table)
[0126] - The relationship between fields: expressed in the form of (source_field, target_field, transform_method), where source_field and target_field are the concatenation of the fields of source_table and target_table (e.g. `source_table.field`)
[0127] Finally, the confidence level of the data lineage relationship for the SQL statement is output (expressed as a percentage, ranging from 0% to 100%).
[0128] ## Enter a description
[0129] Please generate SQL statements, data lineage and confidence levels based on the following typical SQL operation descriptions:
[0130] ## Operation Description
[0131] Complex SQL queries and data target overwrite inserts with multi-table joins
[0132] ## Output format
[0133] 1. Generated SQL statement:
[0134] [Specific SQL statement]
[0135] 2. Data lineage:
[0136] - Blood relationship between tables:
[0137] - Example: (`orders`, `order_archive`)
[0138] - Blood relationship between fields:
[0139] - Example: (`orders.order_id`, `order_archive.order_id`, `directmapping`)
[0140] - (`orders.order_date`, `order_archive.order_date`, `Max`)
[0141] 3. Confidence of data lineage:
[0142] [Confidence Percentage].
[0143] In the task description, the representation of the blood relationship between tables and fields is limited. The diversity of the generated SQL syntax is ensured by replacing the operation description. The output format is also given to avoid redundant output from large models and facilitate subsequent processing.
[0144] To maintain a balanced data set, we generated 60 data items for each type of SQL operation description, for a total of 6,600 data items. These 6,600 data items were sorted by confidence, and the 10% with the lowest confidence were filtered out, ultimately retaining 5,940 data samples.
[0145] For model fine-tuning, lora is used to efficiently fine-tune the GLM4-9b-chat model based on the constructed SQL-DL dataset to enhance the model's ability in data lineage extraction tasks.
[0146] For data lineage analysis, this paper uses the fine-tuned GLM4-9b-chat model to perform data lineage analysis on SQL statements in the domain. The prompt used is as follows:
[0147] # Role: Senior SQL Development Engineer
[0148] ## Task Description
[0149] According to the given SQL statement, the data lineage relationship is given according to the following structure:
[0150] - The relationship between tables: expressed in the form of (source_table, target_table)
[0151] - The relationship between fields: expressed in the form of (source_field, target_field, transform_method), where source_field and target_field are the concatenation of the fields of source_table and target_table (e.g. `source_table.field`)
[0152] Finally, the confidence level of the data lineage relationship for the SQL statement is output (expressed as a percentage, ranging from 0% to 100%).
[0153] ## Enter a description
[0154] Please give the data lineage relationship based on the input SQL statement
[0155] ## Output format
[0156] - Blood relationship between tables:
[0157] - Example: (`orders`, `order_archive`)
[0158] - Blood relationship between fields:
[0159] - Example: (`orders.order_id`, `order_archive.order_id`, `direct mapping`)
[0160] - (`orders.order_date`, `order_archive.order_date`, `Max`)
[0161] By introducing an SQL decomposition module, this paper simplifies complex SQL statements, improves the accuracy of data lineage extraction, and reduces the length of subsequent input into the data lineage extraction module, avoiding errors caused by model input and output limitations. Furthermore, by utilizing a large model to generate and annotate a training dataset for data lineage extraction, this effectively addresses the issue of insufficient annotated data and improves the training effectiveness and performance of the data lineage extraction model.
[0162] S440. Based on the preliminary blood relationship map, traverse all physical source tables, eliminate virtual table redundancy and data node duplication, and generate an optimized blood relationship map.
[0163] According to the present invention, a method for constructing a blood relationship knowledge graph of civil affairs data based on a large model is provided. Based on a preliminary blood relationship graph, all physical source tables are traversed, virtual table redundancy and data node duplication are eliminated, and an optimized blood relationship graph is generated. Specifically, the method includes: constructing an adjacency list of tables and fields according to the preliminary blood relationship graph; marking virtual tables from the adjacency list and filtering physical source tables; initializing the queue, visited node set and merged mapping list required for breadth-first search, and adding the physical source table to the queue; traversing the adjacency list through breadth-first search, merging field mapping relationships, and updating the merged mapping list and visited node set; deduplicating the merged mapping list to remove redundant mapping relationships; and converting the deduplicated mapping list into an optimized blood relationship graph.
[0164] Specifically, the algorithm described below performs virtual table mapping fusion. The FUSE_RELATIONSHIPS function receives a table and field mapping list as input, first generating an adjacency table for the table and fields, and then filtering out physical source tables based on virtual flags. Specifically, a dynamic adjacency table structure is constructed using hash mapping to establish a two-level graph model. The outer hash layer stores vertex attributes using table names as keys, while the inner hash layer records field-level mapping relationships, forming a composite data structure of <source table, target table>-><field mapping set>. Next, a virtual node filtering mechanism is implemented. By parsing the virtual_flag flag in metadata, a whitelist of physical source tables is constructed, automatically excluding non-physical storage objects such as logical views and temporary tables. Next, a BFS graph traversal is performed on these physical source tables, using the physical source tables as the initial node set. A double-ended queue is used for hierarchical expansion, and related target tables are found and merged. Finally, redundant mappings are removed, and the merged table and field mappings are returned.
[0165] In terms of time complexity, the outer loop traverses all physical source tables, and the inner loop traverses the target table in the adjacency list of the current table, so the number of traversals is O(n m), where n is the number of physical source tables and m is the average number of target tables in the adjacency list of each table. In each inner loop, a merge operation is required, with a time complexity of O(k), where k is the number of fields to be merged. Therefore, the total time complexity is O(n m k). In terms of space complexity, the algorithm space is mainly composed of adjacency lists, virtual flags, visited node sets, queues, and merged mapping lists, so the space complexity is O(n + m + k), where n represents the size of the input data, including table mappings, field mappings, and adjacency relations.
[0166] Algorithm 2: Virtual table mapping fusion based on breadth-first search
[0167] Input: table and field mapping list table_map_list = { (source_table, target_table,fields_map), ...}
[0168] Output: merged table and field mapping list fused_map_list
[0169] 1. FUNCTION FUSE_RELATIONSHIPS(table_map_list):
[0170] 2. adjacency = CREATE_ADJACENCY(table_map_list) / / Create an adjacency list of tables and fields
[0171] 3. virtual_flags = MARK_VIRTUAL_TABLES(table_map_list) / / Virtual table mark
[0172] 4. queue = new QUEUE(), visited = new SET(), merged_map = new LIST()
[0173] 5. / / Get all physical source tables
[0174] 6. physical_sources = GET_PHYSICAL_SOURCES(table_map_list, virtual_flags)
[0175] 7. for source_table in physical_sources:
[0176] 8. if source_table not in visited:
[0177] 9. queue.push(source_table) / / Start breadth-first search
[0178] 10. visited.add(source_table)
[0179] 11. while queue is not empty:
[0180] 12. current_table = queue.pop()
[0181] 13. adj_list = adjacency.get(current_table, {}) / / Find the adjacency list of the current table
[0182] 14. for target_table, fields in adj_list.items():
[0183] 15. if target_table not in visited:
[0184] 16. merged_map = MERGE_MAPPING(merged_map,acdefghcurrent_table, target_table, fields)
[0185] 17. visited.add(target_table) / / Mark as visited
[0186] 18. queue.push(target_table) 19.
[0188] 20. final_map = REMOVE_REDUNDANCY(merged_map) / / Remove redundant mappings
[0189] 21.return final_map.
[0190] By introducing a virtual table mapping fusion algorithm based on breadth-first search, the present invention can efficiently process large-scale table and field mapping relationships, automatically exclude non-entity storage objects, accurately merge related target tables, and remove redundant mappings, thereby improving the accuracy and processing efficiency of virtual table mapping.
[0191] S450. Based on the optimized blood relationship map, identify common data nodes and merge field-level data blood relationships to generate the final civil affairs data blood relationship map.
[0192] According to the present invention, a method for constructing a civil affairs data lineage knowledge graph based on a large model is provided. Based on the optimized lineage relationship graph, public data nodes are identified and field-level data lineage is merged to generate the final civil affairs data lineage relationship graph. Specifically, the method includes: traversing all table nodes in the optimized lineage relationship graph to identify public data nodes; for each identified public data node, triggering the field-level alignment process to match field nodes; for field nodes that are successfully matched, merging their field-level data lineage and integrating the merged field-level data lineage into the optimized lineage relationship graph to generate the final civil affairs data lineage relationship graph.
[0193] Specifically, in the process of implementing data table alignment and field merging, the system first builds a table-field two-layer graph structure based on the global metadata warehouse, and establishes a fast mapping of table names to field sets through hash indexes. When performing table alignment, the algorithm traverses all blood relationship graphs to be merged, and identifies common data nodes based on the exact match of the table name hash value. If the fully qualified names of the two tables in the namespace (including database name, schema name, and table name) are exactly the same, they are determined to be the same logical entity. At this time, the system triggers the field-level alignment process: for the table nodes that have been matched successfully, its field subgraph is deeply traversed, and a two-way matching strategy (forward traversal of the source table fields and reverse scanning of the target table fields) is used to find field nodes with the same name. Predefined naming normalization rules (such as unified capitalization and removal of special symbols) are used to eliminate misjudgments caused by differences in naming styles. Finally, the naming of the node fields is compared to align the data.
[0194] By building a two-layer graph structure and hash index, as well as employing a bidirectional matching strategy and naming normalization rules, the present invention can efficiently and accurately align data tables and merge fields. This not only improves the efficiency and accuracy of data processing, but also effectively eliminates misjudgments caused by differences in naming styles, thereby improving the quality and reliability of data integration.
[0195] The following are examples of the present invention in the application of civil affairs data:
[0196] In the application scenario of subsistence allowance fund distribution, by constructing a lineage map of subsistence allowance fund distribution data, we can clearly track the source data, intermediate tables, and final distribution records involved in the fund calculation process, quickly locate the processing links of abnormal data, and assist in auditing and error correction. Combined with the semantic parsing capabilities of large language models for complex SQL statements, we achieve field-level lineage association, preventing the failure of traditional rule matching under dynamic policy adjustments.
[0197] In the application scenario of distributing basic living subsidies to orphans, the subsidy calculation SQL (including subqueries and JOINs) is parsed to extract field-level mapping relationships between the source table (orphan archive table), the intermediate table (subsidy standard table), and the target table (disbursement record table). Temporary views generated during the process (such as the "Orphans Awaiting Review List") are merged to eliminate redundant nodes. A genealogy map is generated to visually display the entire chain of subsidy data, from entry and review to disbursement, assisting in identifying disbursement delays or abnormal amounts.
[0198] It can be seen that the present invention can significantly improve the automation level of civil affairs data management and provide reliable data support for precise rescue, dynamic monitoring and other services.
[0199] The following describes the construction device of the civil affairs data bloodline knowledge graph based on the big model provided by the present invention. The construction device of the civil affairs data bloodline knowledge graph based on the big model described below and the construction method of the civil affairs data bloodline knowledge graph based on the big model described above can be referenced to each other.
[0200] like Figure 7 The present invention provides a large-scale model-based civil affairs data lineage knowledge graph construction device, including:
[0201] A basic entity network construction module 710 is used to extract metadata features of civil affairs data and construct a basic entity network, which includes data tables and fields.
[0202] The data table structure relationship model generation module 720 is used to parse the data table structure based on the basic entity network and generate a data table structure relationship model;
[0203] A preliminary blood relationship map generation module 730 is used to parse complex structured query language statements based on the data table structure relationship model through a large model, identify field-level blood relationship associations, and generate a preliminary blood relationship map;
[0204] An optimized blood relationship map generation module 740 is used to traverse all physical source tables based on the preliminary blood relationship map, eliminate virtual table redundancy and data node duplication, and generate an optimized blood relationship map;
[0205] The final blood relationship map generation module 750 is used to identify common data nodes and merge field-level data blood relationships based on the optimized blood relationship map to generate the final civil affairs data blood relationship map.
[0206] Figure 8 An example of a physical structure diagram of an electronic device is shown below. Figure 8As shown, the electronic device may include: a processor 810 , a communication interface 820 , a memory 830 and a communication bus 840 , wherein the processor 810 , the communication interface 820 and the memory 830 communicate with each other via the communication bus 840 . The processor 810 can call the logical instructions in the memory 830 to execute a method for constructing a civil affairs data bloodline knowledge graph based on a large model, the method including: constructing a basic entity network based on metadata feature extraction of civil affairs data, the basic entity network including data tables and fields; parsing the data table structure based on the basic entity network to generate a data table structure relationship model; parsing complex structured query language statements through the large model based on the data table structure relationship model, identifying field-level bloodline associations, and generating a preliminary bloodline relationship graph; based on the preliminary bloodline relationship graph, traversing all physical source tables, eliminating virtual table redundancy and data node duplication, and generating an optimized bloodline relationship graph; based on the optimized bloodline relationship graph, identifying common data nodes and merging field-level data bloodlines to generate a final civil affairs data bloodline relationship graph.
[0207] Furthermore, the logic instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0208] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the construction method of the civil affairs data bloodline knowledge graph based on the big model provided by the above methods. The method includes: constructing a basic entity network based on metadata feature extraction of civil affairs data, and the basic entity network includes data tables and fields; based on the basic entity network, parsing the data table structure to generate a data table structure relationship model; based on the data table structure relationship model, parsing complex structured query language statements through the big model, identifying field-level bloodline associations, and generating a preliminary bloodline relationship map; based on the preliminary bloodline relationship map, traversing all physical source tables, eliminating virtual table redundancy and data node duplication, and generating an optimized bloodline relationship map; based on the optimized bloodline relationship map, identifying common data nodes and merging field-level data bloodlines to generate the final civil affairs data bloodline relationship map.
[0209] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for constructing a large-model-based civil affairs data bloodline knowledge graph provided by the above-mentioned methods, the method comprising: constructing a basic entity network based on metadata feature extraction of civil affairs data, the basic entity network comprising data tables and fields; parsing the data table structure based on the basic entity network to generate a data table structure relationship model; parsing complex structured query language statements through a large model based on the data table structure relationship model, identifying field-level bloodline associations, and generating a preliminary bloodline relationship graph; based on the preliminary bloodline relationship graph, traversing all physical source tables, eliminating virtual table redundancy and data node duplication, and generating an optimized bloodline relationship graph; based on the optimized bloodline relationship graph, identifying common data nodes and merging field-level data bloodlines to generate a final civil affairs data bloodline relationship graph.
[0210] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. That is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0211] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods of each embodiment or certain portions of the embodiments.
[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for constructing a bloodline knowledge graph of civil affairs data based on a large model, characterized by: include: Extract metadata features of civil affairs data to build a basic entity network, the basic entity network including data tables and fields; Based on the basic entity network, the data table structure is parsed to generate a data table structure relationship model; Based on the data table structure relationship model, the complex structured query language statements are parsed through the large model to identify field-level blood relationship and generate a preliminary blood relationship map; Based on the preliminary blood relationship map, all physical source tables are traversed to eliminate virtual table redundancy and data node duplication, thereby generating an optimized blood relationship map; Based on the optimized blood relationship map, identify common data nodes and merge field-level data blood relationships to generate the final civil affairs data blood relationship map; The method of parsing complex structured query language statements through a large model, identifying field-level blood relationship associations, and generating a preliminary blood relationship map specifically includes: parsing the structured query language statement into an enhanced abstract syntax tree; generating a comprehensive complexity score of the structured query language query based on the enhanced abstract syntax tree; if the complexity score exceeds a preset threshold, decomposing the complex structured query language statement into multiple atomic structured query language query units through the large model; extracting data table generation paths and field-level mapping rules from the decomposed structured query language query units to generate a preliminary blood relationship map; The parsing of the structured query language statement into the enhanced abstract syntax tree specifically includes: preprocessing the structured query language query string; performing preliminary parsing on the preprocessed structured query language query string to generate a basic abstract syntax tree; traversing each child node of the basic abstract syntax tree and performing different processing according to the type of the child node; and returning the enhanced abstract syntax tree after traversing and processing all child nodes.
2. The method for constructing a large-scale model-based civil affairs data lineage knowledge graph according to claim 1 is characterized in that: Generating a comprehensive complexity score of the structured query language query based on the enhanced abstract syntax tree specifically includes: Extracting complexity features from the enhanced abstract syntax tree, normalizing each extracted complexity feature, and calculating the normalized value of each complexity feature; Assign corresponding weights to each normalized complexity feature; The comprehensive complexity score is calculated based on the normalized value of each complexity feature and the corresponding weight.
3. The method for constructing a large-scale model-based civil affairs data lineage knowledge graph according to claim 1 is characterized in that: The process of extracting the data table generation path and field-level mapping rules from the disassembled structured query language query unit to generate a preliminary blood relationship map specifically includes: Input the disassembled multiple atomic structured query language query units into the large model; Parse each structured query language query unit to extract table-level and field-level lineage relationships; The extracted table-level kinship relationship and the field-level kinship relationship are integrated to generate a preliminary kinship relationship map.
4. The method for constructing a large-scale model-based civil affairs data lineage knowledge graph according to claim 1 is characterized in that: Based on the preliminary blood relationship map, all physical source tables are traversed to eliminate virtual table redundancy and data node duplication to generate an optimized blood relationship map, which specifically includes: Based on the preliminary blood relationship map, construct the adjacency list of tables and fields; Marking the virtual table from the adjacency list and filtering the physical source table; Initialize the queue, visited node set, and merged mapping list required for breadth-first search, and add the physical source table to the queue; Traverse the adjacency list through breadth-first search, merge field mapping relationships, and update the merged mapping list and visited node set; De-duplicate the merged mapping list to remove redundant mapping relationships; Convert the deduplicated mapping list into an optimized blood relationship graph.
5. The method for constructing a large-scale model-based civil affairs data lineage knowledge graph according to claim 1 is characterized in that: Based on the optimized blood relationship map, identifying common data nodes and merging field-level data blood relationships to generate the final civil affairs data blood relationship map specifically includes: Traverse all table nodes in the optimized blood relationship graph and identify common data nodes; For each identified common data node, the field-level alignment process is triggered to perform field node matching; For successfully matched field nodes, merge their field-level data lineage; Integrate the merged field-level data lineage into the optimized lineage relationship map to generate the final civil affairs data lineage relationship map.
6. A device for constructing a bloodline knowledge graph of civil affairs data based on a large model, characterized in that: include: A basic entity network construction module is used to extract metadata features of civil affairs data to construct a basic entity network, the basic entity network including data tables and fields; A data table structure relationship model generation module is used to parse the data table structure based on the basic entity network and generate a data table structure relationship model; A preliminary blood relationship map generation module is used to parse complex structured query language statements through a large model based on the data table structure relationship model, identify field-level blood relationship associations, and generate a preliminary blood relationship map; An optimized blood relationship map generation module is used to traverse all physical source tables based on the preliminary blood relationship map, eliminate virtual table redundancy and data node duplication, and generate an optimized blood relationship map; The final blood relationship map generation module is used to identify common data nodes and merge field-level data blood relationships based on the optimized blood relationship map to generate the final civil affairs data blood relationship map; The method of parsing complex structured query language statements through a large model, identifying field-level blood relationship associations, and generating a preliminary blood relationship map specifically includes: parsing the structured query language statement into an enhanced abstract syntax tree; generating a comprehensive complexity score of the structured query language query based on the enhanced abstract syntax tree; if the complexity score exceeds a preset threshold, decomposing the complex structured query language statement into multiple atomic structured query language query units through the large model; extracting data table generation paths and field-level mapping rules from the decomposed structured query language query units to generate a preliminary blood relationship map; The parsing of the structured query language statement into the enhanced abstract syntax tree specifically includes: preprocessing the structured query language query string; performing preliminary parsing on the preprocessed structured query language query string to generate a basic abstract syntax tree; traversing each child node of the basic abstract syntax tree and performing different processing according to the type of the child node; and returning the enhanced abstract syntax tree after traversing and processing all child nodes.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the method for constructing a large-model-based civil affairs data bloodline knowledge graph as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method for constructing a large-model-based bloodline knowledge graph of civil affairs data as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Data analysis-based data blood relationship graph construction method and related equipment
CN116186174A
Data blood relationship analysis method based on SQL (Structured Query Language) analysis
CN118916347A
SQL metadata processing method based on large model, SQL statement generation method, electronic equipment and data medium table
CN119537376A