Water conservancy data quality inspection rule automatic generation method and system based on large language model
By automatically analyzing water conservancy industry documents and generating data quality inspection rules based on a large language model, the problems of high manual participation, low efficiency and insufficient flexibility in the existing technology are solved, and efficient, accurate and flexible water conservancy data quality inspection rules are achieved.
Patent Information
- Application Number
- CN202510394097.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing technology has problems such as high manual participation, low efficiency, insufficient flexibility, and disconnection between industry knowledge and technology in the design of water conservancy data quality inspection rules, resulting in a low level of data quality management.
Using a method based on a large language model, we automatically parse water conservancy industry documents, extract table structure information, and automatically generate data quality inspection rules covering multiple data quality dimensions to convert them into standardized SQL statements.
It significantly improves the efficiency of rule design, reduces manual intervention, eliminates the knowledge gap between industry experts and database technicians, ensures the accuracy and consistency of quality inspection rules, and meets the flexibility and dynamic adaptability requirements of data quality inspection rules in complex water conservancy data governance scenarios.
Smart Images

Figure CN120067090A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of water conservancy data processing, and particularly relates to a method and system for automatically generating water conservancy data quality inspection rules based on a large language model. Background Art
[0002] In modern data governance, data quality management plays a key role in ensuring the reliability of business decisions and data accuracy. However, the design and generation of current water conservancy data quality inspection rules still face many challenges, mainly reflected in the following aspects: 1. Limitations of manual design methods Traditional methods rely on the close cooperation between industry experts and programmers. Industry experts are usually familiar with business logic but lack a background in database technology and are difficult to transform business data logic into specific and executable quality inspection rules; while programmers are proficient in database technology but often have insufficient understanding of the business requirements of specific business fields (such as the water conservancy industry). Due to the large differences in the knowledge backgrounds of both parties and frequent cross - professional communication, information transmission is prone to deviation, resulting in omissions or inaccuracies in the designed quality inspection rules, and at the same time making the overall rule design process inefficient.
[0003] 2. Deficiencies of tool - assisted methods Some existing tools assist in generating data quality inspection rules through a templatized approach. The basic idea is to extract field information from database table design documents and generate corresponding rules based on preset templates. Although this method improves efficiency to a certain extent, the templates are fixed and the rules are rigid, making it difficult to adapt to complex business scenarios and dynamically changing table structures. The preset templates cannot fully reflect the specific requirements of each industry and scenario, resulting in obvious limitations in the flexibility and applicability of the generated rules.
[0004] 3. Disconnection between industry knowledge and technical implementation In the current design process, the knowledge isolation between industry experts and programmers is relatively serious. Industry experts are difficult to transform business requirements into technical rules, and programmers are difficult to accurately capture the subtle differences in business logic during the technical implementation process. This disconnection makes the finally generated data quality inspection rules often unable to fully cover business requirements, thereby affecting the overall effect of data governance.
[0005] In summary, the existing technologies have problems such as high manual participation, low efficiency, lack of flexibility, and disconnection between industry knowledge and technical implementation in the design of water conservancy data quality inspection rules. These problems seriously restrict the improvement of water conservancy data quality management level and the realization of large - scale data governance requirements. Summary of the Invention
[0006] The objective of the present invention is to address the deficiencies existing in the above-mentioned background technology and provide a method and system for automatically generating water conservancy data quality inspection rules based on large language models. It not only significantly improves the automation level and generation efficiency, but also effectively solves the problem of the disconnection between industry knowledge and technical implementation in traditional methods through a dynamic optimization mechanism, meeting the requirements for the accuracy, flexibility, and dynamic adaptability of data quality inspection rules in complex water conservancy data governance scenarios.
[0007] The technical solution adopted by the present invention is as follows: A method for automatically generating water conservancy data quality inspection rules based on large language models includes the following steps: Utilize a large language model to parse documents in the water conservancy industry and extract table structure information; Automatically generate data quality inspection rules for the extracted table structure information, and the data quality inspection rules cover multiple data quality dimensions; Convert the data quality inspection rules into standardized SQL statements.
[0008] In the above technical solution, the document is a Word document in the water conservancy industry, the table structure information includes table name, field identifier, and field description, and the parsing process standardizes the table structure information into JSON format.
[0009] In the above technical solution, the parsing process further extracts the chapter numbers in the document and the code table information of the fields.
[0010] In the above technical solution, the data quality dimensions covered by the automatically generated data quality inspection rules at least include uniqueness, integrity, standardization, validity, and accuracy.
[0011] In the above technical solution, for time series data, the automatically generated data quality inspection rules include dynamic change inspection rules, and the dynamic change inspection rules detect abnormal changes by analyzing the data change rate within a predetermined time window and the fluctuation range of the maximum and minimum values of the data.
[0012] In the above technical solution, for non-time series data, the automatically generated data quality inspection rules include threshold inspection rules, and the threshold inspection rules automatically set a reasonable value range in combination with the characteristics of water conservancy industry data.
[0013] In the above technical solution, the method further includes providing an interactive rule optimization function to enable users to adjust the logic of the data quality inspection rules according to actual needs.
[0014] In the above technical solution, the method further includes using an SQL verification module to verify the generated data quality inspection rules to detect the accuracy and coverage rate of the rules.
[0015] In the above technical solution, the method further includes storing the generated data quality inspection rules in a database and outputting a list of the data quality inspection rules in JSON format or table form.
[0016] The present invention provides a system for automatically generating water conservancy data quality inspection rules based on a large language model, including: A document parsing module, configured to parse water conservancy industry documents using a large language model, extract table structure information, and generate corresponding JSON structured data; A rule generation module, configured to generate data quality inspection rules based on the JSON structured data, where the rules cover multiple data quality dimensions, and convert the generated rules into standardized SQL statements; A rule optimization module, configured to provide an interactive interface for users to optimize and adjust the data quality inspection rules; A rule verification module, configured to verify the accuracy and coverage rate of the data quality inspection rules using SQL; And a storage and output module, configured to store the data quality inspection rules in a database and output them in JSON format or table form.
[0017] The beneficial effects of the present invention are as follows: The present invention uses a large language model to automatically parse table structure information from water conservancy industry documents and automatically generate data quality inspection rules covering multiple dimensions and standardized SQL statements. This not only greatly improves the rule design efficiency, reduces manual intervention, but also effectively eliminates the knowledge gap between industry experts and database technicians, ensuring the accuracy and consistency of the quality inspection rule generation process.
[0018] Furthermore, by using a Word document in the water conservancy industry as the data source and standardizing the extracted table structure information (including table names, field identifiers, and field descriptions) into JSON format, the present invention can ensure the unity and structuring of information extraction. This provides a reliable and standardized data basis for subsequent automatic generation of data quality rules and reduces the error rate in the data preprocessing process.
[0019] Furthermore, on the basis of the standardized JSON data, the present invention further extracts chapter numbers and code table information of fields, so that the generated quality inspection rules can fully consider the information described in detail in the document. This makes the rules able to reflect a more comprehensive business background in the document and helps to generate more refined and business-scenario-specific inspection rules.
[0020] Furthermore, the data quality inspection rules automatically generated by the present invention cover multiple data quality dimensions such as uniqueness, integrity, standardization, validity, and accuracy, ensuring comprehensive monitoring and management of data quality from multiple perspectives. This can not only detect data anomalies but also improve the comprehensiveness and reliability of data governance, providing solid data support for business decisions.
[0021] Furthermore, for time-series data (such as water level, flow rate, etc.), the dynamic change inspection rules generated by the present invention can detect abnormal fluctuations in a timely manner by analyzing the data change rate and the fluctuations between the maximum and minimum values within a predetermined time window. This helps to achieve real-time monitoring and early warning, ensuring that the data remains within a reasonable range during the dynamic change process, thereby improving the real-time and accuracy of data monitoring.
[0022] Furthermore, the threshold inspection rules generated by the present invention for non-time-series data automatically set a reasonable value range in combination with the characteristics of the water conservancy industry, effectively filtering and excluding abnormal data. This can not only adapt to the data characteristics in different business scenarios but also ensure that the generated rules highly match the business requirements of actual water conservancy data, enhancing the scientificity and applicability of data quality management.
[0023] Furthermore, the present invention provides an interactive rule optimization function, enabling users to provide feedback and adjustment on the data quality inspection rules automatically generated according to actual needs. This dynamic optimization mechanism makes up for the deficiencies between manual design and automatic generation, not only reducing the risk of errors caused by understanding deviations but also enabling the rules to be continuously iterated and improved to meet the changing needs in complex water conservancy data governance scenarios.
[0024] Furthermore, the present invention uses an SQL verification module to automatically verify the generated SQL rules, effectively detecting the accuracy and coverage of the generated rules. This verification mechanism ensures the correctness of the automatically generated rules in terms of syntax and logic, reducing the risk of data quality problems caused by incorrect rule design and improving the robustness and reliability of the overall data governance system.
[0025] Furthermore, the present invention stores the generated quality inspection rules in a structured manner in the database and supports output in JSON format or table form, facilitating review, management, and subsequent maintenance. This not only realizes the standardization of rule management but also facilitates cross-departmental communication and tracking of data quality inspection results, thus greatly improving the management efficiency of large-scale data governance. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0027] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, which is convenient for clearly understanding the present invention, but they do not constitute a limitation to the present invention.
[0028] Embodiment 1
[0029] As Figure 1 shown, the present invention provides a method for automatically generating water conservancy data quality inspection rules based on a large language model. The method includes the following steps: a) Extract table structure information from Word documents in the water conservancy industry. The table structure information includes table name, field identifier, field description, chapter number, and field code table information, and convert the information into a standardized JSON format; b) Generate a prompt in a predefined format based on the table structure information in the standardized JSON format. The prompt in the predefined format contains all instructions and parameters required for generating data quality inspection rules and incorporates the knowledge of water conservancy domain experts; c) Submit the generated prompt in the predefined format to the large language model to generate initial data quality inspection rules. The rules cover inspection types such as uniqueness, integrity, normativity, validity, and accuracy. Among them, dynamic change inspection rules are generated for time series data (such as water level, flow rate), including 48-hour data change rate analysis and maximum and minimum value fluctuation detection, and threshold inspection rules are generated for non-time series data to automatically set a reasonable value range. The generated rules are directly converted into standardized SQL statements; d) Use the SQL verification module to automatically verify the generated standardized SQL statements to ensure that the SQL statements meet the business requirements both syntactically and logically; e) Provide an interactive interface to enable users to review and provide feedback on the initially generated rules that fail verification or pass verification but have additional set requirements, and adjust the prompt in the predefined format based on the user feedback; f) Repeat steps (c - e) to form a loop process of generation - feedback - optimization - regeneration until the generated data quality inspection rules meet the predetermined business requirements; g) Store the finally generated data quality inspection rules, including field information, rule description, and corresponding SQL statements, in the database and support output in JSON or table format.
[0030] Specifically, step a) further includes: Convert the Word document (including table structure design and field descriptions) to Markdown format, and split the Markdown content based on the heading level symbols to facilitate the extraction of table structure information and field descriptions related to water conservancy data. Then convert the extracted data to JSON format for subsequent processing.
[0031] The standardized JSON format includes the following fields: Table name, table identifier, chapter number; The information of each field in the table, including field name, field identifier, data type and length, whether null values are allowed, measurement unit, primary key sequence number, and field description (where the field description contains code table information).
[0032] The following is an example of the model Prompt for converting document slices to table JSON: Please return the result strictly in the following JSON format without including any other text.
[0033] If it is not a table structure design document, return: {{ "is_table_structure": false, "message": "Non-table structure design document" }} If it is a table structure design document, return: {{ "is_table_structure": true, "tables": {{ "table_name": "Table name", "table_identifier": "Table identifier", "chapter_ID": "Chapter number", "fields": {{ "field_name": "Field name", "field_identifier": "Field identifier", "field_type": "Type and length", "nullable": "Whether null values are allowed (Y / N)", "unit": "Measurement unit", "key": "Primary key sequence number", "field_description": "Field description (including code table information)" }} }} }} Precautions: 1. When extracting the field description, if the field involves a code table, the code table needs to be converted into a Chinese description, and all the code table information should be extracted completely.
[0034] 2. The chapter number should be the chapter number before the table name. For example, for "6.1.1 Precipitation Monitoring Table", the chapter number is "Section 6.1.1".
[0035] 3. Characteristics of the table structure design document: - Include the table name and table identifier - Include the field definition table - Include the field description section - Include the code table information and threshold description Specifically, in step b) of generating the prompt in the predefined format, the prompt in the predefined format clearly requires the large language model to generate data quality inspection rules based on the data characteristics of the water conservancy industry, and requires that the rules cover five major types: uniqueness, integrity, normativity, validity, and accuracy when generated. Set a large amount of expert knowledge according to the data types for which quality inspection rules are needed, including time series rules and non-time series rules. The rule content includes field name, inspection type, rule description, and SQL implementation.
[0036] Specifically, in step c): For time series data, the generated data quality inspection rules include dynamic change inspection rules, which detect abnormal situations by analyzing the data change rate within 48 hours and the fluctuations between the maximum value and the minimum value.
[0037] For non-time series data, the generated data quality inspection rules include threshold inspection rules, and the threshold inspection rules automatically set a reasonable value range in combination with water conservancy characteristics (such as water level, flow rate, etc.), where the threshold is adjusted according to the measurement unit of the field in the JSON.
[0038] The following is an example of a prompt for converting table JSON information to JSON quality inspection (the prompts for different types of water conservancy data need to be refined and revised): Please generate one or more data quality inspection rules for the fields in the table structure (no inspection rules need to be set for the primary key fields), and return the results in the following format: ```json { "rules": {{ "field_name": "Field Name", "field_identifier": "Field Identifier", "checks": {{ "check_subject": "uniqueness", "description": "Check whether the field is unique.", "sql": "SELECT...FROM {table_identifier} WHERE {field_identifier} =?" }}, {{ "check_subject": "completeness", "description": "Check whether the field is empty.", "sql": "SELECT...FROM {table_identifier} WHERE {field_identifier} IS NULL" }}, {{ "check_subject": "standardization", "description": "Check whether the field format meets the requirements.", "sql": "SELECT...FROM {table_identifier} WHERE {field_identifier} NOT LIKE?" }} }} } Requirements: 1. Check Type: The field check type includes one of the following five: uniqueness, completeness, standardization, validity, accuracy. Each check can have multiple rules.
[0039] 2. Temporal Rules: For numerical monitoring index fields whose field names contain "water level", "flow", "precipitation", etc., a time series rule (check_subject is "validity") must be included.
[0040] Time series rule 1: Analyze the latest 48-hour records and changes, and calculate the data change rate for the previous and next 24 hours.
[0041] Time series rule 2: Check the difference between the maximum and minimum values within 48 hours. Values exceeding the preset common sense are regarded as abnormal.
[0042] 3. Non-time series rule: If the field description contains thresholds, code tables, or specific rules, they must be used as inspection constraints and reflected in the SQL.
[0043] Code table judgment 1: If the field description 'description' contains code constraint descriptions, they must be used as rule constraints, and the content referring to the code should be retained in the 'description' and reflected in the SQL statement. (For example, confidence description: When there are multiple sources for the same type of monitoring sites, the credibility of the site data. 1 represents the most credible site, 0 represents an expired site, and 2 represents others. The description should be retained in full).
[0044] Code table judgment 2: Only detect non-empty records. Pay attention to optimizing the quality inspection rule sql. For example, check whether the value of the flow measurement method field is valid, and it must be one of the values specified in the code table.
[0045] Threshold judgment 1: For the "validity check" of monitoring data, as a water conservancy expert, based on the geographical, climatic, and hydrological characteristics of XX Province as much as possible (such as the water level cannot be higher than 2000 meters above sea level in XX Province and cannot be less than -30 meters; the flow cannot be greater than 80000 cubic meters per second and cannot be less than -80000 cubic meters per second, where the negative number represents the reverse flow; the maximum catchment area of the hydrological station does not exceed 360000 square kilometers; the value range of the cross-sectional water area is limited between 0 and 25000 square meters). Set specific thresholds that conform to common sense for numerical monitoring data items to exclude outlier data. However, pay attention to the measurement unit of this field in the json, and the set value range should be consistent with the measurement unit "unit".
[0046] Threshold judgment 2: If the field description 'description' contains threshold constraints, they must be used as rule constraints, and the specific threshold constraint content should be described and reflected in the SQL statement.
[0047] Do not easily set null value judgment: For those where {nullable} is not N, do not casually set null value judgment rules (such as monitoring methods and monitoring indicators), unless there are several fields that must exist simultaneously, in which case a related null value verification rule can be set.
[0048] Quality inspection rules that water conservancy experts believe need to be set.
[0049] Collection and creation time fields similar to "TM" and "create_time" cannot exceed the current time.
[0050] SQL requirement: Field identifiers (field_identifier) and table identifiers (table_identifier) must be referenced in the SQL statement, and they should be in lowercase English.
[0051] Specifically, in step d), the initially generated data quality inspection rules will be stored in the database according to information such as table name, field name, and inspection type. The storage format includes table name, table identifier, field name, field identifier, quality inspection type, Chinese description of quality inspection, and SQL statement.
[0052] The following is an example code for storing rules in the database.
[0053] def store_rules_to_db(rules, connection): cursor = connection.cursor() for rule in rules: cursor.execute(""" INSERT INTO data_quality_rules (field_name, check_subject, description, sql) VALUES (%s, %s, %s, %s); """, (rule['field_name'], rule['check_subject'], rule['description'], rule['sql'])) connection.commit() Then, the accuracy of the SQL verification rules will be automatically called to ensure that the logic meets the business requirements. The system will automatically execute the corresponding SQL statements to calculate the error ratio under specific data quality inspection rules. For the currently inspected data quality type or metric, for each specific data quality dimension (such as uniqueness, integrity, standardization, validity, or accuracy) set rules, calculate their error rates respectively. For example, for integrity checks, the system will generate SQL statements to count the proportion of missing values, thereby obtaining the data error rate under this rule.
[0054] Specifically, the step e) of providing an interactive interface enables the user to provide feedback on the initially generated data quality inspection rules based on the error rate calculated in step d), and modify the generated prompt in a predefined format accordingly, thereby realizing the dynamic optimization of the prompt in the predefined format, and making the generation rules reach the predetermined accuracy and business requirements through multiple iterations.
[0055] Specifically, in the step g), the generated data quality inspection rules are structured and sorted according to information such as field name, inspection type, rule description, and SQL statement before storage, and stored in the database. At the same time, it supports outputting the complete rule list in JSON format or in tabular forms such as Excel for water conservancy experts or other users to review the automatically generated rule list, which is convenient for review and management.
[0056] Embodiment 2
[0057] The present invention provides a system for automatically generating water conservancy data quality inspection rules based on a large language model. The system includes: A document parsing module for converting a Word document in the water conservancy industry into Markdown format and extracting table structure information to generate standardized JSON format data; A Prompt generation module for generating a Prompt containing water conservancy expert knowledge based on the JSON format data; A large language model module for receiving the generated Prompt and automatically generating data quality inspection rules, including corresponding SQL statements; An interactive interface module for realizing user feedback on the generated rules and dynamically optimizing the Prompt, thereby forming a cycle process of generation - feedback - optimization - regeneration; An SQL verification module for automatically verifying the correctness of the generated SQL statements; and A storage and output module for structurally storing the generated data quality inspection rules in the database and supporting output in JSON or tabular form.
[0058] Specifically, the document parsing module includes: A word2md document conversion module for converting the paragraphs and table content in the Word document into Markdown format, supporting automatic parsing of the title hierarchy; An md_splitter content splitting module for splitting the Markdown content according to the title hierarchy symbols (such as #) and extracting relevant table structure information; The md2json JSON generation module is used to parse Markdown content through a large language model, automatically extract information such as table names, field identifiers, field types, and whether they can be null, and generate it in JSON format.
[0059] Specifically, the Prompt generation module includes an input and preprocessing module, an expert knowledge integration and template design module, and a Prompt generation module.
[0060] The data source of the input and preprocessing module is: receiving the standardized JSON format data generated by the document parsing module, which contains the table name, field identifier, field description, chapter number, and field code table information extracted from the water conservancy industry documents.
[0061] The preprocessing process of the input and preprocessing module includes: preprocessing the JSON data, sorting out and summarizing the basic information of each field, and providing a structured basis for generating the Prompt.
[0062] The functions of the expert knowledge integration and template design module include expert knowledge integration and template design.
[0063] Among them, expert knowledge integration is to combine the knowledge provided by water conservancy industry experts to clarify the requirements for each data quality inspection type, specifically including: Inspection types: uniqueness, integrity, standardization, validity, and accuracy.
[0064] Time series rules: For time series data such as water level, flow rate, and precipitation, it is required to generate dynamic change inspection rules, such as the data change rate within 48 hours and abnormal fluctuation detection.
[0065] Non-time series rules: For other data, generate rules through threshold inspection, automatically set a reasonable range based on water conservancy characteristics, and the threshold is dynamically adjusted according to the measurement unit of the field in the JSON.
[0066] Template design is to design a set of prompt templates in a predefined format, which clearly requires the large language model to generate data quality inspection rules based on the above expert knowledge. The template content includes: Task description: The generated rules are required to cover field names, inspection types, rule descriptions, and SQL implementations.
[0067] Data background: Includes table structure data, specific field information, and water conservancy business scenario background.
[0068] Rule logic requirements: Describe in detail the generation conditions of time series rules (such as 48-hour change rate, abnormal fluctuation) and non-time series rules (such as threshold detection).
[0069] The Prompt generation module is used to automatically generate Prompts: According to the template and the preprocessed JSON data, it automatically assembles a Prompt text containing water conservancy expert knowledge, providing clear instructions for the subsequent large language model to generate rules.
[0070] Example description: For example, the Prompt will clearly state "Please generate data quality inspection rules covering five inspection types: uniqueness, integrity, standardization, validity, and accuracy. Among them, generate rules for the change rate within 48 hours and abnormal fluctuation detection for time-series data; generate threshold inspection rules for non-time-series data, and the threshold is adjusted according to the measurement unit of the data."
[0071] Specifically, the large language model module includes a Prompt receiving and parsing module, a rule generation processing module, and an output and feedback module. Among them, the Prompt receiving and parsing module: Receives the prompt in a predefined format output by the Prompt generation module. This prompt contains detailed task requirements, data background, and rule generation logic. It uses the large language model to deeply parse the Prompt to understand the water conservancy business background, requirements for each inspection type, and details of time-series and non-time-series rule generation.
[0072] The rule generation processing module automatically generates data quality inspection rules based on the parsing results, specifically including: Inspection type rules: Generate inspection rules covering aspects such as uniqueness, integrity, standardization, validity, and accuracy.
[0073] Time-series rule generation: Generate dynamic change rules for fields including water level, flow rate, precipitation, etc., and detect anomalies by analyzing the data change rate within 48 hours and the fluctuations of the maximum and minimum values.
[0074] Non-time-series rule generation: For other data fields, generate threshold inspection rules and automatically set a reasonable value range that conforms to water conservancy characteristics and is adjusted according to the measurement unit.
[0075] Rule formatting: Organize the generated rules into a standardized format, usually structured JSON data, and attach the corresponding SQL statements to ensure that the generated results can be directly used for data quality inspection execution.
[0076] The output and feedback module outputs the generated rules (including field names, inspection types, rule descriptions, and SQL implementations) to downstream modules for subsequent verification, storage, and optimization. After user feedback, the large language model module can regenerate revised rules according to the adjusted Prompt, realizing a cycle of generation - feedback - optimization - regeneration to ensure that the rules are continuously improved and meet actual business requirements.
[0077] Specifically, the interactive interface module aims to provide an interactive platform for users, enabling them to review the automatically generated quality inspection rules for water conservancy data and provide feedback on the rule logic. After receiving the user feedback, the system dynamically adjusts the prompt of the generated rules based on the feedback, realizing a cyclic process of generation - feedback - optimization - regeneration.
[0078] The implementation solution of the interactive interface module in this embodiment is as follows: Users view the generated rules and the corresponding SQL statements through a graphical interface or a Web interface. The interface supports inputting feedback opinions, such as adjustment suggestions for rule descriptions, inspection logics, timing rules, or threshold settings.
[0079] After receiving the user feedback, the system automatically adjusts the original predefined format prompt, incorporates the modification opinions put forward by the user, and simultaneously updates the rule generation requirements related to the characteristics of water conservancy data in the expert knowledge base.
[0080] The modified Prompt will be submitted to the large language model again to generate a new rule version. This process loops continuously until the generated rules fully meet the business requirements.
[0081] The interactive interface module significantly reduces the errors caused by cross - professional communication, shortens the rule optimization cycle; improves the accuracy and applicability of the generated rules, ensuring that the generated rules can truly reflect the business logic of the water conservancy industry; forms a dynamically adaptive rule generation closed - loop, continuously improving the quality of the quality inspection rules.
[0082] Specifically, the main task of the SQL verification module is to automatically call and execute the SQL statements generated by the large language model to verify the correctness and logical consistency of the SQL statements in the actual data environment.
[0083] The implementation solution of the SQL verification module in this embodiment is as follows: The system automatically extracts the SQL statements included in the generated rules and executes them in the test database environment. By executing the SQL statements, the system calculates the data error rate under this quality inspection rule (for example, for integrity checks, counts the proportion of missing records) to determine whether the rule logic meets the expectations. Compare the execution result with the expected rule effect. If a syntax error or logical flaw is found in the SQL statement, an error report or prompt message will be automatically generated for users and developers to refer to, so as to further optimize or regenerate the rules.
[0084] The SQL verification module ensures the correctness and executability of the automatically generated rules in terms of technical implementation; reduces the risk of data quality inspection failure caused by SQL errors; improves the robustness and reliability of the overall data governance system.
[0085] Specifically, the storage and output module is responsible for storing the generated water conservancy data quality inspection rules in the database in a structured format and supporting the export of rule lists in multiple formats (such as JSON or tables), which is convenient for subsequent management, review, and use.
[0086] The implementation solution of the storage and output module in this embodiment is as follows: The storage and output module includes a rule storage module and a rule output module: The rule storage module classifies and stores the generated rules in the database according to information such as table name, field name, inspection type, rule description, and SQL statement.
[0087] The rule output module provides the function of exporting the complete rule list in JSON format or Excel (.xlsx) format. The exported content includes field name, rule type, rule description, and SQL statement, which is convenient for water conservancy experts or other users to review and share the automatically generated rules.
[0088] The storage and output module realizes the long-term structured storage and efficient management of rules, which is convenient for subsequent data quality monitoring work; supports multiple output formats, meets the rule review and application requirements in different usage scenarios; improves cross-departmental collaboration efficiency, and ensures the standardization and convenient sharing of data quality inspection rules.
[0089] Embodiment 3
[0090] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for automatically generating water conservancy data quality inspection rules based on the large language model described in the above technical solution.
[0091] Embodiment 4
[0092] The present invention provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method for automatically generating water conservancy data quality inspection rules based on the large language model described in the above technical solution.
[0093] By introducing large model technology, the present invention constructs an intelligent system that can not only understand the business logic of industry experts but also handle the technical details of database table structures, realizing the automatic generation of data quality inspection rules. The specific objectives are as follows: (1) Improve efficiency and reduce manual intervention The present invention automatically parses industry design documents (such as WORD format) through a large model, automatically extracts table structures and field information, combines with preset water conservancy expert knowledge information, realizes the generation of quality inspection rules from design documents, and automates the whole process of storing the quality inspection rules in the database in a set format. In an intelligent way, it greatly improves the efficiency of generating water conservancy data quality inspection rules and solves the problems of manual tediousness and low efficiency in the prior art.
[0094] (2) Improve flexibility and adaptability Through the powerful understanding ability of the large language model, the present invention can automatically generate corresponding quality inspection rules according to water conservancy business requirements and table structures, and can quickly adjust the rules when business requirements change to meet the needs of dynamic data governance. The system can automatically adapt to changes in various database types and table structures, improving the adaptability and flexibility of rule generation.
[0095] (3) Bridge the gap between industry knowledge and technical implementation One of the purposes of the present invention is to intelligently bridge the gap between industry knowledge and database technology implementation by introducing large language model technology. The large model can automatically parse the provided water conservancy business knowledge and database design documents, deeply understand the business logic and data constraints therein, and automatically generate quality inspection rules that meet technical standards based on these understandings. This process eliminates misunderstandings and omissions in manual conversion, ensuring that technicians can accurately generate high-quality data quality inspection rules based on business requirements, thus achieving seamless docking between industry knowledge and technical implementation.
[0096] Based on the large language model, the present invention realizes the automatic extraction of table structure information from Word documents in the water conservancy industry, including table names, field identifiers, field descriptions, etc., and generates standardized JSON format data, supporting the extraction of chapter numbers and the parsing of field code table information.
[0097] Based on the parsed table structure JSON data, the present invention automatically generates data quality inspection rules, covering five types: uniqueness, integrity, normativity, validity, and accuracy. For time series data (such as water level and flow rate), dynamic change inspection rules are generated, including the analysis of data change rate within 48 hours and the detection of the fluctuation range of maximum and minimum values. For non-time series rules such as threshold checks, reasonable ranges are automatically set in combination with water conservancy characteristics to ensure the accuracy and industry applicability of rule generation. The generated rules are directly converted into standardized SQL statements and are consistent with the rule descriptions, meeting the requirements of data quality management.
[0098] The present invention provides an interactive rule optimization function, which supports users to dynamically adjust the generated quality inspection rule logic according to actual needs to ensure a high degree of matching between the rules and the business scenario. At the same time, combined with the SQL verification module, it automatically detects the accuracy and coverage rate of the rules to ensure that the generated quality inspection rules are logically reasonable and feasible and meet the high standards of data governance requirements.
[0099] The present invention realizes the structured storage of generated rules, including field information, rule descriptions, and SQL statements, and supports the normalized output of the rule list in JSON or table form for easy review and management.
[0100] Through modular design, the present invention constructs a full-process automated system from document parsing to rule generation, verification, storage, and output, significantly improving the efficiency and adaptability of rule generation.
[0101] Compared with the prior art, the present invention significantly improves the efficiency and accuracy in the generation of water conservancy data quality inspection rules. The prior art mainly relies on the manual cooperation of industry experts and programmers, with a cumbersome process and prone to errors. By introducing a large language model and combining partial industry knowledge provided by experts, the present invention automatically completes table structure parsing and rule generation. The large model can expand based on expert knowledge to generate more comprehensive and accurate quality inspection rules, thus significantly reducing the depth and frequency of manual intervention and greatly improving efficiency.
[0102] Compared with the traditional templatized method, the present invention has higher flexibility in the generation of dynamic time-series rules and non-time-series rules, and can quickly expand to generate scientific and reasonable quality inspection rules by combining a small amount of industry-specific knowledge. This system can not only quickly respond to changes in business requirements but also significantly reduce the difficulty of cross-professional communication, meeting the requirements of high efficiency and dynamics in complex water conservancy data governance scenarios.
[0103] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0104] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0105] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0107] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention and the claims. All of these fall within the protection scope of the present invention.
[0108] What is not described in detail in this specification belongs to the known prior art of those skilled in the art.
Claims
1. A method for automatically generating water conservancy data quality inspection rules based on a large language model, characterized by: The following steps are involved: Use a large language model to parse documents in the water conservancy industry and extract table structure information; Automatically generate data quality check rules for the table structure information based on the extracted table structure information, wherein the data quality check rules cover multiple data quality dimensions; The data quality check rules are converted into standardized SQL statements.
2. The method according to claim 1, characterized in that: The document is a Word document of the water conservancy industry, the table structure information includes a table name, a field identifier and a field description, and the parsing process standardizes the table structure information into a JSON format.
3. The method according to claim 2, characterized in that: The parsing process further extracts the chapter numbers and code table information of the fields in the document.
4. The method according to claim 1, characterized in that: The data quality dimensions covered by the automatically generated data quality check rules include at least uniqueness, completeness, standardization, validity and accuracy.
5. The method according to claim 1, characterized in that: For time series data, the automatically generated data quality check rules include dynamic change check rules, which detect abnormal changes by analyzing the data change rate within a predetermined time window and the fluctuation range of the maximum and minimum values of the data.
6. The method according to claim 1, characterized in that: For non-time series data, the automatically generated data quality check rules include threshold check rules, and the threshold check rules automatically set a reasonable value range in combination with the characteristics of water conservancy industry data.
7. The method according to claim 1, characterized in that: The method also includes providing an interactive rule optimization function to enable a user to adjust the logic of the data quality check rule according to actual needs.
8. The method according to claim 1, characterized in that: The method further comprises verifying the generated data quality check rules by using an SQL verification module to detect the accuracy and coverage of the rules.
9. The method according to claim 1, characterized in that: The method also includes storing the generated data quality check rules in a database, and outputting a list of the data quality check rules in JSON format or in a table format.
10. A water conservancy data quality inspection rule automatic generation system based on a large language model, characterized in that: include: The document parsing module is used to parse water conservancy industry documents using a large language model, extract table structure information and generate corresponding JSON structured data; A rule generation module, used to generate data quality inspection rules based on the JSON structured data, wherein the rules cover multiple data quality dimensions, and convert the generated rules into standardized SQL statements; A rule optimization module, used for providing an interactive interface to enable a user to optimize and adjust the data quality inspection rules; A rule verification module, used to verify the accuracy and coverage of the data quality check rules using SQL; And a storage output module, used for storing the data quality inspection rules in a database and outputting them in JSON format or table form.
Citation Information
Patent Citations
Table data processing method based on natural language dialogue
CN116737909A
Quality inspection rule determination method and device, equipment and storage medium
CN116975044A
Knowledge base construction method, video automatic production method and software product
CN117952203A
Large language model retrieval enhancement generation method based on hierarchical information expansion
CN118779425A
Petroleum data quality inspection method
CN118886781A
Cited By
Government-affair-oriented Internet of Things data sharing and data quality inspection method, system and device based on large model, and storage medium
CN120670415A
Automatic data quality inspection system and method based on large model and data flow arrangement
CN120893585A