Data quality probing method and system based on large language model AI intelligent agent

By automatically identifying data types, matching governance rules and generating analysis scripts in data quality evaluation based on large language models, the problems of low efficiency, poor accuracy and data security in the existing technology are solved, and efficient, accurate and secure data quality analysis is achieved.

CN119938658APending Publication Date: 2025-05-06HANGZHOU SHUZHENG TECH CO LTD

Patent Information

Application Number
CN202510015641.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing technology relies on manual operations in data quality evaluation and optimization, which is inefficient and difficult to guarantee accuracy. Moreover, the application of large language model technology in data governance scenarios has problems with data security and low intelligence.

Method used

AI agents based on large language models are used to automatically identify the data types in the database, match corresponding governance rules, generate scripts for data quality analysis, and calculate the quality score of data in multiple dimensions to achieve an intelligent closed loop throughout the process.

Benefits of technology

It improves the efficiency and accuracy of data quality analysis, reduces manual intervention and errors, enhances the intelligence of data governance, and solves data security problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938658A_ABST
    Figure CN119938658A_ABST
Patent Text Reader

Abstract

The invention discloses a data quality probing method and system based on a large language model AI intelligent agent. The method comprises the following steps: firstly, automatically identifying data types in a database; matching a corresponding governance rule according to the data type; and finally, generating a script for data quality analysis, executing a data analysis task, outputting problem data, and calculating a quality score of the data according to an analysis result in combination with multiple dimensions. The system of the invention is designed to carry out the above method. Through the intelligent agent technology based on the large language model, the problems that in traditional data quality analysis, key steps are low in efficiency and low in intelligent degree are solved, intelligent closed loop of the whole process is achieved, and the efficiency and accuracy of data quality analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology and relates to a data quality detection method and system based on a large language model AI agent. Background Art

[0002] Data quality assessment and optimization are crucial to ensure effective use of data. However, traditional data quality exploration methods usually rely on manual operations, which not only require a lot of manpower to analyze the type and characteristics of data, but also require manual selection of governance rules based on data characteristics. This approach has many problems, such as missing Chinese comments in table fields and irregular naming, which requires data developers to invest a lot of energy in analyzing the type of data. After determining the data type, they also need to manually select matching governance rules based on the data type. The whole process is cumbersome, inefficient, and difficult to guarantee accuracy.

[0003] Large Language Model Technology: Large Language Model (LLM) is a type of natural language processing technology based on deep learning. It uses a large-scale dataset to train a model on large-scale network parameters, and has powerful language understanding and generation capabilities. LLM relies on ultra-large-scale parameters and hierarchical structures, and can learn complex language patterns, contextual semantic relationships, and reasoning capabilities from massive text data. Its core technology is based on the Transformer's self-attention mechanism, which allows the model to capture global context dependencies when generating text, thereby improving language understanding capabilities.

[0004] Lora: Lora (Low-Rank Adaptation) is a lightweight model fine-tuning technology that aims to reduce the computational cost of fine-tuning large language models for specific tasks. Lora reduces the number of parameters required in the fine-tuning process by limiting the update of model parameters to the range of low-rank matrices, while retaining the powerful capabilities of large language models. It locally tunes the model in a modular manner by inserting low-rank matrices between existing model layers, and can achieve efficient optimization for specific tasks without significantly increasing the computational burden. Taking the data type recognition scenario as an example, Lora can be used to fine-tune a large language model on a specific data set, thereby improving the model's ability to recognize data types. Through specific learning of data patterns, Lora fine-tuning training enhances the model's ability to perceive and classify data in different business scenarios, and improves the accuracy of data recognition of large language models in data governance scenarios.

[0005] Patent No. CN116991932A "A data analysis and governance system based on artificial intelligence and its method" describes a technical method of collecting and cleaning data, then classifying the target data and judging whether the data is abnormal, and then further analyzing and improving the abnormal data. However, according to the description in specification

[0074] , whether the data is abnormal is determined by calculating the fluctuation range of the data. This method of determining abnormalities is simple and crude. This patent adopts authoritative standards such as national, local and industry standards as the basis, which is more authoritative in judging whether the data is abnormal, and the granularity of the analysis is also higher.

[0006] Patent No. CN 118297069 B "Data Management System, Method, Device and Medium Based on Natural Language Processing" proposes a data management system based on natural language processing, which mainly uses part-of-speech recognition to extract medical entities in text, then calculates entity similarity based on text similarity, and obtains medical unified coding based on entity matching. In general, the management content is simple.

[0007] Patent No. CN 118379026 A "An Intelligent Data Governance Management System Based on Consultation Planning" describes a data governance consultation planning system based on question-and-answer format. The system allows users to input questions, and if relevant data governance analysis results are available, the results are returned; if not, data governance tasks are generated through information extraction, data quality is checked, and analysis reports are output. However, the patent mainly describes the main route of technical implementation, and data quality-related content is only used as a submodule, and the implementation principle of the submodule is not explained in detail. It is only mentioned in the specification

[0016] that the data quality submodel is trained through the data governance dimension knowledge base, but the specific process of data quality analysis is not explained.

[0008] Patent No. CN 116501829 A "A Data Governance Method and System Based on Artificial Intelligence Large Language Model Platform" classifies data by expression form

[0007] , and performs feature recognition on the classified data

[0008] ,

[0009] proposes a data selection module for comparing text data and filtering out pre-selected text data according to thresholds,

[0010] proposes planning different processing schemes and processing flows according to pre-selected data, and finally

[0011] calculates the selection rate and performs priority evaluation through a priority evaluation module. The research direction is mainly to plan different processing schemes and flows for pre-selected data. However, the article does not mention how to plan the processing scheme, and lacks the specific processing flow of each processing scheme.

[0009] Patent CN117421328A "Data governance method, system, device and readable medium based on artificial intelligence" proposes a third-party large language model interface service (such as ChatGPT, Wenxin large language model, etc.)

[0054] , which generates ETL processes based on user input, generates corresponding SQL instructions for each process node, and then verifies through small batch data sampling and processing. If it fails, it will be regenerated until it succeeds. There are data security issues when using a third-party large language model interface, and

[0054] mentioned that it relies on user-guided input of demand descriptions, cannot automatically identify data and automatically match governance rules, and has a low level of intelligence.

[0010] Although the title of patent CN 118312998 A "A Large Language Model System and Method for Data Governance Analysis" has a high degree of similarity, the core content of the research is to analyze the layout and calculate the semantic sensitivity coefficient of each layout element in the analysis results, and determine whether to encrypt the document based on the overall sensitivity coefficient.

[0011] In summary, the existing technical solutions have the following problems:

[0012] 1. Summary of the problem from the perspective of small models: In order to distinguish them from large language models, the patents that do not mention the use of large language model technical solutions are collectively referred to as small models. For example, the above-mentioned patents CN116991932A, CN118297069B, and CN118312998A all use small model solutions. Compared with large language models, small models have fewer parameters and lower model complexity. They are suitable for processing relatively simple tasks, but have weak learning capabilities for complex data patterns and features. In data governance scenarios, they have problems such as inaccurate data type recognition, inaccurate rule recommendations, and SQL generation errors. Therefore, through the above patent analysis, it can be found that the data governance business scenarios that the small model can handle are very limited. For example, CN116991932A judges whether the data is abnormal based on the coarse-grained data fluctuation range rather than based on specific data; CN 118297069B only achieves the acquisition of medical unified coding after medical entity extraction, and the governance result is relatively simple; CN118312998A extracts layout elements through layout analysis algorithms and performs sensitivity analysis, and determines whether to encrypt documents based on the sensitivity coefficient. In general, due to the insufficient capabilities of the small model, the data governance scenarios that can be achieved based on the small model technology are relatively simple.

[0013] 2. Summary of the problem from the perspective of the large language model: There are data security issues when calling third-party interfaces, and most inventions do not disclose a detailed description of the technical solution for the data quality analysis process. For example, patents CN118379026A, CN116501829A, and CN117421328A use the implementation of the large language model, but the key technical implementation routes of CN118379026A and CN116501829A are not described in detail. For example, CN 118379026A mentions the data quality module, but does not elaborate on the specific quality analysis process; CN116501829A mentions planning different processing solutions and processes for pre-selected data, but does not elaborate on the technical implementation instructions; and the large language model in CN117421328A calls a third-party interface service. This method has a large data security problem because it needs to transmit metadata or data samples to a third-party server, and it relies on user-guided input of demand descriptions, and has a low degree of intelligence. Summary of the invention

[0014] In view of the deficiencies in the prior art, the present invention provides a data quality detection method and system based on a large language model AI agent.

[0015] In a first aspect, the present invention provides a data quality detection method based on a large language model AI agent, the method comprising:

[0016] Automatically identify data types in the database;

[0017] Match the corresponding governance rules according to the data type;

[0018] Generate scripts for data quality analysis and execute data analysis tasks, output problem data, and calculate the data quality score based on the analysis results and multiple dimensions.

[0019] Furthermore, by automatically analyzing the fields, table structures, and remark information in the database, combined with data sampling and large language model element exploration agents, data types can be quickly identified;

[0020] Furthermore, the automated analysis is specifically:

[0021] Collect metadata information including table name, field name and remarks from the metadata database;

[0022] Preprocess the collected metadata information;

[0023] The metadata rule library is used to determine the data type of keywords in table names, field names, and notes.

[0024] Furthermore, for fields whose types cannot be determined by rule matching, data content detection is used, specifically:

[0025] For data sampling, a certain amount of non-empty and non-repetitive data is randomly selected for analysis, together with the task description prompt words and the few-shot learning samples as the input of the feature exploration agent. The feature exploration agent analyzes the input samples and outputs the data type of the samples;

[0026] The element detection agent is an AI model built based on a large language model and determines the data type according to the field data characteristics.

[0027] Furthermore, the matching process is assisted by task description prompts and corresponding example data; and the few-sample learning capability of the large language model is utilized to obtain a data governance rule set applicable to the selected field by analyzing a small number of sample sets.

[0028] Furthermore, the script is used to generate an intelligent agent, which automatically generates a data quality SQL analysis script corresponding to the current data based on the matching governance rules, combined with the task prompt words and metadata information;

[0029] Automatically execute the above SQL analysis script to filter out data with quality problems from the database;

[0030] The data quality score is calculated based on the six dimensions of completeness, relevance, standardization, consistency, uniqueness and accuracy of the screened data.

[0031] Furthermore, it also includes generating a data quality analysis report, specifically:

[0032] Generate each chapter in the report in sequence according to the preset outline framework, covering different dimensions of data quality;

[0033] After automatic generation of each chapter, it is integrated into a complete data quality analysis document.

[0034] In a second aspect, the present invention provides a data quality detection system based on a large language model AI agent, comprising:

[0035] Data exploration module, used to automatically identify data types in the database;

[0036] The rule recommendation module is used to match the corresponding governance rules according to the data type;

[0037] The quality analysis module is used to generate scripts for data quality analysis and execute data analysis tasks, output problem data, and calculate the quality score of the data based on the analysis results and multiple dimensions.

[0038] In a third aspect, the present invention provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction implements the steps of the above method when executed by the processor.

[0039] In a fourth aspect, the present invention provides a readable storage medium, on which a program or instruction is stored, and the program or instruction implements the steps of the above method when executed by a processor.

[0040] Beneficial effects of the present invention: The present invention uses the intelligent agent technology based on the large language model to make up for the problems of low efficiency and low intelligence of key steps in traditional data quality analysis, realizes the intelligent closed loop of the whole process, and improves the efficiency and accuracy of data quality analysis. Specifically:

[0041] The element exploration agent in the technical solution of the present invention reduces human intervention through automatic data type recognition, improves efficiency and eliminates errors in the manual analysis process.

[0042] The rule recommendation agent is based on the Few-Shot Learning technology of the large language model. It not only solves the problems of low efficiency and poor applicability of traditional manually configured rules, but also optimizes the model recommendation effect by adjusting the prompt words without retraining the model.

[0043] The present invention also introduces a script generation agent based on a large language model, which can automatically convert the governance rules described in natural language into SQL analysis scripts, reducing the tedious steps of manual SQL writing, improving work efficiency, and reducing possible errors in the manual writing process, thereby ensuring the accuracy and efficiency of data quality analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A system composition diagram of an embodiment of the present application;

[0045] Figure 2 A schematic diagram of a method according to an embodiment of the present application;

[0046] Figure 3 A schematic diagram of data exploration in an embodiment of the present application;

[0047] Figure 4 A schematic diagram of generating an element exploration agent for an embodiment of the application;

[0048] Figure 5 A schematic diagram of rule recommendations for application examples;

[0049] Figure 6 A schematic diagram of quality analysis of an embodiment of the application. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0051] like Figure 1 As shown, a data quality exploration system based on a large language model AI agent provided in an embodiment of the present application consists of four modules: a data exploration module, a rule recommendation module, a quality analysis module and a report generation module.

[0052] The data exploration module is used to automatically identify data types. This module automatically analyzes the fields, table structures, and notes in the database, and combines data sampling and large language model element exploration agent technology to quickly identify data types (such as mobile phone numbers, addresses, or time, etc.). Compared with the traditional method that relies on manual judgment, the data exploration module greatly improves recognition efficiency and accuracy.

[0053] The rule recommendation module is used to match the corresponding governance rules according to the data type. It uses the Few-Shot Learning method based on the large language model to automatically recommend the corresponding data governance rules according to the data type. Through the few-shot learning technology, the module can optimize the recommendation strategy for data governance needs in different fields without retraining, significantly improving the flexibility and accuracy of rule recommendations.

[0054] The quality analysis module is used to generate scripts for data quality analysis and perform data analysis tasks, output problem data, and calculate the quality score of data based on multiple dimensions based on the analysis results. It generates SQL analysis scripts through a script generation agent based on large language model technology to achieve multi-dimensional analysis of data quality, including data integrity, relevance, standardization, consistency, uniqueness and accuracy, reducing the complexity of manually writing SQL and improving the efficiency of data governance.

[0055] The report generation module automatically generates data quality reports covering multiple dimensions based on the analysis results. The report content is comprehensive and the information is accurate, which simplifies the traditional manual writing process, helps users quickly understand the data quality status, and provides a basis for subsequent data governance decisions.

[0056] like Figure 2As shown, an embodiment of the present application provides a data quality detection method based on a large language model AI agent. The method performs comprehensive data detection and quality analysis on the data in the database through the collaborative work of multiple large language model agents.

[0057] First, the feature exploration agent automatically identifies the data types in the database;

[0058] Secondly, the rule recommendation agent matches the corresponding governance rules according to the data type;

[0059] Then, the script generation agent converts natural language into SQL analysis statements, generates scripts for data quality analysis, executes data analysis tasks, and outputs problematic data; based on the analysis results, the quality score of the data is calculated based on six dimensions: data integrity, relevance, standardization, consistency, uniqueness, and accuracy;

[0060] Finally, a data quality analysis report is generated.

[0061] This embodiment greatly improves the intelligence level of data quality analysis and the efficiency of data governance analysis through the collaborative work of various large language model agents, providing an important basis for subsequent data governance work.

[0062] Furthermore, automatic data type identification: before doing data quality analysis, the data type is analyzed (such as determining whether the field is a mobile phone number, date, address, etc.), and then the corresponding governance rules are matched based on the analysis results for analysis. This process often relies on manual analysis, and the type of the field is determined by the table, field, and remarks, or when the table and field information cannot be determined, the field content data needs to be sampled and analyzed to manually determine the data type. The significance of automatic data type identification is to use artificial intelligence technology to automate the tedious manual analysis process, reduce manual intervention, reduce the risk of human error, and significantly improve efficiency and accuracy.

[0063] Further, intelligent rule recommendation: Under the current technical background, the recommendation of governance rules mainly relies on manual configuration. This process is not only inefficient, but also difficult to flexibly respond to the complex and changing data governance needs of individual industries because the governance personnel are limited by the knowledge background of specific industries. To address this challenge, this embodiment proposes an intelligent rule recommendation system based on Few-Shot Learning of a large language model, which aims to automatically recommend corresponding governance rules by analyzing the data characteristics of the data. Based on the Few-Shot Learning capability of the large language model, the model recommendation results can be optimized by dynamically adjusting the learning samples in the Prompt and optimizing the task prompt words without retraining the model. This method significantly improves the applicability and accuracy of governance rule recommendations, enabling it to better adapt to data governance needs in different fields.

[0064] Further, intelligent data analysis script generation: In data governance work, traditional quality inspection and analysis of data usually rely on manual writing of SQL, which is not only time-consuming and error-prone, but also limits the efficiency and accuracy of data governance. In order to solve this problem, the embodiment introduces a script generation agent (NL2SQL) based on a large language model, which realizes the automatic generation of governance rules described in natural language to SQL analysis scripts. Based on the powerful semantic understanding and generation capabilities of the large language model, the agent can accurately parse the governance rules described in natural language and generate corresponding SQL analysis scripts. This simplifies the steps of writing SQL in data analysis and greatly improves the work efficiency of data governance.

[0065] Furthermore, efficient data quality analysis report generation: Existing data quality analysis reports mainly rely on manual writing, which is inefficient and prone to errors. By introducing an intelligent report generation mechanism based on a large language model, a data quality assessment report covering six dimensions of completeness, relevance, standardization, consistency, uniqueness and accuracy is automatically generated based on the data quality calculation results. This intelligent report generation process not only simplifies the complex operation steps of manual report writing, but also ensures the comprehensiveness and consistency of information, greatly improving the efficiency and accuracy of report generation, and helping users quickly obtain comprehensive quality insights.

[0066] Furthermore, data security and privacy protection: Some existing solutions rely on third-party interface services for data processing, which poses data security risks. This embodiment completely eliminates the reliance on external interface services by training and deploying large language models locally, ensuring that the entire data processing process is carried out in a safe and controllable environment.

[0067] An embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction implements the steps of the above method when executed by the processor.

[0068] The embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored, and the steps of the above method are implemented when the program or instruction is executed by a processor.

[0069] like Figure 3 As shown, in some embodiments, the data exploration module is mainly divided into two parts, namely metadata exploration and data content exploration.

[0070] Furthermore, metadata exploration refers to collecting metadata information such as table names, field names, and notes from the metadata database. Then, the collected metadata is preprocessed, such as removing special characters and other cleaning operations. Then, the metadata rule library is used to make a preliminary judgment on the keywords in the table names, field names, and notes information.

[0071] Specifically, the rule base stores a large number of keywords related to common data types, such as "mobile phone number", "address", "time", "date", etc. During the processing, the system scans the field name, table name and remark information to determine whether they contain the above keywords. If the match is successful, the field type is directly determined to be the corresponding data type.

[0072] Furthermore, for fields whose types cannot be determined by rule matching (such as table field names that are not standardized and notes are missing), data content exploration is used, that is, the system will sample data, randomly extract a certain amount of non-empty and non-repetitive data for analysis, and use it together with the task description prompt words and few-shot learning samples as input for the feature exploration agent. The feature exploration agent analyzes the input samples and outputs the data type of the samples.

[0073] After the data content detection is completed, the metadata detection results are merged with the content detection results as the input of the next module (rule recommendation module).

[0074] The feature detection agent is an AI model built based on a large language model to determine the data type according to the characteristics of field data. It can identify the potential patterns in the data by fine-tuning the training based on the large language model Lora, such as the character structure of the phone number, the format of the date, the geographic information of the address, etc., so as to realize automatic identification of data types, significantly improve the efficiency and accuracy of data type identification, and reduce the risk of manual intervention and human error, such as Figure 4 Compared with the solution of calling third-party model interfaces, this solution based on local training and deployment of models can better ensure data security.

[0075] like Figure 5 As shown, in some embodiments, rule recommendations are based on the few-shot learning capability of the large language model, which can significantly improve the accuracy and efficiency of governance rule matching. Compared with the traditional approach of manually configuring governance rules, rule recommendations based on the large language model can automatically recommend corresponding governance rule sets based on data types (field Chinese names), thereby avoiding the problem of low efficiency when manually configuring governance rules.

[0076] In the rule recommendation module, the input data mainly comes from the results of the data exploration module, assisted by task description prompts and corresponding sample data. Based on these inputs, the system can use the few-shot learning capability of the large language model to derive a set of data governance rules applicable to a specific field by analyzing a small number of sample sets.

[0077] Another important advantage of this module is that it can adjust the model's recommendation capabilities without retraining the model. This process relies on the flexibly adjustable prompt. By dynamically adjusting a small number of example samples and task prompt words, the recommendation strategy can be flexibly improved, thereby continuously improving the accuracy and applicability of the recommendation.

[0078] like Figure 6 As shown, in some embodiments, traditional data quality analysis usually relies on manually written SQL scripts, which is not only inefficient, but also requires analysts to have strong SQL skills and is easily affected by human errors, resulting in inaccurate analysis results and low efficiency. Therefore, the quality analysis module in this embodiment introduces script generation agent technology and uses a large language model to generate SQL scripts based on governance rules, which greatly optimizes the process of data quality analysis.

[0079] First, the script generation agent is used to automatically generate the data quality SQL analysis script corresponding to the current data based on the governance rules recommended by the rule recommendation module, combined with the task prompt words and metadata information. This process converts complex governance rule text into SQL analysis scripts through natural language processing technology, which not only simplifies the difficulty of writing analysis scripts, avoids possible deviations or errors when manually writing SQL, but also greatly improves the efficiency of writing analysis scripts.

[0080] Then, the above SQL analysis script is automatically executed to filter out data with quality problems from the database.

[0081] Finally, the data quality score is calculated based on the six dimensions of completeness, relevance, standardization, consistency, uniqueness and accuracy of the screened data, providing a data basis for the subsequent generation of data quality reports.

[0082] In summary, the quality analysis module uses large language model technology to realize the efficiency and intelligence of the data quality analysis process, and improves the accuracy and work efficiency of data governance.

[0083] In some embodiments, the quality report generation realizes the automatic generation of quality reports through a large language model, which significantly improves the efficiency and accuracy of data quality analysis report generation. First, the module will generate the various chapters in the report in sequence according to the preset outline framework, covering different dimensions of data quality, such as completeness, relevance, standardization, consistency, uniqueness and accuracy.

[0084] During the generation process, each chapter takes the writing requirements of the current chapter and the corresponding quality analysis data as the input of this module. This module generates a summary report based on the analysis data and combines it with the queried quality analysis data as the content of the current chapter.

[0085] Finally, after automatic generation of each chapter, a complete data quality analysis document is integrated. This report not only includes data quality scores from various dimensions, but also provides in-depth insights through the analysis results of the large language model, helping users to fully understand the advantages and disadvantages of data quality.

[0086] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0087] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0088] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A data quality detection method based on a large language model AI agent, characterized in that The method includes: Automatically identify data types in the database; Match the corresponding governance rules according to the data type; Generate scripts for data quality analysis and execute data analysis tasks, output problem data, and calculate the data quality score based on the analysis results and multiple dimensions.

2. According to claim 1, a data quality detection method based on a large language model AI agent is characterized by: By automatically analyzing the fields, table structures and annotation information in the database, combined with data sampling and large language model element detection agents, data types can be quickly identified.

3. The data quality detection method based on a large language model AI agent according to claim 2 is characterized by: The automated analysis is specifically: Collect metadata information including table name, field name and remarks from the metadata database; Preprocess the collected metadata information; The metadata rule library is used to determine the data type of keywords in table names, field names, and notes.

4. The data quality detection method based on a large language model AI agent according to claim 3 is characterized by: For fields whose types cannot be determined through rule matching, data content detection is used, specifically: For data sampling, a certain amount of non-empty and non-repetitive data is randomly extracted for analysis, and together with the task description prompt words and the few-shot learning samples, it is used as the input of the feature exploration agent. The feature exploration agent analyzes the input samples and outputs the data type of the samples; The element detection agent is an AI model built based on a large language model and determines the data type according to the field data characteristics.

5. The data quality detection method based on a large language model AI agent according to claim 1 is characterized in that: The matching process is assisted by task description prompts and corresponding example data; and the few-sample learning ability of the large language model is utilized to obtain a data governance rule set applicable to the selected field by analyzing a small number of sample sets.

6. The data quality detection method based on a large language model AI agent according to claim 1, characterized in that: Use scripts to generate intelligent agents, based on matching governance rules, combined with task prompts and metadata information, to automatically generate data quality SQL analysis scripts corresponding to the current data; Automatically execute the above SQL analysis script to filter out data with quality problems from the database; The data quality score is calculated based on the six dimensions of completeness, relevance, standardization, consistency, uniqueness and accuracy of the screened data.

7. A data quality detection method based on a large language model AI agent according to any one of claims 1 to 6, characterized in that: It also includes generating data quality analysis reports, specifically: Generate each chapter in the report in sequence according to the preset outline framework, covering different dimensions of data quality; After automatic generation of each chapter, it is integrated into a complete data quality analysis document.

8. A data quality detection system based on a large language model AI agent, characterized in that: include: Data exploration module, used to automatically identify data types in the database; The rule recommendation module is used to match the corresponding governance rules according to the data type; The quality analysis module is used to generate scripts for data quality analysis and execute data analysis tasks, output problem data, and calculate the quality score of the data based on the analysis results and multiple dimensions.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Data governance method and system based on artificial intelligence large language model platform

    CN116501829A

  • Data analysis and treatment system and method based on artificial intelligence

    CN116991932A

  • Data management method, system and device based on artificial intelligence and readable medium

    CN117421328A

  • Data governance system, method, device and medium based on natural language processing

    CN118297069B

  • Large data governance analysis model system and method

    CN118312998A

Cited By

  • Data processing method and device, electronic equipment, storage medium and program product

    CN122132466A

  • AI model iteration-oriented data quality closed-loop management and intelligent evaluation system

    CN122132777A