Intelligent data quality rule generation method and system based on large model
Through the intelligent data quality rule generation method based on large models, the problems of low rule generation efficiency, relying on expert experience, low conversion efficiency and insufficient dynamic evolution capabilities of the rule base in the existing technology are solved, efficient and accurate rule generation and maintenance are achieved, cost reduction and adapting to the needs of data changes.
Patent Information
- Application Number
- CN202510228174.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems such as inefficient efficiency, relying on expert experience, low conversion efficiency of natural language rules, and lack of dynamic evolution capabilities in the generation and management of data quality rules.
The intelligent data quality rule generation method based on large models is adopted, and technical means such as metadata-driven dynamic rule construction, multi-modal rule conversion, rule confidence evaluation and rule base self-optimization are used. Specifically, it includes: metadata acquisition and analysis, multimodal rule conversion engine, rule confidence evaluation module and rule base self-optimization module.
The efficiency of rule generation is significantly improved, the accuracy of natural language to code conversion is high, the cost of rule maintenance is reduced, and the rule base can better adapt to the dynamic changes of data and meet the needs of large-scale data processing.
Smart Images

Figure CN120196620A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of data governance and artificial intelligence, and particularly relates to an intelligent data quality rule generation method and system combining large language models and metadata analysis. Background Art
[0002] In today's digital age, data has become the core asset of enterprises and organizations, and high - quality data is crucial for decision - making, business operations, and strategic planning. However, with the continuous growth of data scale and the increasing diversification of data sources, data quality problems have become more prominent. The generation and management of data quality rules are the key links to ensure data quality, but there are many defects in the existing technologies.
[0003] Rule generation depends on expert experience: Existing technologies such as CN114003710A only implement rule template matching, which highly depends on expert experience. This makes the rule generation process inefficient and difficult to adapt to complex and changing data environments. The data structures and relationships vary greatly in different business scenarios, and it is difficult for experts to cover all situations comprehensively, resulting in insufficient rule coverage. Through actual measurement, traditional methods can only cover 63% of field associations, and a large number of data association relationships cannot be effectively constrained, making it difficult to guarantee data quality.
[0004] Low conversion efficiency between natural language rules and executable code: Technologies such as US20220365925A1 require manual coding when converting natural language rules into executable code. This not only increases labor costs and time costs but also is prone to errors, with a conversion error rate as high as 38%. The manual coding process is cumbersome, requires professional programming knowledge, and is easily affected by human factors, resulting in the generated code being unable to accurately execute the intention of natural language rules.
[0005] The rule base lacks the ability of dynamic evolution: WO2022156680A1 adopts a static rule set and cannot update rules in a timely manner according to data changes and newly emerging problems. This makes the rule base gradually out of touch with the actual data situation, requiring 15 person - hours of manual maintenance per month, and the speed of discovering new rules is slow. It takes 72 hours from problem discovery to new rule generation, making it difficult to meet the rapidly changing data business requirements. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to overcome the defects of the existing technology and provide an intelligent data quality rule generation method and system based on large models.
[0007] To solve the above - mentioned technical problems, the present invention provides the following technical solutions:
[0008] An intelligent data quality rule generation method and system based on large models of the present invention includes the following steps:
[0009] Metadata-driven Dynamic Rule Construction: Using data lineage graph parsing technology, automatically extract feature vectors such as table structure, field associations, and business constraints to construct a context for rule generation; input the extracted feature vectors into a large model for rule reasoning; construct a rule generation context matrix CTX = αSchema + βLineage + γ*Constraint, where α = 0.6, β = 0.3, γ = 0.1;
[0010] Multi-modal Rule Conversion: Design a cross-modal conversion model based on deep learning to achieve automatic conversion of natural language rules into executable code; adopt a three-layer conversion architecture, including a syntax parser, a semantic disambiguation model, and a code generator;
[0011] Rule Confidence Evaluation: Use a hybrid evaluation algorithm to calculate the rule confidence, and the formula is confidence = α model probability + β historical accuracy + γ expert score, where α + β + γ = 1; adopt a dynamic weight adjustment formula β = 1 / (1 + e^(-kt)), k is the learning rate parameter, and t is the number of times the rule is used;
[0012] Rule Base Self-optimization: Build a rule life cycle management system to achieve a closed-loop of abnormal pattern detection, rule evaluation, and version iteration; use an improved version of the IsolationForest-based abnormal pattern detection algorithm.
[0013] As a preferred technical solution of the present invention, in the metadata-driven dynamic rule construction method, when extracting feature vectors through data lineage graph parsing technology, the method of the MetadataParser class is adopted;
[0014] including extract_features for extracting features;
[0015] _calculate_field_correlation for calculating field associations;
[0016] _mine_business_constraints for mining business constraints;
[0017] And return a FeatureSet containing feature vectors and constraints.
[0018] As a preferred technical solution of the present invention, in the multi-modal rule conversion, the data for training the cross-modal conversion model includes 500,000 groups of NL-SQL / Python paired samples.
[0019] As a preferred technical solution of the present invention, in the rule confidence evaluation, in the test scenario of the financial risk control data set, the adopted hybrid evaluation formula is Confidence = 0.4P_model + 0.3Accuracy + 0.2Expert + 0.1Freshness, and the learning rate parameter k in the dynamic weight adjustment algorithm takes the value of 0.1.
[0020] As a preferred technical solution of the present invention, in the rule library self-optimization mechanism, the rule life cycle management system performs state conversion according to the rule confidence. When the rule confidence > 0.8, it is in the Active state. When the confidence is < 0.6 for 3 consecutive times or a new abnormal pattern is found, it enters the Deprecated state, and then the rule version is iteratively updated.
[0021] An intelligent data quality rule generation system based on a large model according to the present invention includes:
[0022] Metadata collection and analysis module: used to collect the metadata of the data, and through the data lineage graph parsing technology, extract feature vectors such as table structure, field association, and business constraints, and construct the rule generation context;
[0023] Multimodal rule conversion engine module: Develop a cross-modal conversion model based on deep learning to realize the automatic conversion of natural language rules into executable code, and internally includes a syntax parser, a semantic disambiguation model, and a code generator;
[0024] Rule confidence evaluation module: Use a hybrid evaluation algorithm and a dynamic weight adjustment formula to evaluate the confidence of the generated rules;
[0025] Rule library self-optimization module: Construct a rule life cycle management system, and use an abnormal pattern detection algorithm to realize the abnormal pattern detection, evaluation, and version iteration of the rule library.
[0026] As a preferred technical solution of the present invention, the metadata collection and analysis module can collect the metadata of the order table and analyze the association relationship with the logistics table in the e-commerce data governance scenario.
[0027] As a preferred technical solution of the present invention, the multimodal rule conversion engine module can convert the natural language rule "The order creation time should be earlier than the logistics delivery time" into the SQL check statement "SELECT * FROM orders o JOIN logistics l ON o.order_id = l.order_id WHERE o.create_time > l.delivery_time" in the e-commerce data governance scenario;
[0028] In the financial risk control data scenario, the rule that "the ID number should conform to the GB11643 format" can be converted into a Python verification function.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] 1. The rule generation efficiency of the present invention is greatly improved: compared with the traditional manual method, the rule generation efficiency is increased by 85%. The traditional method can only generate 12 rules per hour, while the present invention can generate 85 rules per hour, an increase of 608%, greatly improving the generation speed of data quality rules and meeting the needs of large-scale data processing.
[0031] 2. The present invention has a high accuracy in converting natural language to code: the accuracy of converting natural language to code reaches 92.7% (F1 value), showing a significant improvement compared with the 38% conversion error rate in the prior art. The code generation speed reaches 200 rules per minute (tested on AWS c5.4xlarge instances), realizing the efficient and accurate conversion of natural language rules to executable code.
[0032] 3. The rule maintenance cost of the present invention is significantly reduced: the rule maintenance cost is reduced by 70%, from 15 person-hours per month to 4.5 person-hours per month. The speed of discovering new rules is increased by 3 times, from 72 hours to 24 hours, effectively reducing the manual maintenance workload, improving the update efficiency of the rule base, and enabling the rule base to better adapt to the dynamic changes of data. Description of the Drawings
[0033] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0034] Figure 1 : Overall system flow chart;
[0035] Figure 2 : Multimodal conversion process timing diagram;
[0036] Figure 3 : Rule life cycle state diagram; Detailed Embodiments
[0037] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0038] Embodiment 1
[0039] As Figures 1-3 shown, the present invention provides an intelligent data quality rule generation method and system based on a large model, including:
[0040] Metadata-driven Dynamic Rule Construction Method: Adopt data lineage graph parsing technology to automatically extract feature vectors such as table structure, field association, and business constraints, and construct a rule generation context. By developing a data lineage graph parsing algorithm, key features of data can be accurately obtained. On this basis, construct a rule generation context matrix CTX = αSchema + βLineage + γ*Constraint (α = 0.6, β = 0.3, γ = 0.1, the optimal weight combination determined through a large number of experiments). Input the feature matrix into a large model for rule reasoning, thereby generating quality rules that better conform to the actual data situation.
[0041] Multi-modal Rule Conversion Engine: Develop a cross-modal conversion model based on deep learning to achieve automatic conversion from natural language to executable code. Design a three-layer conversion architecture, including a syntax parser, a semantic disambiguation model, and a code generator. The syntax parser performs syntax analysis on natural language rules, the semantic disambiguation model eliminates semantic ambiguities, and the code generator generates executable Python or SQL code based on the analysis results. The training data contains 500,000 groups of NL-SQL / Python paired samples to improve the accuracy and stability of the conversion.
[0042] Rule Confidence Evaluation Model: Design a hybrid evaluation algorithm, confidence = α model probability + β historical accuracy + γ expert score, where α + β + γ = 1. Through the dynamic weight adjustment formula β = 1 / (1 + e^(-kt)) (k is the learning rate parameter, t is the number of times the rule is used), dynamically adjust the evaluation weights according to the usage of the rule. In the financial risk control dataset test, adopt the hybrid evaluation formula Confidence =
[0043] 0.4P_model + 0.3Accuracy + 0.2Expert + 0.1Freshness, effectively reducing the rule misjudgment rate.
[0044] Rule Base Self-optimization Mechanism: Construct a rule life cycle management system to achieve a closed-loop of anomaly pattern detection -> rule evaluation -> version iteration. Use an improved version of the anomaly pattern detection algorithm based on Isolation Forest to timely detect anomaly patterns in data. When the confidence of the rule is <0.6 for 3 consecutive times or a new anomaly pattern is found, evaluate and iterate the version of the rule to ensure that the rule base always adapts to the changes in data.
[0045] Among them Figure 1 : System Overall Flowchart: It shows the complete process from metadata collection, through data lineage analysis, rule feature extraction, large model reasoning, natural language rule generation, multi-modal conversion, rule testing, confidence evaluation, to rule base update. Each link is closely connected to form an organic whole.
[0046] Figure 2 : Multi-modal conversion process timing diagram: It details the chronological order and operation process from when the user submits natural language rules to syntax parsing, semantic disambiguation, code generation, and finally returning Python / SQL code during the multi-modal rule conversion, clearly presenting the working mechanism of the cross-modal conversion model.
[0047] Figure 3 : Rule lifecycle state diagram: It shows the different states of rules in the lifecycle, including Candidate, Active, Deprecated, Updated, etc., as well as the conversion conditions between states. For example, when the confidence level > 0.8, the rule is in the Active state; when the confidence level < 0.6 for three consecutive times or a new abnormal pattern is discovered, it enters the Deprecated state and is then updated, intuitively demonstrating the operating logic of the rule library self-optimization mechanism.
[0048] In specific usage, e-commerce data governance:
[0049] 1. Collect metadata of the order table (including fields such as user ID, product ID, transaction time, etc.)
[0050] 2. Discover the association relationship with the logistics table through data lineage analysis
[0051] 3. Generate a rule: "The order creation time should be earlier than the logistics delivery time"
[0052] 4. Automatically convert it into an SQL check statement:
[0053] SELECT * FROM orders o
[0054] JOIN logistics l ON o.order_id = l.order_id
[0055] WHERE o.create_time > l.delivery_time
[0056] Financial risk control data:
[0057] 1. Analyze the field constraints of the customer information table
[0058] 2. Automatically generate a rule: "The ID number should conform to the GB11643 format"
[0059] 3. Convert it into a Python validation function:
[0060] def validate_id(id_num):
[0061] # Implement the validation algorithm
[0062] return check_sum(id_num) == int(id_num[-1]).
[0063] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for generating intelligent data quality rules based on a large model, characterized in that: The following steps are involved: Metadata-driven dynamic rule construction: Use data lineage graph analysis technology to automatically extract feature vectors such as table structure, field association, and business constraints to build rule generation context; input the extracted feature vectors into the big model for rule reasoning; build the rule generation context matrix CTX = αSchema + βLineage + γ*Constraint, where α = 0.6, β = 0.3, γ = 0.1; Multimodal rule conversion: Design a cross-modal conversion model based on deep learning to achieve automatic conversion of natural language rules to executable code; Adopt a three-layer transformation architecture, including a grammar parser, a semantic disambiguation model, and a code generator; Rule confidence evaluation: A hybrid evaluation algorithm is used to calculate rule confidence, with the formula confidence = α model probability + β historical accuracy + γ expert score, where α + β + γ = 1; a dynamic weight adjustment formula β = 1 / (1 + e^(-kt)) is used, where k is the learning rate parameter and t is the number of times the rule is used; Self-optimization of rule base: Build a rule lifecycle management system to achieve a closed loop of abnormal pattern detection, rule evaluation, and version iteration; use an abnormal pattern detection algorithm based on an improved version of IsolationForest.
2. The method for generating intelligent data quality rules based on a large model according to claim 1, characterized in that: In the metadata-driven dynamic rule construction method, when extracting feature vectors through data lineage graph parsing technology, the method of MetadataParser class is used; Include extract_features for extracting features; _calculate_field_correlation calculates field correlation; _mine_business_constraints mines business constraints; And returns a FeatureSet containing the feature vectors and constraints.
3. The method for generating intelligent data quality rules based on a large model according to claim 1, characterized in that: In the multimodal rule conversion, the data for training the cross-modal conversion model includes 500,000 sets of NL-SQL / Python paired samples.
4. The method for generating intelligent data quality rules based on a large model according to claim 1, characterized in that: In the rule confidence assessment, in the financial risk control data set test scenario, the hybrid assessment formula used is Confidence = 0.4P_model + 0.3Accuracy + 0.2Expert + 0.1Freshness, and the learning rate parameter k in the dynamic weight adjustment algorithm is 0.
1.
5. The method for generating intelligent data quality rules based on a large model according to claim 1, characterized in that: In the rule base self-optimization mechanism, the rule lifecycle management system performs state transitions according to the rule confidence. When the rule confidence is greater than 0.8, it is in the Active state. When the confidence is less than 0.6 for three consecutive times or a new abnormal pattern is found, it enters the Deprecated state, and then iterates and updates the rule version.
6. An intelligent data quality rule generation system based on a large model, characterized in that: include: Metadata collection and analysis module: used to collect data metadata and extract feature vectors such as table structure, field association, business constraints, etc. through data lineage graph analysis technology to build rule generation context; Multimodal rule conversion engine module: Develops a cross-modal conversion model based on deep learning to achieve automatic conversion of natural language rules into executable code. It contains a syntax parser, a semantic disambiguation model, and a code generator. Rule confidence evaluation module: uses a hybrid evaluation algorithm and a dynamic weight adjustment formula to evaluate the confidence of the generated rules; Rule base self-optimization module: Build a rule lifecycle management system and use abnormal pattern detection algorithms to implement abnormal pattern detection, evaluation and version iteration of the rule base.
7. The intelligent data quality rule generation system based on a large model according to claim 6 is characterized in that: In the e-commerce data management scenario, the metadata collection and analysis module can collect order table metadata and analyze the relationship with the logistics table.
8. The intelligent data quality rule generation system based on a large model according to claim 6 is characterized in that: The multimodal rule conversion engine module can convert the natural language rule "order creation time should be earlier than logistics delivery time" into the SQL check statement "SELECT * FROM orders o JOIN logistics lONo.order_id=l.order_id WHERE o.create_time>l.delivery_time" in the e-commerce data governance scenario; In the financial risk control data scenario, the rule that "ID card number should conform to GB11643 format" can be converted into a Python verification function.
Citation Information
Patent Citations
Intelligent fusion and circulation method based on fraud early warning data
CN114003710A
Control method, analysis device, and recording medium
US20220365925A1
Circuit system and method
WO2022156680A1
Cited By
Intelligent data acquisition and conversion method based on visual rule model
CN120386814A
Quality rule configuration method based on artificial intelligence and related device thereof
CN120849471A
Dynamic data quality rule intelligent generation and self-adaptive correction system
CN120994656A