Data quality intelligent auditing system and method based on dynamic rule base

Through the intelligent audit system based on a dynamic rule base, the difficulties faced by traditional systems in adapting to changing business scenarios and dynamic data types have been solved, rapid rule development and real-time verification have been achieved, the false alarm rate has been reduced, and the efficiency of multi-source data governance and verification accuracy have been improved.

CN120763153APending Publication Date: 2025-10-10ZHUMADIAN POWER SUPPLY ELECTRIC POWER OFHENAN
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202510787614.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-10-10

Smart Images

  • Figure CN120763153A_ABST
    Figure CN120763153A_ABST
Patent Text Reader

Abstract

The invention discloses a data quality intelligent auditing system and method based on a dynamic rule base, and belongs to the technical field of data auditing, and the system comprises a rule base construction module which is used for analyzing business scene parameters through a scene analysis unit according to business scene demands and data type features to generate a rule configuration instruction; the multi-source monitoring engine module is connected to the rule base construction module and is used for collecting multi-source data in real time and loading corresponding checking rules; the automatic verification execution module is used for executing normalized quality verification on the multi-source data based on the verification rule base; and the feedback optimization module analyzes a rule hit rate and a false alarm rate in a verification result through a reinforcement learning algorithm, and dynamically iteratively updates a rule threshold value and a logic combination in the rule base. By constructing a full-automatic process of rule generation, execution, feedback and updating, the problems that a traditional system depends on manual intervention, response is slow, the industry average rule updating period is 3-7 days, and real-time updating is achieved through the scheme are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data auditing, and in particular relates to a data quality intelligent auditing system and method based on a dynamic rule library. Background Art

[0002] Currently, data quality audits are conducted to verify and control data quality during the data calculation process throughout the data lifecycle. In the end-to-end data visualization process of wireless networks, from the call order layer to the middle layer and finally to the application layer, multiple stages of calculation are required to obtain the final result. Auditing network data quality at each stage ensures the calculations for the next stage and the final data presentation. Traditional systems use fixed rule bases, which is like giving all patients the same medication. Financial data requires timeliness, medical data emphasizes privacy, and e-commerce data requires consistency; traditional systems cannot achieve differentiation. Multi-source, heterogeneous data is even more challenging. SQL and NoSQL have fundamentally different validation logic, and forcing unified processing is highly inefficient. Traditional audit systems use static rule templates (such as predefined SQL scripts) that cannot adapt to changing business scenarios and dynamic data types (structured, semi-structured, and streaming data), resulting in rigid rules. Audit rules must be independently developed for heterogeneous data sources (relational databases, NoSQL, and APIs), resulting in a large amount of duplicated development work, difficulty in cross-source data consistency verification, and inefficient multi-source data governance. Reliance on manual analysis reports → manual rule adjustments leads to delays in detecting new data anomalies (average response cycle > 72 hours), high false alarm rates (industry average false alarm rate exceeds 35%), and delayed rule updates. Summary of the Invention

[0003] The purpose of the present invention is to overcome the deficiencies of the prior art and to provide a data quality intelligent audit system and method based on a dynamic rule base, which solves the problems in the above-mentioned background technology.

[0004] The object of the present invention is achieved like this:

[0005] The application discloses a data quality intelligent auditing system based on a dynamic rule base, which comprises a rule base construction module, a multi-source monitoring engine module, an automatic checking execution module and a feedback optimization module.

[0006] Further, the scene analysis unit generates rule configuration instructions in the following manner: extracting key entity attributes and constraint conditions in the business scene, and generating a rule judgment tree through logical operator combination. The rule generation efficiency is improved, and the rule development speed is improved by 99% through rule judgment tree generation. The traditional rule development time is 4 hours per rule, and the rule development time of the application is 2 minutes per rule. Precise business matching, business scene parameter analysis accuracy is greater than or equal to 92%, and data accidents caused by rule mismatch in sensitive fields are avoided. Code error rate is zero, and the dynamic compiler automatically generates executable code to eliminate human coding errors.

[0007] Further, the multi-source monitoring engine comprises a data adapter supporting unified access of relational databases, NoSQL databases and API stream data; and a rule loader automatically matching the checking rule set in the rule base according to the data source type. By setting the unified interface of the data adapter, 15+ data source types are supported, the implementation period is shortened by 83%, and the implementation period is shortened from 6 months to 1 month. By setting the rule loader, millisecond-level rule matching is realized, the average response is 8ms, and the real-time auditing delay problem is solved.

[0008] Further, the working logic of the automatic checking execution module is as follows: starting a checking task according to a preset period or an event triggering mechanism, and outputting an auditing report containing data anomaly marks and error distribution statistics. The triggering mechanism is innovative, the dual mode of event driving and period scanning is adopted, and the abnormality capturing rate is improved. The report value is upgraded, the structured auditing report contains an error distribution heat map, the positioning problem efficiency is improved, and the fault troubleshooting time is shortened.

[0009] A data quality intelligent auditing method based on dynamic rule base, applied to the above system, comprising the following steps:

[0010] Receiving service scenario configuration instructions, generating rule configuration instructions by analyzing service scenario parameters, and dynamically compiling executable rules to build a personalized verification rule set;

[0011] Polling the access state of multi-source data, loading corresponding rules into the memory execution environment;

[0012] Triggering the data quality verification process, and outputting a structured audit report;

[0013] Analyzing the rule hit rate and false positive rate in the verification result through reinforcement learning algorithm, and dynamically updating the rule base iteratively.

[0014] Further, the specific operation of dynamically generating the verification rule set includes: analyzing the entity relationship topology in the business scenario, and the relationship topology model is:

[0015] G=(V,E,Φ)

[0016] In the formula, V={v i} is the entity node set; v i =entity name, data source type, C fields ; E={e jk} is the constraint edge set; e jk =<v j →v k ,τ,φ>; Φ={φ i} is the global constraint condition; the entity constraint is converted into SQL / NoSQL query statement, and the relationship topology model is used to realize the accurate conversion of business logic to rules.

[0017] Further, the data quality verification process includes: performing parallel verification on data integrity, consistency and timeliness through a multi-dimensional verification model, and optimizing the execution order through a rule dependency topology graph; the test objective function of the multi-dimensional verification model is:

[0018]

[0019] s.t.Re sourceUsage≤R max

[0020] In the formula, n is the total number of elements to be monitored; i is the element number; w i is the weight of the ith element in the system; v i is the monitored entity; s.t. is the constraint condition; ResourceUsage is the resource usage; Rmax is the maximum resource limit; the above formula is used to ensure that the key business is prioritized and the hardware resources are not wasted.

[0021] The weight distribution strategy is:

[0022]

[0023] The dependency graph adjacency matrix of the rule-dependent topology optimization model is defined as:

[0024] A=[a ij ] n×n

[0025] in, The execution order optimization condition is: the acyclic constraint det(IA) ≥ 0 (guaranteed directed acyclic graph), where det is the matrix determinant, I is the identity matrix, and A is the adjacency matrix;

[0026] Critical path acceleration: for the longest path L critical Prioritize the allocation of computing resources. The resource allocation formula is:

[0027]

[0028] Where R alloc is the amount of allocated resources; R max is the maximum available resource; T base is the single-thread execution time; T target is the target execution time; C core The number of CPU cores.

[0029] Furthermore, the rule update strategy includes: when a rule has not triggered an anomaly for N consecutive cycles, the execution priority of the rule is automatically downgraded; when a new data anomaly pattern is detected, a rule addition suggestion is generated and sent for manual review;

[0030] The health evaluation function of the rule degradation decision model is:

[0031]

[0032] Where H(r) is the rule health; α is the accuracy weight; β is the missed detection penalty weight; TP(r) is the number of times rule r correctly triggers anomalies; FP(r) is the number of times rule r incorrectly triggers anomalies or false alarms; N miss The number of times a true anomaly is detected for the rule; N total N is the total number of data samples checked for rule r; th is the missed detection cycle threshold;

[0033] The parameter constraints are:

[0034] α+β=1,α≥0.7,accuracy weight dominates; N miss ≥N th , continuous miss cycle threshold Nth ≥5;

[0035] The feature vector of the new anomaly detection model is defined as:

[0036] F abnormal = <f1, f2, f3> = <null value rate, standard deviation multiple, update time deviation> F abnormal is an anomaly feature vector;

[0037] The anomaly separation condition is:

[0038] ||F abnormal -μ known ||2>λ·σ known

[0039] In the formula, μ known is the mean vector of the known mode; σ known is the standard deviation of the known mode; ||·||2 is the Euclidean distance; λ is the confidence coefficient, λ≥2.58, corresponding to the 99% confidence interval.

[0040] The beneficial effects of the present application: by constructing the "rule generation, execution, feedback, and update" fully automated process, solving the slow response caused by the dependence of the traditional system on manual intervention, the industry average rule update cycle is 3-7 days, and the present scheme is updated in real time. It can adapt to variable business scenarios and dynamic data types, and optimizes the rule threshold through reinforcement learning to reduce the system false positive rate, realizes cross-source data consistency checking, and improves the efficiency of multi-source data management. Through rule decision tree generation, the development speed of new rules is improved by 99%, and the traditional 4 hours per rule is reduced to 2 minutes per rule. Precise business matching, business scenario parameter analysis accuracy ≥92%, avoids data accidents caused by rule mismatch in sensitive areas. Code error rate is zero, dynamic compiler automatically generates executable code, eliminating human coding errors. By setting the data adapter unified interface to support 15+ data source types, the implementation cycle is shortened by 83%, from 6 months to 1 month. Through the rule loader, millisecond-level rule matching is realized, with an average response of 8ms, solving the problem of real-time audit delay. Through the dual mode of event-driven and periodic scanning, the abnormal capture rate is improved. The value of the report is upgraded, and the structured audit report contains an error distribution heat map, which improves the positioning problem efficiency and shortens the troubleshooting time. BRIEF DESCRIPTION OF DRAWINGS

[0041] Figure 1 is a system flowchart of the present application. DETAILED DESCRIPTION

[0042] The present application will be further described in detail below with reference to the accompanying drawings, it should be pointed out that it is only for more clearly illustrating and explaining the present application.

[0043] Example 1

[0044] As Figure 1 shown, the embodiment discloses a data quality intelligent audit system based on dynamic rule base, which comprises: a rule base construction module for generating rule configuration instructions by analyzing business scenario parameters through a scenario analysis unit according to business scenario requirements and data type characteristics, and converting the instructions into executable rules by a dynamic compiler to dynamically develop a personalized audit rule base; the rule base construction module comprises a rule template library for storing the mapping relationship between predefined data types and rule logic. A multi-source monitoring engine module connected to the rule base construction module is used to collect multi-source data in real time and load corresponding audit rules; an automatic verification execution module performs normal quality verification on multi-source data based on the audit rule base; a feedback optimization module analyzes the rule hit rate and false positive rate in the verification result through a reinforcement learning algorithm to dynamically and iteratively update the rule threshold and logic combination in the rule base. By constructing a "rule generation, execution, feedback, and update" fully automated process, the slow response caused by the dependence of traditional systems on manual intervention is solved, the industry average rule update cycle is 3-7 days, and the present scheme is updated in real time. It can adapt to variable business scenarios and dynamic data types, optimize the rule threshold through reinforcement learning, reduce the system false positive rate, realize cross-source data consistency verification, and improve multi-source data governance efficiency.

[0045] A data quality intelligent audit method based on a dynamic rule base, applied to the above system, comprising the following steps: receiving a business scenario configuration instruction, generating a rule configuration instruction by analyzing business scenario parameters, and dynamically compiling executable rules to construct a personalized audit rule set; polling the multi-source data access state, loading corresponding rules to the memory execution environment; triggering the data quality verification process, outputting a structured audit report; analyzing the rule hit rate and false positive rate in the verification result through a reinforcement learning algorithm to dynamically and iteratively update the rule base.

[0046] Embodiment 2

[0047] As Figure 1As shown, this embodiment discloses a data quality intelligent audit system based on a dynamic rule base, which includes: a rule base construction module, which is used to parse business scenario parameters through a scenario analysis unit according to business scenario requirements and data type characteristics to generate rule configuration instructions, and a dynamic compiler converts the instructions into executable rules to dynamically develop a personalized verification rule base; the rule base construction module includes a rule template library, which is used to store the mapping relationship between predefined data types and rule logic. A multi-source monitoring engine module is connected to the rule base construction module and is used to collect multi-source data in real time and load corresponding verification rules; an automatic verification execution module performs normalized quality verification on multi-source data based on the verification rule base; a feedback optimization module analyzes the rule hit rate and false alarm rate in the verification results through a reinforcement learning algorithm, and dynamically iterates and updates the rule thresholds and logic combinations in the rule base. By building a fully automated process of "rule generation, execution, feedback, and update", the slow response caused by the traditional system's reliance on manual intervention is solved. The industry average rule update cycle is 3-7 days, and this solution is updated in real time. It can adapt to changing business scenarios and dynamic data types, optimize rule thresholds through reinforcement learning, reduce the system's false alarm rate, achieve cross-source data consistency verification, and improve the efficiency of multi-source data governance.

[0048] To achieve optimal results, the scenario analysis unit generates rule configuration instructions by extracting key entity attributes and constraints from the business scenario and generating a rule decision tree through logical operator combination. This improves rule generation efficiency. By generating a rule decision tree, the development of new rules is accelerated by 99%, from a traditional 4-hour process to 2 minutes per rule per rule. Precise business matching is achieved, with a business scenario parameter parsing accuracy of ≥92%, preventing data incidents caused by rule mismatches in sensitive areas. Code errors are reduced to zero, with a dynamic compiler automatically generating executable code, eliminating manual coding errors.

[0049] To achieve optimal results, the multi-source monitoring engine includes a data adapter that supports unified access to relational databases, NoSQL databases, and API streaming data; and a rule loader that automatically matches verification rule sets in the rule library based on the data source type. By setting up a unified data adapter interface to support 15+ data source types, the implementation cycle was shortened by 83%, from 6 months to 1 month. The rule loader achieves millisecond-level rule matching, with an average response time of 8ms, addressing real-time audit delays.

[0050] To achieve optimal results, the automated verification execution module's operating logic is to initiate verification tasks based on a preset cycle or event trigger mechanism, and output an audit report containing data anomaly markers and error distribution statistics. This innovative trigger mechanism, featuring both event-driven and periodic scanning modes, improves the anomaly capture rate. Furthermore, the report's value is enhanced with a structured audit report that includes an error distribution heat map, improving problem location efficiency and shortening troubleshooting time.

[0051] An intelligent data quality audit method based on a dynamic rule base is applied to the above-mentioned system and includes the following steps: receiving business scenario configuration instructions, generating rule configuration instructions by parsing business scenario parameters, and dynamically compiling them into executable rules to build a personalized verification rule set; polling and monitoring the access status of multi-source data, and loading corresponding rules into the memory execution environment; triggering the data quality verification process and outputting a structured audit report; analyzing the rule hit rate and false alarm rate in the verification results through a reinforcement learning algorithm, and dynamically iterating and updating the rule base.

[0052] To achieve better results, the specific operations of dynamically generating the verification rule set include: parsing the entity relationship topology in the business scenario, and the relationship topology model is:

[0053] G=(V,E,Φ)

[0054] Where, V={v i} is the entity node set; v i =Entity name, data source type, C fields ; E={e jk} is the constraint edge set; e jk = <v j →v k ,τ,φ>;Φ={φ i} is a global constraint condition; entity constraints are converted into SQL / NoSQL query statements, and the relational topology model is used to achieve accurate conversion of business logic into rules.

[0055] To achieve better results, the data quality verification process includes: parallel verification of data integrity, consistency, and timeliness through a multi-dimensional verification model, and optimizing the execution order through a rule-dependent topology graph; the verification objective function of the multi-dimensional verification model is:

[0056]

[0057] stResourceUsage≤Rmax

[0058] Where n is the total number of elements to be monitored; i is the element number; w i is the weight of the i-th element in the system; v i is the monitored entity; st is the constraint condition; ResourceUsage is the resource usage; Rmax is the maximum resource limit; the above formula is used to ensure priority management of key businesses and achieve zero waste of hardware resources.

[0059] The weight distribution strategy is:

[0060]

[0061] The dependency graph adjacency matrix of the rule-dependent topology optimization model is defined as:

[0062] A=[a ij ] n×n

[0063] in, The execution order optimization condition is: the acyclic constraint det(IA) ≥ 0 (guaranteed directed acyclic graph), where det is the matrix determinant, I is the identity matrix, and A is the adjacency matrix;

[0064] Critical path acceleration: for the longest path L critical Prioritize the allocation of computing resources. The resource allocation formula is:

[0065]

[0066] Where R alloc is the amount of allocated resources; R max is the maximum available resource; T base is the single-thread execution time; T target is the target execution time; C core The number of CPU cores.

[0067] To achieve better results, the rule update strategy includes: automatically downgrading the execution priority of a rule if it has not triggered an exception for N consecutive cycles; generating rule addition suggestions and sending them for manual review when new data anomaly patterns are detected;

[0068] The health evaluation function of the rule degradation decision model is:

[0069]

[0070] Where H(r) is the rule health; α is the accuracy weight; β is the missed detection penalty weight; TP(r) is the number of times rule r correctly triggers anomalies; FP(r) is the number of times rule r incorrectly triggers anomalies or false alarms; N miss The number of times a true anomaly is detected for the rule; N total N is the total number of data samples checked for rule r; th is the missed detection cycle threshold;

[0071] The parameter constraints are:

[0072] α+β=1,α≥0.7,accuracy weight dominates; N miss ≥N th , continuous miss cycle threshold N th ≥5;

[0073] The feature vector of the new anomaly detection model is defined as:

[0074] F abnormal =<f1,f2,f3> = <empty value rate, standard deviation multiple, update time deviation> F abnormal is the abnormal feature vector;

[0075] The abnormal separation conditions are:

[0076] ||F abnormal -μ known ||2>λ·σ known

[0077] Where μ known is the mean vector of the known pattern; σ known is the standard deviation of the known pattern; ||·||2 is the Euclidean distance; λ is the confidence coefficient, λ≥2.58 corresponds to a 99% confidence interval.

[0078] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes based on the technical solutions and concepts of the present invention within the technical scope disclosed by the present invention, and they should be covered by the scope of protection of the present invention.

Claims

1. A data quality intelligent audit system based on a dynamic rule base, characterized by: include: The rule base construction module is used to generate rule configuration instructions by parsing business scenario parameters through the scenario analysis unit according to business scenario requirements and data type characteristics, and convert the instructions into executable rules by the dynamic compiler to dynamically develop a personalized verification rule base; A multi-source monitoring engine module, connected to the rule base construction module, for collecting multi-source data in real time and loading corresponding verification rules; An automatic verification execution module performs normalized quality verification on multi-source data based on the verification rule base; The feedback optimization module analyzes the rule hit rate and false alarm rate in the verification results through the reinforcement learning algorithm, and dynamically iterates and updates the rule thresholds and logical combinations in the rule base.

2. The data quality intelligent audit system based on a dynamic rule base according to claim 1 is characterized by: The scenario analysis unit generates rule configuration instructions in the following manner: extracting key entity attributes and constraints in the business scenario, and generating a rule decision tree through logical operator combination.

3. The data quality intelligent audit system based on a dynamic rule base according to claim 1 is characterized by: The multi-source monitoring engine includes: a data adapter that supports unified access to relational databases, NoSQL databases and API stream data; a rule loader that automatically matches the verification rule set in the rule library according to the data source type.

4. The data quality intelligent audit system based on a dynamic rule base according to claim 1 is characterized by: The working logic of the automatic verification execution module is: starting the verification task according to a preset cycle or event trigger mechanism, and outputting an audit report containing data anomaly marks and error distribution statistics.

5. The data quality intelligent audit system based on a dynamic rule base according to any one of claims 1 to 4, characterized in that: The rule base construction module includes a rule template library for storing mapping relationships between predefined data types and rule logic.

6. A data quality intelligent audit method based on a dynamic rule base, characterized in that: The system according to any one of claims 1 to 5 comprises the steps of: Receive business scenario configuration instructions, generate rule configuration instructions by parsing business scenario parameters, and dynamically compile them into executable rules to build a personalized verification rule set; Poll and monitor the access status of multi-source data and load corresponding rules into the memory execution environment; Trigger the data quality verification process and output a structured audit report; The rule hit rate and false alarm rate in the verification results are analyzed through reinforcement learning algorithms, and the rule base is updated dynamically and iteratively.

7. The data quality intelligent audit method based on the dynamic rule base according to claim 6 is characterized in that: The specific operations of dynamically generating the verification rule set include: parsing the entity relationship topology in the business scenario. The relationship topology model is: G=(V,E,Φ) Where, V={v i } is the entity node set; v i =Entity name, data source type, C fields ; E={e jk } is the constraint edge set; e jk = <v j →v k ,τ,φ>;Φ={φ i } is a global constraint condition; entity constraints are converted into SQL / NoSQL query statements.

8. The data quality intelligent audit method based on a dynamic rule base according to claim 6 is characterized in that: The data quality verification process includes: parallel verification of data integrity, consistency, and timeliness through a multi-dimensional verification model, and optimizing the execution order through a rule-dependent topology graph; the verification objective function of the multi-dimensional verification model is: stResourceUsage≤Rmax Where n is the total number of elements to be monitored; i is the element number; w i is the weight of the i-th element in the system; v i is the monitored entity; st is the constraint; ResourceUsage is the resource usage; Rmax is the maximum resource limit; The weight distribution strategy is: The dependency graph adjacency matrix of the rule-dependent topology optimization model is defined as: in, The execution order optimization condition is: the acyclic constraint det(IA) ≥ 0 (guaranteed directed acyclic graph), where det is the matrix determinant, I is the identity matrix, and A is the adjacency matrix; Critical path acceleration: for the longest path L critical Prioritize the allocation of computing resources. The resource allocation formula is: Where R alloc is the amount of allocated resources; R max is the maximum available resource; T base is the single-thread execution time; T target is the target execution time; C core The number of CPU cores.

9. The data quality intelligent audit method based on a dynamic rule base according to claim 6 is characterized in that: The rule update strategy includes: automatically downgrading the execution priority of a rule when it has not triggered an anomaly for N consecutive cycles; generating a rule addition suggestion and sending it for manual review when a new data anomaly pattern is detected; The health evaluation function of the rule degradation decision model is: Where H(r) is the rule health; α is the accuracy weight; β is the missed detection penalty weight; TP(r) is the number of times rule r correctly triggers anomalies; FP(r) is the number of times rule r incorrectly triggers anomalies or false alarms; N miss The number of times a true anomaly is detected for the rule; N total N is the total number of data samples checked for rule r; th is the missed detection cycle threshold; The parameter constraints are: α+β=1,α≥0.7,accuracy weight dominates; N miss ≥N th , continuous miss cycle threshold N th ≥5; The feature vector of the new anomaly detection model is defined as: F abnormal =<f1,f2,f3> = <empty value rate, standard deviation multiple, update time deviation> F abnormal is the abnormal feature vector; The abnormal separation conditions are: ||F abnormal -m known ||2>l·s known Where μ known is the mean vector of the known pattern; σ known is the standard deviation of the known pattern; ||·||2 is the Euclidean distance; λ is the confidence coefficient, λ≥2.58 corresponds to a 99% confidence interval.

Citation Information

Cited By

  • System and method for efficiently decoding RRC (Radio Resource Control) message based on dynamic configuration

    CN120915862A

  • Dynamic data quality rule intelligent generation and self-adaptive correction system

    CN120994656A

  • Data quality management system and method based on data contract and manual cooperation

    CN121279756A

  • Large language model output quality monitoring method and device, program and storage medium

    CN121327447A

  • Automatic quality inspection method, system and equipment based on low-efficiency land use data and medium

    CN121413624A