Fourth-order intelligent engine driven closed-loop data quality management method and system

The closed-loop data quality management method driven by a four-stage intelligent engine solves the problems of fragmented requirements, rigid rules, and lack of governance closed loop in power data quality management. It achieves high efficiency in intelligent clustering of requirements, full lifecycle autonomy of rules, and verification execution, thereby improving the systematic nature and business integration capabilities of data quality management.

CN122132389APending Publication Date: 2026-06-02STATE GRID INFORMATION & TELECOMM BRANCH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID INFORMATION & TELECOMM BRANCH
Filing Date
2026-02-05
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In power data quality management, there are problems such as fragmented requirements, rigid rules, inefficient execution, and lack of governance closed loop. It is difficult to achieve unified governance and dynamic adaptability across disciplines and departments, resulting in low efficiency of data quality management and limited business integration.

Method used

A closed-loop data quality management method based on a fourth-order intelligent engine is adopted. Natural language processing technology is used to intelligently cluster requirements, and a governance priority decision matrix is ​​generated by combining the full-link impact scope. Quality verification rules are constructed using a dual-engine driven mode of artificial intelligence generation and low-code orchestration. Distributed task scheduling and multi-dimensional root cause diagnosis are carried out based on dynamic baseline thresholds to establish a mapping relationship between governance effectiveness and business value.

Benefits of technology

It has achieved intelligent and collaborative decision-making in demand management, full lifecycle autonomy in rule management, high efficiency in verification and execution, and automation of governance closed loop, which has improved the systematic nature and business integration capabilities of data quality management and formed a closed-loop governance method driven by a four-stage intelligent engine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132389A_ABST
    Figure CN122132389A_ABST
Patent Text Reader

Abstract

This invention provides a closed-loop data quality management method and system driven by a fourth-order intelligent engine. Based on the acquired business data governance requirements, natural language processing technology is used to intelligently cluster the requirements, and a governance priority decision matrix is ​​generated by combining the full-link impact scope and economic impact assessment. Quality verification rules are constructed according to the business data governance requirements through a dual-engine driven mode combining artificial intelligence generation and low-code orchestration. Based on the governance priority decision matrix and data lineage, the quality verification rules are aggregated and packaged into rule verification task groups. Based on dynamic baseline thresholds, a distributed task scheduling framework is used to execute verification tasks on incremental data in each rule verification task group. Based on the full-link impact scope, multi-dimensional root cause diagnosis and automated governance are performed on the verified abnormal data. Based on the governed business data, a mapping relationship between governance effectiveness and business value is established to generate a value assessment report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power grids, and specifically to a closed-loop data quality management method and system based on a fourth-order intelligent engine. Background Technology

[0002] With the deepening of digital transformation, data has become a core production factor, playing a crucial role in business decision-making, trend prediction, and intelligent applications. In the power industry, big data applications built on data platforms involve data sharing across businesses, systems, levels, regions, and departments, placing higher demands on data quality. However, due to the dispersed historical construction and inconsistent standards of business systems, data platforms face problems such as poor data standardization, lack of integrity, and consistency conflicts when integrating multi-source heterogeneous data, resulting in low credibility of cross-professional data and limited application effectiveness.

[0003] Currently, power data quality management mainly relies on manual verification or static rule validation, lacking systematic evaluation methods and dynamic control capabilities, making it difficult to meet the needs of real-time and automated governance. Therefore, there is an urgent need to build a unified data quality management methodology, through a standardized evaluation system, the construction of an intelligent rule base, and a closed-loop control mechanism, to improve data quality and support the efficient sharing and value mining of cross-domain data.

[0004] Currently, various industries have developed research methods and systems for data quality management. By establishing unified data quality standards and verification rules, they can systematically measure the accuracy, completeness, and consistency of data quality. These methods also incorporate data quality tools to promptly detect data problems and notify relevant personnel, enabling rapid response and handling of quality issues. Furthermore, data quality assessment reports allow users to intuitively understand the trends in data quality problems, providing strong support for business decision-making. Some methods also integrate data visualization techniques, such as parallel coordinate graphs and scatter plot matrices, to help users analyze data quality problems more intuitively.

[0005] However, existing methods still have some limitations: 1. Lack of Requirements Management and Insufficient Collaboration: The current system lacks a unified governance requirements management mechanism, resulting in fragmented quality information across disciplines and departments, making it difficult to form a holistic view. Governance requirements are often submitted repeatedly, and there is a lack of effective priority assessment and coordination mechanisms, affecting the rational allocation of resources and the efficiency of problem-solving.

[0006] 2. Rigid rule management and weak adaptability: Data quality rules mostly rely on manual static configuration and cannot dynamically adapt to new data quality problems (such as abnormal data association graphs, semantic conflicts, etc.), thus failing to achieve adaptive optimization and iteration of rules.

[0007] 3. Low execution efficiency and high resource consumption: Massive rules typically require the creation of independent scheduling tasks, leading to task bloat and scheduling redundancy, which puts significant pressure on the data platform's computing and storage resources. Furthermore, the mechanisms for processing and verifying incremental data are not robust enough to effectively handle scenarios with rapidly growing data volumes and high real-time requirements.

[0008] 4. Lack of a closed-loop governance system, and limited automation and business integration: In the problem-solving phase, data transmission anomalies often rely on manual intervention by implementation personnel, while source-end business quality issues depend on manual troubleshooting and repair by frontline staff. Automation and intelligence levels need improvement, especially when dealing with complex data scenarios and dynamic data environments, where existing methods are insufficiently adaptable and lack end-to-end autonomous closed-loop capabilities. Furthermore, the lack of in-depth analysis and business integration capabilities for data quality issues prevents the provision of forward-looking optimization suggestions and decision support to business users, limiting the initiative and strategic value of quality management.

[0009] Therefore, it is of great significance to address the issues of fragmented requirements, rigid rules, inefficient execution, and lack of governance loop in power data quality management. Summary of the Invention

[0010] To address the problems of fragmented requirements, rigid rules, inefficient execution, and lack of governance closure in existing power data quality management technologies, this invention proposes a closed-loop data quality management method and system driven by a fourth-order intelligent engine.

[0011] Firstly, a closed-loop data quality management method based on a fourth-order intelligent engine is provided, including: Based on the acquired business data governance needs, natural language processing technology is used to intelligently cluster the needs, and a governance priority decision matrix is ​​generated by combining the determination of the full-link impact scope and economic impact assessment. The full-link impact scope is determined based on the data to be governed through metadata lineage graph. Based on the acquired business data governance needs, quality verification rules are constructed using a dual-engine driven model that combines artificial intelligence generation and low-code orchestration. Based on the governance priority decision matrix and data lineage correlation, the quality verification rules are aggregated and packaged into rule verification task groups. Based on the dynamic baseline threshold, a distributed task scheduling framework is used to execute verification tasks on the incremental data in each rule verification task group. The dynamic baseline threshold is generated by obtaining the historical data quality fluctuation pattern. Based on the full-link impact scope, the abnormal data identified is subjected to multidimensional root cause diagnosis and automated governance, and a value assessment report is generated based on the governance effectiveness and business value mapping relationship established on the governed business data.

[0012] Secondly, a closed-loop data quality management system based on a fourth-order intelligent engine is provided, including: The clustering module is used to intelligently cluster the requirements based on the acquired business data governance needs using natural language processing technology, and generate a governance priority decision matrix by combining the full-link impact scope and economic impact assessment. The full-link impact scope is determined based on the data to be governed through metadata lineage graph. The building module is used to construct quality verification rules based on the acquired business data governance needs through a dual-engine driven mode that combines artificial intelligence generation and low-code orchestration. The verification module is used to aggregate and package quality verification rules into rule verification task groups based on the governance priority decision matrix and data lineage correlation, and to perform verification tasks on incremental data in each rule verification task group using a distributed task scheduling framework based on a dynamic baseline threshold. The dynamic baseline threshold is generated by obtaining historical data quality fluctuation patterns. The governance module is used to perform multidimensional root cause diagnosis and automated governance of the abnormal data identified based on the full-link impact scope, and to generate a value assessment report based on the mapping relationship between governance effectiveness and business value established by the governed business data.

[0013] In another aspect, this application also provides an electronic device, comprising: at least one processor and a memory; the memory and the processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, a closed-loop data quality management method based on a fourth-order intelligent engine, as described above, is implemented.

[0014] In another aspect, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements a closed-loop data quality management method based on a fourth-order intelligent engine as described above.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides a closed-loop data quality management method and system driven by a fourth-order intelligent engine. The method uses natural language processing technology to intelligently cluster the acquired business data governance requirements, and generates a governance priority decision matrix by combining the full-link impact scope and economic impact assessment. The full-link impact scope is determined based on the data to be governed through metadata lineage mapping. Based on the acquired business data governance requirements, quality verification rules are constructed using a dual-engine driven mode combining artificial intelligence generation and low-code orchestration. These rules are then aggregated and packaged into rule verification task groups based on the governance priority decision matrix and data lineage correlation. Based on dynamic baseline thresholds, a distributed task scheduling framework is used to execute verification tasks on incremental data in each rule verification task group. Furthermore, based on the full-link impact scope, multi-dimensional root cause diagnosis and automated governance are performed on the verified abnormal data. Finally, a value assessment report is generated based on the mapping relationship between governance effectiveness and business value established using the governed business data. This method solves the systemic defects of existing data quality management, such as vague and scattered requirement management, fragmented rule management, high coupling of verification execution, lack of governance closed loop, and limited business integration, thus forming a closed-loop governance method driven by a fourth-order intelligent engine. Attached Figure Description

[0016] Figure 1 The flowchart shows the closed-loop data quality management method based on a fourth-order intelligent engine according to the present invention. Figure 2 This is a schematic diagram of the execution flow of the closed-loop data quality management method based on a fourth-order intelligent engine driven by the present invention. Figure 3 This is a schematic diagram of AI-driven collaborative decision-making for the closed-loop data quality management method based on a fourth-order intelligent engine, as described in this invention. Figure 4 This is a schematic diagram of the dual-engine driven rule management of the closed-loop data quality management method based on a fourth-order intelligent engine according to the present invention. Figure 5 This is a schematic diagram illustrating the quality check based on a dynamic baseline engine and intelligent task scheduling in the closed-loop data quality management method driven by a fourth-order intelligent engine according to the present invention. Figure 6 This is a schematic diagram illustrating the collaborative autonomy based on value community and intelligent diagnosis in the closed-loop data quality management method driven by a fourth-order intelligent engine according to the present invention. Figure 7 This is a schematic diagram of the structure of the closed-loop data quality management system based on a fourth-order intelligent engine according to the present invention. Figure 8 This is a schematic diagram of an electronic device structure according to the present invention. Detailed Implementation

[0017] This invention provides a closed-loop data quality management method and system driven by a four-stage intelligent engine. It aims to optimize the core deficiencies in existing data quality management, such as missing requirement management, fragmented rule management, high coupling of verification and execution, and broken governance loops. By constructing a "four-stage intelligent engine" architecture, it achieves a paradigm shift: 1. Addressing the pain points of fragmented requirement information, duplicate submissions, and ambiguous priorities, a requirement collaborative decision-making engine is built to achieve intelligent requirement clustering and governance priority assessment; 2. Addressing the pain points of scattered rules, high conflict rates, and low reuse rates, a pioneering dual-track rule factory engine is created, integrating AI natural language to SQL conversion and low-code orchestration to achieve full rule lifecycle management. 3. To address the system resource contention and task delays caused by the execution of massive rules, a dynamic verification platform engine is designed. Based on incremental computing and distributed scheduling, it achieves parallel verification of millions of rules at the minute level, and introduces a time-series prediction dynamic baseline to capture data anomalies in real time; 4. To address the lag in governance response and the recurrence of problems, a collaborative autonomous engine connecting technical governance and business value is created. Through a three-layer linkage model of "intelligent diagnosis - private domain autonomy - value transformation", it achieves a fundamental transformation from passive alarm to proactive governance, and from technical control to business empowerment, solving the core pain points of lack of governance closed loop, low business participation, and difficulty in realizing value.

[0018] To better understand the present invention, the following description, in conjunction with the accompanying drawings and embodiments, will further illustrate the content of the present invention.

[0019] Example 1: A closed-loop data quality management method driven by a fourth-order intelligent engine, such as... Figure 1 As shown, it includes: Step 1: Based on the acquired business data governance requirements, use natural language processing technology to intelligently cluster the requirements, and combine the full-link impact scope and economic impact assessment to generate a governance priority decision matrix; Step 2: Based on the acquired business data governance requirements, construct quality verification rules using a dual-engine driven mode that combines AI generation and low-code orchestration. Step 3: Based on the governance priority decision matrix and data lineage correlation, aggregate and package the quality verification rules into rule verification task groups, and use a distributed task scheduling framework to perform verification tasks on the incremental data in each rule verification task group based on the dynamic baseline threshold; Step 4: Based on the full-link impact scope, perform multidimensional root cause diagnosis and automated governance on the abnormal data identified, and generate a value assessment report based on the governance effectiveness and business value mapping relationship established based on the governed business data.

[0020] Among them, the full-link impact range is determined based on the data to be governed through metadata lineage mapping; the dynamic baseline threshold is generated by obtaining historical data quality fluctuation patterns.

[0021] In this embodiment, in step 1, based on the acquired business data governance requirements, natural language processing technology is used to perform intelligent clustering of requirements, and the full-link impact range of the data to be governed is determined through an integrated metadata lineage graph. Then, in the process of generating a governance priority decision matrix in conjunction with economic impact assessment, clustering and merging can be performed by calculating the similarity between semantic vectors extracted using natural language processing technology to achieve intelligent requirement clustering. This provides a foundation for the subsequent generation of the governance priority decision matrix and the diagnosis and governance of subsequent abnormal data. Specifically, this includes: Based on the constructed natural language processing requirement parsing engine, a hybrid model of BERT and bidirectional long short-term memory network is used to identify and extract semantic vectors from the acquired business data governance requirements; The unstructured requirements in the semantic vectors are converted into high-dimensional semantic vectors, and the similarity between the vectors is calculated. When the similarity exceeds a preset threshold, the vectors are merged to generate a clustering analysis report. Construct the upstream and downstream impact chain of data assets based on metadata lineage graphs to determine the full-link impact scope of data to be governed; A tiered impact indicator system is constructed, which includes direct economic losses, regulatory risks, and business downtime costs. Based on the tiered impact indicator system, a weighted scoring algorithm is used to calculate the score of each business data governance requirement. Based on the scoring results, the business data governance requirements are classified into levels to generate multi-level importance requirements. A priority decision matrix is ​​generated based on the clustering analysis report, the full-link impact scope, and the multi-level importance requirements.

[0022] BERT can be understood as a pre-trained language model based on the Transformer architecture.

[0023] In one specific implementation, NLP (Natural Language Processing) technology is used to automatically parse the original requirements submitted by business departments (such as missing equipment measurement data), identify keywords, and associate them with similar historical requirements to generate clustering analysis reports, achieving intelligent requirement clustering. Furthermore, a metadata lineage graph is integrated to visually display the systems, tables, fields, and upstream and downstream impacts involved in the requirements (e.g., abnormal transformer file data will lead to deviations in electricity settlement), assisting in quickly reaching a governance consensus across departments. Based on historical governance data, an economic impact assessment algorithm is constructed (e.g., methods such as classifying the degree of economic impact according to data quality issues), automatically outputting a requirement priority matrix to guide the precise allocation of governance resources. Through these methods, the problems of information fragmentation, duplicate submissions, and ambiguous priorities inherent in traditional requirement management processes that rely on manual collection of Excel spreadsheets can be solved, thereby building an AI-driven requirement collaboration platform.

[0024] In this embodiment, during the process of constructing quality verification rules in step 2 based on the acquired business data governance requirements using a dual-engine driven mode combining artificial intelligence generation and low-code orchestration, the development efficiency can be improved and the technical threshold lowered by using a dual-engine driven mode that combines artificial intelligence generation and low-code orchestration for the logical requirements in the acquired business data governance requirements. Specifically, this includes: Based on the complex logical requirements in the acquired business data governance needs, the natural language descriptions are parsed by the deep learning model in the artificial intelligence generation engine and automatically converted into a preset code framework, and the association rules recommended based on the knowledge graph are used to assist in the verification. Based on the general logical requirements in the acquired business data governance needs, the system generates rule parameters in response to user configuration operations through a visual drag-and-drop interface, and automatically performs field type matching degree detection and threshold validity verification during the configuration process. The preset code framework and the rule parameters are used to construct quality verification rules.

[0025] In one specific embodiment, based on the input natural language requirements, artificial intelligence technology is used to automatically parse them into technical rule logic, generate executable SQL / Python code drafts, and recommend existing relevant rules based on knowledge graphs (such as automatically associating voltage anomaly rules with current fluctuation detection). After manual verification, the rules are stored in the database to prevent rule blind spots and improve development efficiency. Secondly, a visual drag-and-drop interface is provided with built-in quality verification templates (uniqueness, code value compliance, etc.), allowing business personnel to directly configure rules and lower the technical threshold.

[0026] In this embodiment, after constructing the quality verification rules in step 2, the quality verification rules can be stored in the corresponding rule base, and real-time scanning and other operations can be performed on the rule base to ensure the real-time performance and security of the rules. Specifically, this includes: The system scans the rule base corresponding to the quality verification rules in real time and uses logical reasoning algorithms to identify conflicting rules. Based on data lineage tracking and source metadata changes, when a field in the source system changes, an automatic failure rule warning is triggered and optimization suggestions are generated.

[0027] In one specific embodiment, the rule base is scanned in real time to detect conflicting rules (such as rule A requiring non-empty rules and rule B allowing empty rules) and invalid rules (source system field changes), and automatic alerts are issued and optimization solutions are recommended.

[0028] In summary, by using an AI rule generator, low-code rule orchestration, and real-time rule health monitoring, we can solve the problems of inconsistent rule standards and low development efficiency, thereby creating a dual-engine rule tool that combines "AI + low-cost code".

[0029] In this embodiment, after obtaining the governance priority decision matrix and quality verification rules through the aforementioned steps 1 and 2, the quality verification rules can be grouped based on the priority decision matrix and quality verification rules, combined with a dynamic baseline threshold generated based on historical data quality fluctuation patterns. Incremental data can then be checked based on the corresponding priorities, reducing verification time and further improving problem location efficiency. Specifically, this includes: A time series prediction algorithm is used to learn the fluctuation patterns of historical data quality indicators, establish a dynamic pass rate baseline, and automatically train a model to generate a dynamic threshold range corresponding to the quality verification rules. Based on the priorities in the governance priority decision matrix, combined with data lineage and business scenarios, multiple quality verification rules with strong correlations are grouped to generate multiple rule verification task groups. Using a distributed task scheduling framework, computing resources are dynamically allocated to each rule verification task group according to rule priority and computational complexity. Incremental data is extracted through preset extraction technology, and verification tasks are executed based on the real-time quality indicators of the incremental data and the dynamic threshold range.

[0030] In one specific embodiment, based on a time-series prediction algorithm, the system automatically learns the historical data quality fluctuation patterns, generates a dynamic pass rate threshold, and issues immediate alarms for fluctuations exceeding the threshold. A distributed task scheduling framework is employed to intelligently group and package large-scale rules, dynamically allocate computing resources according to priority, and introduce incremental verification technology to scan only changed data (such as newly added device records for the day), reducing verification time. This method changes the inefficient traditional "one rule, one task" model, constructs an intelligent verification platform, reduces full-scale verification time, and improves problem location efficiency.

[0031] In this embodiment, after abnormal data is detected through the verification task based on step 3, the abnormal data can be traced back to its source and diagnosed using the full-link impact scope obtained in step 1. Then, automated governance is performed based on the diagnosis results, and a value assessment report is generated based on the governed business data. This achieves automatic analysis and closed-loop processing of abnormal root causes, specifically including: When an abnormal alarm is detected, the full-link context information of the abnormal data is automatically associated within the full-link impact range. The context information includes checking recent version changes of the system to which the data belongs, ETL task logs, and source operation records. Based on the aforementioned metadata lineage map, upstream and downstream dependency issues are located within the entire influence chain, and a root cause analysis report is generated. The root cause analysis report is parsed, and the pre-set programmable problem library is matched according to the parsed fault type to trigger an automatic repair script.

[0032] In this embodiment, after generating a value assessment report by establishing a mapping relationship between governance effectiveness and business value based on the governed business data, a six-dimensional quality profile can be constructed for each piece of business data for visualization. A private governance portal can also be opened to business departments, thereby surpassing the traditional ranking and reporting model and constructing a governance value transformation engine. Specifically, this includes: A six-dimensional quality profile is constructed for each business data item, and the profile is visualized using a topology diagram. Open the private governance portal to business departments and subordinate units to customize governance goals and indicators; Based on the obtained economic impact assessment results and demand analysis report, a mapping model between governance effectiveness and business value is established, and a quarterly value white paper quantifying business revenue is generated based on the mapping model.

[0033] In one specific implementation, a quality profile is built for each table in the resource catalog, integrating six-dimensional health indices such as completeness and validity. Weaknesses are visually displayed through a topology diagram. Private dashboards are made available to lower-level units, supporting customized governance goals. When an abnormal alarm occurs (such as a 30% drop in the pass rate of line current data in a certain province), the system automatically correlates the recent version of the system to which the data belongs, task logs, and source operation records, outputting a root cause analysis report and performing automated processing. Furthermore, a governance effectiveness-business value mapping model is established, generating quarterly value white papers to drive proactive participation from business departments in governance. This surpasses the traditional ranking and reporting model, constructing a governance value transformation engine.

[0034] This invention has the following innovative features: 1. Intelligent demand decision-making It pioneered a three-dimensional collaborative decision-making model integrating "intelligent demand clustering, impact projection, and value quantification." Through NLP technology, it automatically parses unstructured demands and accurately matches them with historically similar demands, solving the problem of information fragmentation. It integrates metadata lineage graphs for dynamic visualization of the impact chain, breaking down cross-departmental collaboration barriers. An economic impact assessment algorithm maps data quality issues to economic loss levels, automatically generating a resource allocation priority matrix. This completely transforms the traditional manual Excel-based data aggregation model, enabling full autonomy from fragmented demand reporting to scientific decision-making.

[0035] 2. Rule-based full lifecycle autonomy We build a rule factory driven by a dual engine of "AI natural language to SQL + low-code visualization" to achieve full lifecycle autonomous management from automatic requirement parsing, code generation, knowledge graph recommendation to conflict self-checking and health monitoring, thus solving the pain points of inconsistent rule standards, low development efficiency and high conflict rate.

[0036] 3. Lightweight Distributed Intelligent Verification The design incorporates a distributed scheduling framework of "rule intelligent grouping - incremental scanning - dynamic resource allocation". By combining time-series prediction to generate dynamic baselines, it scans only changing data, enabling minute-level parallel verification and predictive alarms for millions of rules, thus breaking through the high-coupling execution bottleneck of the traditional "one rule, one task" approach.

[0037] 4. Business collaboration and autonomous closed loop Through the mechanism of "private domain quality profiling + target customization + value mapping", business departments are transformed from "those being governed" to "co-formulators of governance", and a closed loop of intelligent diagnosis and automated handling of anomalies is established, realizing a fundamental transformation from technical control to business empowerment, and from passive alarms to proactive governance.

[0038] Example 2: The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention.

[0039] like Figure 2 As shown, Figure 2 This is a schematic diagram illustrating the process execution of the closed-loop data quality management system method driven by a fourth-order intelligent engine according to the present invention, including: 1. AI-driven demand-driven collaborative decision-making module This invention employs an intelligent clustering and collaborative decision-making platform to achieve demand management, such as... Figure 3 As shown, the specific implementation method is as follows: First, we collect data governance needs from business departments through online system submissions or offline collection of Excel files for import.

[0040] Second, intelligent demand clustering is implemented. This involves constructing an NLP demand parsing engine and a high-efficiency vector retrieval engine to automate the merging and analysis of unstructured demands. The NLP demand parsing engine is built using a BERT+BiLSTM hybrid model. By training a dedicated thesaurus of power business entity words and combining it with a regular expression template library, it accurately identifies typical problem expressions such as "missing equipment measurement data." Based on the FAISS framework's demand similarity calculation model, with an initial dynamic clustering threshold of 0.85, unstructured demands are converted into 384-dimensional semantic vectors and then intelligently merged. This automatically generates a clustering analysis report containing statistics on similar demands, rankings of high-frequency issues, and impact maps of related systems, providing data insights for governance decisions.

[0041] Thirdly, metadata visualization is implemented. The system integrates with the data platform's metadata service and utilizes the Neo4j graph database to construct a lineage network of all system data tables and fields. By developing positive impact tracing and reverse attribution analysis algorithms, a bidirectional analysis model for data quality requirements is built. A visualization method combining Echarts relationship diagrams and Sankey diagrams is used to clearly display the impact scope of problematic fields and the potential root causes of anomalies.

[0042] Fourth, conduct economic impact assessments and establish a three-tiered impact indicator system that includes direct economic losses, regulatory risks, and business downtime costs, transforming abstract data quality issues into comparable business value impacts. A weighted scoring algorithm (economic loss weight 0.6, impact scope 0.3, urgency level 0.1) automatically generates a four-quadrant priority matrix, classifying issues into four levels: urgent and important (to be addressed within 48 hours), important but not urgent (included in iterations), urgent but not important (standardized processing), and routine maintenance, providing a quantitative basis for governance resource allocation.

[0043] 2. Dual-engine driven rule management module This invention achieves parallel rule management through a dual-engine driven rule system, such as... Figure 4 As shown, the specific implementation method is as follows: First, there's the AI ​​engine. This engine is built upon deep learning to perform natural language parsing. Through pre-trained power industry-specific models, it accurately identifies key elements and logical relationships or unstructured data standards in user-input business requirements. This data is automatically converted into structured rule expressions and generates executable SQL / Python code frameworks. Simultaneously, leveraging the semantic association capabilities of a domain knowledge graph, it recommends existing detection rules to assist in manual verification. Once confirmed, these rules are stored in a unified rule base. This process effectively solves the inefficiencies and logical blind spots inherent in traditional manual rule writing, enabling rapid and accurate transformation of business requirements into technical rules.

[0044] Secondly, the low-code engine allows business users to configure rules independently through a visual drag-and-drop interface. Users can select data fields, operators (such as > and contain), and thresholds to intuitively build complete, unique, valid, accurate, and consistent rules (such as customer age ≥ 18). The system has a built-in intelligent validation mechanism that automatically performs multiple checks before saving: including field type and operator matching checks (such as disabling numeric comparisons for date fields) and threshold validity verification (such as thresholds containing both numeric and character types), and provides correction suggestions (such as selecting date-type operators). This real-time feedback improves the efficiency and accuracy of rule configuration, truly achieving zero-coding business rule management.

[0045] Third, real-time monitoring of rule health is implemented. Logical reasoning algorithms identify conflicts between rules (e.g., rule A requires non-empty rules, while rule B allows empty rules). This, combined with data lineage tracing to track the impact of upstream metadata changes on rules (e.g., changes to source system fields), automatically triggers alerts and generates optimization suggestions. Simultaneously, the effectiveness of rules is evaluated based on historical execution data, intelligently recommending redundant rule merging and performance optimization solutions. This ensures the rule base remains healthy at all times, preventing data quality incidents caused by rule failures or conflicts, and achieving continuous self-optimization capabilities for data quality management.

[0046] 3. Quality verification module based on dynamic baseline engine and intelligent task scheduling This invention performs quality checks based on a dynamic baseline engine and intelligent task scheduling, such as... Figure 5 As shown, the specific implementation is as follows: First, a dynamic baseline engine is built to generate thresholds, enabling intelligent detection and predictive alerts for data quality anomalies. Using time series forecasting algorithms (such as ARIMA, LSTM, or Prophet), the system automatically learns fluctuation patterns based on historical data quality indicators (such as data integrity, accuracy, and timeliness) to establish a dynamic pass rate baseline. The system automatically trains the model daily, generating dynamic thresholds (such as ±3σ ranges). When real-time data quality indicators exceed the threshold range, an alarm is triggered. Simultaneously, the dynamic thresholds support manual calibration and intervention to ensure business adaptability. For example, if the pass rate of line current data in a certain province drops sharply by 30% within 24 hours, the system automatically marks it as an anomaly and pushes it to the verification work order system.

[0047] Second, distributed task scheduling and intelligent rule grouping. Through the distributed task scheduling framework, large-scale verification rules are optimized according to the following strategies: (1) Intelligent rule grouping: Based on data lineage, business scenario or execution frequency, rules with strong correlation are packaged (such as merging the compliance verification of fields in the same data table into one task) to reduce redundant scanning. (2) Dynamic resource allocation: Based on rule priority (such as core business data rules taking priority) and computational complexity, cluster resources are dynamically allocated to ensure that high-priority tasks are completed quickly. (3) Incremental verification technology: Through CDC tools or incremental snapshots or log comparisons based on data update timestamps, only data newly added or modified on the same day (such as incremental equipment ledger records) are scanned to avoid full table scanning and reduce verification time.

[0048] 4. Collaborative Autonomous Module Based on Value Community and Intelligent Diagnosis This invention promotes collaborative and self-consistent data governance based on a value community and intelligent diagnostics, such as... Figure 6 As shown, the specific implementation is as follows: First, intelligent diagnosis and autonomous closed-loop processing enable automatic root cause analysis and closed-loop handling of anomalies. When the system detects an anomaly and receives an alarm work order, it automatically associates the full-link context information of the anomaly data and performs in-depth root cause analysis and closed-loop processing. Specifically, this includes: Data link tracing: checking recent version changes of the system to which the anomaly data belongs, ETL task logs (e.g., whether task failures or delays occurred), and source-end operation records (e.g., manual data entry or interface changes) to quickly locate faults caused by changes or operations. Lineage correlation analysis: using the lineage relationship network in the AI-driven demand collaborative decision-making module, locating upstream and downstream dependency issues (e.g., anomalies in certain archive data leading to incorrect calculations of measurement data indicators) to accurately diagnose the root cause of the problem. Automated report output and closed-loop handling: integrating the analysis results to generate a structured report (including anomaly timeline, impact scope, possible causes, and remediation suggestions). The system directly pushes this report to the relevant responsible persons and can trigger automatic repair scripts for clear and programmable issues (e.g., data re-collection, task rerun), completing an autonomous closed loop from "discovery-diagnosis-handling-verification".

[0049] Second, establish a private domain quality profiling and target collaboration mechanism to drive proactive business participation. Moving beyond traditional global reporting models, this involves building a perceptible and manageable governance space for each business unit, transforming technical governance capabilities into business empowerment tools. (1) Data quality profiling: A comprehensive quality evaluation system is constructed for each core table in the resource catalog, and the "6-dimensional health index" is continuously integrated and visualized: completeness, accuracy, timeliness, consistency, uniqueness, and effectiveness. The health status and weaknesses of data assets are intuitively displayed through a topology diagram.

[0050] (2) Private Domain Autonomous Portal: Open up exclusive private governance portals to each business department and subordinate unit, with dashboards focusing on displaying the health of data assets directly related to their KPIs. Support business leaders of each unit to customize governance goals (e.g., to support precision marketing, the accuracy rate of the 'customer profile table' tags is required to be increased to 95% next quarter), realizing the transformation from "I have to do it" to "I want to do it".

[0051] (3) Governance Effectiveness-Business Value Mapping Model: Combining the economic impact assessment and demand report of the AI-driven demand collaborative decision-making module, a governance effectiveness-business value mapping model is established. The system regularly generates quarterly value white papers to quantify the business benefits brought about by improved data quality (such as cost savings and efficiency improvements). This surpasses the traditional ranking and reporting model, builds a governance value transformation engine, and drives business departments to actively participate in governance.

[0052] Example 3: Based on the same inventive concept, this invention also provides a closed-loop data quality management system driven by a fourth-order intelligent engine, such as... Figure 7As shown, it includes: The clustering module is used to intelligently cluster the requirements based on the acquired business data governance needs using natural language processing technology, and generate a governance priority decision matrix by combining the full-link impact scope and economic impact assessment. The full-link impact scope is determined based on the data to be governed through metadata lineage graph. The building module is used to construct quality verification rules based on the acquired business data governance needs through a dual-engine driven mode that combines artificial intelligence generation and low-code orchestration. The verification module is used to aggregate and package quality verification rules into rule verification task groups based on the governance priority decision matrix and data lineage correlation, and to perform verification tasks on incremental data in each rule verification task group using a distributed task scheduling framework based on a dynamic baseline threshold. The dynamic baseline threshold is generated by acquiring historical data quality fluctuation patterns. The governance module is used to perform multidimensional root cause diagnosis and automated governance of the abnormal data identified based on the full-link impact scope, and to generate a value assessment report based on the mapping relationship between governance effectiveness and business value established by the governed business data.

[0053] Preferably, the clustering module is specifically used for: Based on the constructed natural language processing requirement parsing engine, a hybrid model of BERT and bidirectional long short-term memory network is used to identify and extract semantic vectors from the acquired business data governance requirements; The unstructured requirements in the semantic vectors are converted into high-dimensional semantic vectors, and the similarity between the vectors is calculated. When the similarity exceeds a preset threshold, the vectors are merged to generate a clustering analysis report. Construct the upstream and downstream impact chain of data assets based on metadata lineage graphs to determine the full-link impact scope of data to be governed; A tiered impact indicator system is constructed, which includes direct economic losses, regulatory risks, and business downtime costs. Based on the tiered impact indicator system, a weighted scoring algorithm is used to calculate the score of each business data governance requirement. Based on the scoring results, the business data governance requirements are classified into levels to generate multi-level importance requirements. A priority decision matrix is ​​generated based on the clustering analysis report, the full-link impact scope, and the multi-level importance requirements.

[0054] Preferably, the building module is specifically used for: Based on the complex logical requirements in the acquired business data governance needs, the natural language descriptions are parsed by the deep learning model in the artificial intelligence generation engine and automatically converted into a preset code framework, and the association rules recommended based on the knowledge graph are used to assist in the verification. Based on the general logical requirements in the acquired business data governance needs, the system generates rule parameters in response to user configuration operations through a visual drag-and-drop interface, and automatically performs field type matching degree detection and threshold validity verification during the configuration process. The preset code framework and the rule parameters are used to construct quality verification rules.

[0055] Preferably, the verification module is specifically used for: A time series prediction algorithm is used to learn the fluctuation patterns of historical data quality indicators, establish a dynamic pass rate baseline, and automatically train a model to generate a dynamic threshold range corresponding to the quality verification rules. Based on the priorities in the governance priority decision matrix, combined with data lineage and business scenarios, multiple quality verification rules with strong correlations are grouped to generate multiple rule verification task groups. Using a distributed task scheduling framework, computing resources are dynamically allocated to each rule verification task group according to rule priority and computational complexity. Incremental data is extracted through preset extraction technology, and verification tasks are executed based on the real-time quality indicators of the incremental data and the dynamic threshold range.

[0056] Preferably, the governance module performs multidimensional root cause diagnosis and automated governance of the identified abnormal data based on the full-link impact range, including: When an abnormal alarm is detected, the full-link context information of the abnormal data is automatically associated within the full-link impact range. The context information includes checking recent version changes of the system to which the data belongs, ETL task logs, and source operation records. Based on the aforementioned metadata lineage map, upstream and downstream dependency issues are located within the entire influence chain, and a root cause analysis report is generated. The root cause analysis report is parsed, and the pre-set programmable problem library is matched according to the parsed fault type to trigger an automatic repair script.

[0057] Preferably, the system further includes a drive module for: A six-dimensional quality profile is constructed for each business data item, and the profile is visualized using a topology diagram. Open the private governance portal to business departments and subordinate units to customize governance goals and indicators; Based on the obtained economic impact assessment results and demand analysis report, a mapping model between governance effectiveness and business value is established, and a quarterly value white paper quantifying business revenue is generated based on the mapping model.

[0058] Preferably, the system further includes an early warning module, used for: The system scans the rule base corresponding to the quality verification rules in real time and uses logical reasoning algorithms to identify conflicting rules. Based on data lineage tracking and source metadata changes, when a field in the source system changes, an automatic failure rule warning is triggered and optimization suggestions are generated.

[0059] Example 4 like Figure 8 As shown, the present invention also provides an electronic device, which may be a computer device, a microcontroller device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, processor, and transceiver component are connected via a bus; the memory can be used to store executable programs, and an exemplary executable program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, which can be accessed and / or modified when instructions are executed.

[0060] The processor may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, and it is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to realize the corresponding method flow or corresponding function, so as to realize the steps of a closed-loop data quality management method based on a fourth-order intelligent engine driven in the above embodiments.

[0061] Example 5 Based on the same inventive concept, this invention also provides a readable storage medium, specifically an electronic device readable storage medium (Memory). This readable storage medium is a memory device within an electronic device used to store programs and data. It is understood that the storage medium here can include both built-in storage media within the electronic device and extended storage media supported by the electronic device. The storage medium provides storage space, which stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more executable programs (including program code). It should be noted that the storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. Loading and executing one or more instructions stored in the storage medium by the processor can implement the steps of a closed-loop data quality management method based on a fourth-order intelligent engine driven in the above embodiments.

[0062] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0063] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0064] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.

[0065] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0066] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of the claims of the present invention pending approval.

Claims

1. A closed-loop data quality management method based on a fourth-order intelligent engine, characterized in that, include: Based on the acquired business data governance needs, natural language processing technology is used to intelligently cluster the needs, and a governance priority decision matrix is ​​generated by combining the full-link impact scope and economic impact assessment. The full-link impact scope is determined based on the data to be governed through metadata lineage graph. Based on the acquired business data governance needs, quality verification rules are constructed using a dual-engine driven model that combines artificial intelligence generation and low-code orchestration. Based on the governance priority decision matrix and data lineage correlation, the quality verification rules are aggregated and packaged into rule verification task groups. Based on the dynamic baseline threshold, a distributed task scheduling framework is used to execute verification tasks on the incremental data in each rule verification task group. The dynamic baseline threshold is generated by obtaining the historical data quality fluctuation pattern. Based on the full-link impact scope, the abnormal data identified is subjected to multidimensional root cause diagnosis and automated governance, and a value assessment report is generated based on the governance effectiveness and business value mapping relationship established on the governed business data.

2. The method according to claim 1, characterized in that, The governance requirements based on acquired business data utilize natural language processing technology for intelligent clustering of requirements, and combine this with end-to-end impact assessment and economic impact evaluation to generate a governance priority decision matrix, including: Based on the constructed natural language processing requirement parsing engine, a hybrid model of BERT and bidirectional long short-term memory network is used to identify and extract semantic vectors from the acquired business data governance requirements; The unstructured requirements in the semantic vectors are converted into high-dimensional semantic vectors, and the similarity between the vectors is calculated. When the similarity exceeds a preset threshold, the vectors are merged to generate a clustering analysis report. Construct the upstream and downstream impact chain of data assets based on metadata lineage graphs to determine the full-link impact scope of data to be governed; A tiered impact indicator system is constructed, which includes direct economic losses, regulatory risks, and business downtime costs. Based on the tiered impact indicator system, a weighted scoring algorithm is used to calculate the score of each business data governance requirement. Based on the scoring results, the business data governance requirements are classified into levels to generate multi-level importance requirements. A priority decision matrix is ​​generated based on the clustering analysis report, the full-link impact scope, and the multi-level importance requirements.

3. The method according to claim 1, characterized in that, The quality verification rules are constructed using a dual-engine driven model combining artificial intelligence generation and low-code orchestration, based on the acquired business data governance requirements. This includes: Based on the complex logical requirements in the acquired business data governance needs, the natural language descriptions are parsed by the deep learning model in the artificial intelligence generation engine and automatically converted into a preset code framework, and the association rules recommended based on the knowledge graph are used to assist in the verification. Based on the general logical requirements in the acquired business data governance needs, the system generates rule parameters in response to user configuration operations through a visual drag-and-drop interface, and automatically performs field type matching degree detection and threshold validity verification during the configuration process. The preset code framework and the rule parameters are used to construct quality verification rules.

4. The method according to claim 1, characterized in that, The process of aggregating and packaging quality verification rules into rule verification task groups based on the governance priority decision matrix and data lineage correlation, and executing verification tasks on incremental data in each rule verification task group using a distributed task scheduling framework based on a dynamic baseline threshold, includes: A time series prediction algorithm is used to learn the fluctuation patterns of historical data quality indicators, establish a dynamic pass rate baseline, and automatically train a model to generate a dynamic threshold range corresponding to the quality verification rules. Based on the priorities in the governance priority decision matrix, combined with data lineage and business scenarios, multiple quality verification rules with strong correlations are grouped to generate multiple rule verification task groups. Using a distributed task scheduling framework, computing resources are dynamically allocated to each rule verification task group according to rule priority and computational complexity. Incremental data is extracted through preset extraction technology, and verification tasks are executed based on the real-time quality indicators of the incremental data and the dynamic threshold range.

5. The method according to claim 1, characterized in that, The process of performing multidimensional root cause diagnosis and automated management of the abnormal data identified based on the full-link impact range includes: When an abnormal alarm is detected, the full-link context information of the abnormal data is automatically associated within the full-link impact range. The context information includes checking recent version changes of the system to which the data belongs, ETL task logs, and source operation records. Based on the aforementioned metadata lineage map, upstream and downstream dependency issues are located within the entire influence chain, and a root cause analysis report is generated. The root cause analysis report is parsed, and the pre-set programmable problem library is matched according to the parsed fault type to trigger an automatic repair script.

6. The method according to claim 1, characterized in that, After generating a value assessment report by establishing a mapping relationship between governance effectiveness and business value based on the governed business data, the process also includes: A six-dimensional quality profile is constructed for each business data item, and the profile is visualized using a topology diagram. Open the private governance portal to business departments and subordinate units to customize governance goals and indicators; Based on the obtained economic impact assessment results and demand analysis report, a mapping model between governance effectiveness and business value is established, and a quarterly value white paper quantifying business revenue is generated based on the mapping model.

7. The method according to claim 1, characterized in that, After constructing quality verification rules based on the acquired business data governance requirements using a dual-engine driven model combining artificial intelligence generation and low-code orchestration, the process also includes: The system scans the rule base corresponding to the quality verification rules in real time and uses logical reasoning algorithms to identify conflicting rules. Based on data lineage tracking and source metadata changes, when a field in the source system changes, an automatic failure rule warning is triggered and optimization suggestions are generated.

8. A closed-loop data quality management system driven by a fourth-order intelligent engine, characterized in that, include: The clustering module is used to intelligently cluster the requirements based on the acquired business data governance needs using natural language processing technology, and generate a governance priority decision matrix by combining the full-link impact scope and economic impact assessment. The full-link impact scope is determined based on the data to be governed through metadata lineage graph. The building module is used to construct quality verification rules based on the acquired business data governance needs through a dual-engine driven mode that combines artificial intelligence generation and low-code orchestration. The verification module is used to aggregate and package quality verification rules into rule verification task groups based on the governance priority decision matrix and data lineage correlation, and to perform verification tasks on incremental data in each rule verification task group using a distributed task scheduling framework based on a dynamic baseline threshold. The dynamic baseline threshold is generated by obtaining historical data quality fluctuation patterns. The governance module is used to perform multidimensional root cause diagnosis and automated governance of the abnormal data identified based on the full-link impact scope, and to generate a value assessment report based on the mapping relationship between governance effectiveness and business value established by the governed business data.

9. An electronic device, characterized in that, include: At least one processor and memory; The memory and processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the closed-loop data quality management method based on a fourth-order intelligent engine as described in any one of claims 1 to 7 is implemented.

10. A readable storage medium, characterized in that, It contains an execution program, which, when executed, implements the closed-loop data quality management method based on a fourth-order intelligent engine as described in any one of claims 1 to 7.