Data quality improvement method and system based on multi-dimensional governance strategy
Through multi-dimensional governance strategies, including dynamic primary key recognition, multi-mode completion and full-link traceability management, the problems of duplicate data recognition, missing value completion and wrong data correction in the existing technology are solved, and quality optimization and efficient management of the entire life cycle of data are achieved.
Patent Information
- Application Number
- CN202510431574.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-29
AI Technical Summary
In the existing data governance technology, duplicate data recognition relies on a single primary key to process complex related data, missing value filling methods are single, and completion strategies that do not distinguish between different scenarios, error data correction lacks a hierarchical processing mechanism, governance process lacks traceability management, and data change history cannot be traced.
A multi-dimensional governance strategy is adopted to construct a dynamic primary key duplicate data identification and merging mechanism through metadata blood analysis. A multi-modal incomplete data completion strategy based on field importance, combined with a three-level error processing process of automatic correction, rule warning and manual intervention to realize full-link governance traceability management.
Significantly improve data availability, reduce manual intervention costs, and is suitable for automated data cleaning and quality management in large-scale heterogeneous data environments.
Smart Images

Figure CN120386779A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data governance, and specifically to a method and system for improving data quality based on multi-dimensional governance strategies. Background Art
[0002] The existing data governance technologies have the following problems:
[0003] 1. Duplicate data identification depends on a single primary key and cannot handle complex associated data;
[0004] 2. The method for filling missing values is single, and the filling strategies for different scenarios are not distinguished;
[0005] 3. The correction of incorrect data lacks a hierarchical processing mechanism, and important data is easily deleted by mistake;
[0006] 4. The governance process lacks traceability management and cannot trace the data change history. Summary of the Invention
[0007] The technical task of the present invention is to address the above deficiencies, and provide a method and system for improving data quality based on multi-dimensional governance strategies, which can significantly improve data availability, reduce the cost of manual intervention, and are applicable to automated data cleaning and quality management in large-scale heterogeneous data environments.
[0008] The technical solution adopted by the present invention to solve its technical problems is:
[0009] A method for improving data quality based on multi-dimensional governance strategies, including:
[0010] Constructing a duplicate data identification and merging mechanism with a dynamic primary key through metadata lineage analysis;
[0011] A multi-mode incomplete data filling strategy based on field importance, and realizing the classification filling of incomplete data by combining logical reasoning and statistical methods;
[0012] Including a three-level error handling process of automatic correction, rule warning, and manual intervention to achieve forced compliance and intelligent correction;
[0013] Full-link governance traceability based on blockchain technology to achieve full-link traceability management of the data governance process.
[0014] This method realizes the quality optimization of the entire data life cycle through a multi-dimensional collaborative governance strategy of noise data deletion, duplicate data merging, incomplete data filling, incorrect data correction, problem data archiving, and governance traceability management.
[0015] Further, the multi-mode incomplete data filling includes:
[0016] Filling through field logical relationship calculation;
[0017] Complete using a time series prediction model;
[0018] Verify and complete by associating with external data sources.
[0019] Furthermore, the triggering conditions of the three-level error handling process include:
[0020] Automatic correction: field format errors, simple logic conflicts;
[0021] Rule warning: numerical out-of-bounds but can be automatically corrected;
[0022] Manual intervention: key business field anomalies and cannot be automatically repaired.
[0023] Furthermore, the implementation of this method includes the following steps:
[0024] 1) Noise data governance, including establishing a rule engine library and a dynamic matching processor;
[0025] 2) Duplicate data governance, including:
[0026] Lineage analyzer: Trace the data generation link through metadata;
[0027] Composite primary key generator: Generate a dynamic primary key by combining business timestamps, device IDs, etc.;
[0028] Intelligent merge decision tree: Select a retention strategy based on data timeliness and source credibility;
[0029] 3) Incomplete data governance, including field importance assessment and multi-mode filling;
[0030] 4) Error data governance, including a three-level processing mechanism:
[0031] Level 1: Automatic correction, including: format conversion, outlier replacement;
[0032] Level 2: Rule warning, including: triggering the business review process;
[0033] Level 3: Manual intervention, including: generating a data correction work order;
[0034] 5) Governance traceability.
[0035] Furthermore, the establishment of the rule engine library: stores constraint rules such as field formats and value ranges;
[0036] The dynamic matching processor: real-time detects and deletes data that does not conform to the rules.
[0037] Furthermore, the field importance assessment is based on the field usage frequency and business impact score to establish a field importance assessment matrix;
[0038] The multi-mode filling includes:
[0039] Logical association filling: generating missing values through the calculation relationship between fields;
[0040] Time series interpolation filling: using the ARIMA model to predict missing time series;
[0041] Multi-source verification filling: associating with external data sources for cross-verification.
[0042] Furthermore, the governance traceability includes:
[0043] Traceability blockchain: recording the data cleaning operation logs, including timestamps, operators, and modified content;
[0044] Version control system: saving the states before and after data modification and supporting version backtracking.
[0045] The present invention also claims a data quality improvement system based on a multi-dimensional governance strategy, including:
[0046] Noise data governance module, including a rule engine library and a dynamic matching processor;
[0047] Duplicate data governance module, including a lineage analyzer, a composite primary key generator, and an intelligent merge decision tree;
[0048] Incomplete data governance module, including field importance evaluation and multi-mode filling, for implementing multi-mode incomplete data completion;
[0049] Error data governance module, including a three-level processing mechanism;
[0050] Governance traceability module, for full-link governance traceability based on blockchain technology;
[0051] This system can implement the above-mentioned data quality improvement method based on a multi-dimensional strategy.
[0052] The present invention also claims a data quality improvement device based on a multi-dimensional governance strategy, including: at least one memory and at least one processor;
[0053] The at least one memory is used to store machine-readable programs;
[0054] The at least one processor is used to call the machine-readable programs to implement the above-mentioned method.
[0055] The present invention also claims a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the above-mentioned method is implemented.
[0056] Compared with the prior art, a data quality improvement method and system based on multi-dimensional governance strategies of the present invention have the following beneficial effects:
[0057] This method constructs a multi-dimensional data quality evaluation and governance system to achieve intelligent duplicate identification and merging based on data lineage; develops classification and completion strategies for different missing scenarios, establishes a hierarchical correction and compliance verification mechanism for incorrect data, and creates a full-process governance traceability management system; through multi-dimensional collaborative governance strategies such as noise data deletion, duplicate data merging, incomplete data completion, incorrect data correction, problem data archiving, and governance traceability management, it realizes the quality optimization of the entire data life cycle. The present invention can significantly improve data availability and reduce the cost of manual intervention, and is applicable to data governance scenarios in fields such as finance, healthcare, and the Internet of Things. Description of the Drawings
[0058] Figure 1 is a flowchart of the data quality improvement method based on multi-dimensional strategies provided by an embodiment of the present invention. Detailed Embodiments
[0059] The present invention will be further described below in conjunction with the drawings and specific embodiments.
[0060] An embodiment of the present invention provides a data quality improvement method based on multi-dimensional governance strategies, including:
[0061] Constructing a duplicate data identification and merging mechanism with dynamic primary keys through metadata lineage analysis;
[0062] A multi-mode incomplete data completion strategy based on field importance, combining logical reasoning and statistical methods to achieve classification and completion of incomplete data;
[0063] A three-level error handling process including automatic correction, rule warning, and manual intervention to achieve forced compliance and intelligent correction;
[0064] Full-link governance traceability based on blockchain technology to achieve full-link traceability management of the data governance process.
[0065] Among them, the multi-mode incomplete data completion includes:
[0066] Completing through field logical relationship calculation;
[0067] Completing using a time series prediction model;
[0068] Verifying and completing by associating with external data sources.
[0069] The triggering conditions of the three-level error handling process include:
[0070] Automatic correction: field format error, simple logical conflict;
[0071] Rule warning: The value is out of bounds but can be automatically corrected;
[0072] Manual intervention: Key business fields are abnormal and cannot be automatically repaired.
[0073] This method constructs a multi-dimensional data quality evaluation and governance system, realizes intelligent duplicate identification and merging based on data lineage, develops classification completion strategies for different missing scenarios, establishes a hierarchical correction and compliance verification mechanism for error data, and creates a whole-process governance traceability management system; through multi-dimensional collaborative governance strategies such as noise data deletion, duplicate data merging, incomplete data completion, error data correction, problem data archiving, and governance traceability management, the quality optimization of the entire data life cycle is achieved.
[0074] The specific implementation process of this method is as follows:
[0075] 1. Noise data governance, including establishing a rule engine library and a dynamic matching processor.
[0076] Establish a rule engine library: Store constraint rules such as field formats and value ranges;
[0077] Dynamic matching processor: Detect and delete data that does not conform to the rules in real time.
[0078] 2. Duplicate data governance, including:
[0079] Lineage relationship analyzer: Trace the data generation link through metadata;
[0080] Composite primary key generator: Generate a dynamic primary key by combining business timestamps, device IDs, etc.;
[0081] Intelligent merge decision tree: Select a retention strategy based on data timeliness and source credibility.
[0082] 3. Incomplete data governance, including field importance assessment and multi-mode filling.
[0083] Field importance assessment: Establish a field importance assessment matrix based on field usage frequency and business impact score.
[0084] Multi-mode filling, including:
[0085] Logical association filling: Generate missing values through the calculation relationship between fields;
[0086] Time series interpolation filling: Use the ARIMA model to predict time series missing values;
[0087] Multi-source verification filling: Associate external data sources for cross-verification.
[0088] 4. Error data governance, including a three-level processing mechanism:
[0089] Level 1: Automatic correction, including: format conversion and outlier replacement.
[0090] Level 2: Rule warning, including: triggering the business review process.
[0091] Level 3: Manual intervention, including: generating data correction work orders.
[0092] 5. Governance traceability, including:
[0093] Traceability blockchain: Records the data cleaning operation logs (timestamp, operator, modified content).
[0094] Version control system: Saves the states before and after data modification and supports version backtracking.
[0095] The following effect comparison was obtained through test verification of this method (test environment: 1 million medical data):
[0096] Index Traditional method This method Data availability 72% 95% Processing efficiency 1200 pieces / second 5800 pieces / second Proportion of manual intervention 35% 8% Traceability query efficiency None <200ms
[0097] This method significantly improves data availability and reduces the cost of manual intervention, and is applicable to data governance scenarios in fields such as finance, healthcare, and the Internet of Things.
[0098] The embodiment of the present invention also provides a data quality improvement system based on a multi-dimensional governance strategy, including:
[0099] 1. Noise data governance module, including a rule engine library and a dynamic matching processor.
[0100] Establish a rule engine library: Store constraint rules such as field formats and value ranges;
[0101] Dynamic matching processor: Real-time detects and deletes data that does not conform to the rules.
[0102] 2. Duplicate data governance module, including a lineage analyzer, a composite primary key generator, and an intelligent merge decision tree.
[0103] Lineage analyzer: Traces the data generation link through metadata;
[0104] Composite primary key generator: Generates a dynamic primary key by combining the business timestamp, device ID, etc.;
[0105] Intelligent merge decision tree: Selects a retention strategy based on data timeliness and source credibility.
[0106] 3. Incomplete data governance module, including field importance assessment and multi-mode filling, for implementing multi-mode incomplete data completion.
[0107] Field importance assessment: Builds a field importance assessment matrix based on field usage frequency and business impact score.
[0108] Multi - mode filling, including:
[0109] Logical association filling: generating missing values through the calculation relationship between fields;
[0110] Time - series interpolation filling: using the ARIMA model to predict missing values in time series;
[0111] Multi - source verification filling: associating external data sources for cross - verification.
[0112] 4. Error data governance module, including a three - level processing mechanism:
[0113] Level 1: Automatic correction, including: format conversion, outlier replacement.
[0114] Level 2: Rule warning, including: triggering the business review process.
[0115] Level 3: Manual intervention, including: generating a data correction work order.
[0116] 5. Governance traceability module, used for full - link governance traceability based on blockchain technology; including:
[0117] Traceability blockchain: recording data cleaning operation logs (timestamp, operator, modified content).
[0118] Version control system: saving the pre - and post - modification states of data, supporting version backtracking.
[0119] This system, through the intelligent identification and merging mechanism of duplicate data based on metadata lineage relationship; combined with the strategy of classifying and complementing incomplete data using logical reasoning and statistical methods; including a hierarchical error handling system with mandatory compliance and intelligent correction; and full - link traceability management of the governance process; can implement the data quality improvement method based on multi - dimensional governance strategies described in the above embodiments.
[0120] An embodiment of the present invention also provides a data quality improvement device based on multi - dimensional governance strategies, including: at least one memory and at least one processor;
[0121] The at least one memory is used to store machine - readable programs;
[0122] The at least one processor is used to call the machine - readable program to implement the data quality improvement method based on multi - dimensional governance strategies described in the above embodiments.
[0123] An embodiment of the present invention also provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor is caused to execute the data quality improvement method based on a multi-dimensional governance strategy in the above embodiment. Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program code stored in the storage medium.
[0124] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0125] Examples of the storage medium for providing the program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0126] In addition, it should be clear that not only can the functions of any one of the above embodiments be realized by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.
[0127] In addition, it can be understood that the program code read from the storage medium is written into a memory provided in an expansion board inserted into the computer or a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is caused to execute part and all of the actual operations, thereby realizing the functions of any one of the above embodiments.
[0128] The present invention has been described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that the code review means in different above embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.
Claims
1. A method for improving data quality based on multi-dimensional governance strategies, characterized in that, Including: A duplicate data identification mechanism that constructs a dynamic primary key through metadata lineage analysis; A multi-mode incomplete data completion strategy based on field importance; A three-level error handling process including automatic correction, rule warning, and manual intervention; Full-link governance traceability based on blockchain technology.
2. The data quality improvement method based on multi-dimensional governance strategies according to claim 1, wherein, The multi-mode incomplete data completion includes: Completion through calculation of field logical relationships; Completion using a time series prediction model; Completion by verifying and associating with external data sources.
3. A method for improving data quality based on a multi-dimensional governance strategy according to claim 1, characterized in that The triggering conditions of the three-level error handling process include: Automatic correction: field format errors, simple logical conflicts; Rule warning: numerical out-of-bounds but can be automatically corrected; Manual intervention: key business field anomalies and cannot be automatically repaired.
4. A method for improving data quality based on a multi-dimensional governance strategy according to claim 1 or 2 or 3, characterized in that The implementation of this method includes the following steps: 1) Noise data governance, including establishing a rule engine library and a dynamic matching processor; 2) Duplicate data governance, including: Lineage analyzer: generating a data generation link through metadata tracing; Composite primary key generator: generating a dynamic primary key by combining business timestamps, device IDs, etc.; Intelligent merge decision tree: selecting a retention strategy based on data timeliness and source credibility; 3) Incomplete data governance, including field importance assessment and multi-mode filling; 4) Error data governance, including a three-level processing mechanism: Level 1: Automatic correction, including: format conversion, outlier replacement; Level 2: Rule warning, including: triggering a business review process; Level 3: Manual intervention, including: generating a data correction work order; 5) Governance traceability.
5. A method for improving data quality based on a multi-dimensional governance strategy according to claim 4, characterized in that, The establishment of the rule engine library: storing constraint rules such as field formats and value range; The dynamic matching processor: detecting and deleting data that does not conform to the rules in real time.
6. The data quality improvement method based on a multi-dimensional governance strategy according to claim 4, characterized in that The field importance assessment is based on the field usage frequency and business impact score to establish a field importance assessment matrix; The multi-mode filling includes: Logical association filling: generating missing values through the calculation relationship between fields; Time series interpolation filling: predicting time series missing values using the ARIMA model; Multi-source verification filling: associating with external data sources for cross-verification.
7. A method for improving data quality based on a multi-dimensional governance strategy according to claim 4, characterized in that The governance traceability includes: Traceability blockchain: recording data cleaning operation logs, including timestamps, operators, and modified content; Version control system: saving the state before / after data modification and supporting version backtracking.
8. A data quality improvement system based on a multi-dimensional governance strategy, characterized in that, Including: Noise data governance module, including establishing a rule engine library and a dynamic matching processor; Duplicate data governance module, including a lineage analyzer, a composite primary key generator, and an intelligent merge decision tree; Incomplete data governance module, including field importance assessment and multi-mode filling, used to implement multi-mode incomplete data completion; Error data governance module, including a three-level processing mechanism; Governance traceability module, used for full-link governance traceability based on blockchain technology; This system can implement the data quality improvement method based on multi-dimensional strategies described in any one of claims 1 to 7.
9. A data quality improvement device based on a multi-dimensional governance strategy, characterized in that, Including: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.
10. A computer-readable medium, characterized in that, Computer instructions are stored on the computer-readable medium, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.