Data quality detection and automatic correction device and method
Through data quality detection and automatic correction devices, data quality problems are automatically processed, and a variety of algorithms and scheduling strategies are used to solve the problem of automatic data quality correction and improve data processing and analysis efficiency.
Patent Information
- Application Number
- CN202510179106.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, data quality problems are difficult to automatically correct, resulting in manual corrections being time-consuming and labor-intensive and error-prone, and data analysis results are inaccurate.
Provides a device and method for data quality detection and automatic correction, including input module, configuration module, scheduling module, execution module, audit module and output module. Data quality problems are discovered and corrected through automated processes, supports a variety of data sources and types, uses SQL, programming languages and artificial intelligence algorithms to check and correct, and optimize task execution in combination with scheduling strategies.
It realizes automatic discovery and correction of data quality problems, reduces manual intervention, improves data processing and analysis efficiency, and has good scalability and accuracy.
Smart Images

Figure CN120256415A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data quality detection, and more specifically, to an apparatus and method for data quality detection and automatic correction. Background Art
[0002] Due to various errors that may occur during data acquisition, transmission, processing, storage, etc., data quality problems occur frequently. Data quality problems include, but are not limited to, consistency, uniqueness, integrity, accuracy, timeliness, relevance, etc. Data quality problems may lead to inaccurate and unreliable data analysis results, and even cause significant losses.
[0003] To address these problems, there are currently some methods for evaluating and studying data quality. However, most of these methods focus on how to evaluate the quality of a data set, and do not involve the problem of correcting incorrect data. In practical applications, manually correcting data quality problems is often time-consuming and laborious, and is prone to errors. Summary of the Invention
[0004] The present invention aims to solve the drawbacks in the background art, and provides an apparatus and method for data quality detection and automatic correction, which can achieve automatic correction of data on the premise of detecting data quality problems, is beneficial to improving the efficiency of data quality detection and correction; and can automatically schedule the work tasks of data quality detection and correction to quickly process the quality problems of a large amount of data and improve the efficiency of data processing and analysis.
[0005] To achieve the above object, the present invention provides an apparatus for data quality detection and automatic correction, including: An input module: for inputting a data set to be detected; A configuration module: for configuring a plurality of work tasks for the data set to be detected, each work task being used to process a type of data quality problem, and the configuration content of each work task including fields to be checked, check rules, correction methods, priorities; A scheduling module: for setting a scheduling method and scheduling the configured work tasks according to the scheduling method; An execution module: for executing the work tasks according to the scheduling result of the scheduling module, and checking and correcting the data set to be detected according to the configuration content of the work tasks to obtain a result data set, and generating a corresponding work report after the work tasks are completed; An auditing module: for receiving the result data set and the work report, and auditing and correcting the result data set and the work report; An output module: for outputting the result data set and the work report after the auditing and correction by the auditing module are completed.
[0006] Further, the configuration module includes: When configuring each of the work tasks, several of the fields to be inspected are selected based on the understanding of the dataset to be detected and the business and uses to which the dataset to be detected belongs, and from the perspectives of importance, severity of data quality problems, and field relevance.
[0007] Further, the configuration module includes: When configuring each of the work tasks, the preset inspection rules are selected or new inspection rules are created according to the specifications of the fields to be inspected.
[0008] Further, the configuration module includes: When configuring each of the work tasks, according to the data standards and specifications of the fields to be inspected, the preset correction methods are selected for the data in the fields to be inspected with problems, or new correction methods are created.
[0009] Further, the correction methods specifically include: For problems with inconsistent data formats, SQL or other programming languages are used for format conversion; For uniqueness problems, a duplicate removal algorithm is used to delete duplicate data; For missing value problems, standard values are queried from the basic database based on the values of relevant fields for filling, or artificial intelligence algorithms are used to generate data for filling; For accuracy problems, artificial intelligence algorithms are used to identify and correct problem data; For timeliness problems, timestamps are used for detection and corresponding processing.
[0010] Further, the scheduling module includes: Different scheduling methods are selected according to the resource status of the system, the available time of data, the urgency of data requirements, and the logical order of data inspection.
[0011] Further, the scheduling methods include the determination method, the priority method, and the workload method; The determination method means that the work tasks are executed according to the specified cycle and time points; The priority method means that the work tasks are executed according to the priorities in the configuration content of the work tasks. For work tasks at the same priority, it is decided whether to execute them simultaneously or randomly select some of the work tasks to execute first according to the available resource status of the system; The workload method means that the workload of the work tasks is evaluated before execution, and the work tasks with low workload are executed first.
[0012] Further, the work report includes inspection results and correction results, and the specific content includes: the forms and fields inspected, the names and meanings of the inspection rules used, the quantity and proportion of unqualified records found, the names and meanings of the correction methods used, and the quantity and proportion of the records corrected.
[0013] Further, the audit module includes auditing the inspection results and the correction results, and includes: For each record with data corrected, determine whether the inspection results and the correction results meet the expected values. If they meet the expected values, accept the corresponding inspection rules and correction methods; otherwise, reject the corresponding inspection rules and correction methods. Statistically analyze all the corrected records related to the same inspection rule. If the acceptance ratio is greater than the preset threshold, accept the inspection rule; otherwise, the inspection rule needs to be revised or a new inspection rule needs to be used, and the work task needs to be reconfigured and scheduled for execution.
[0014] In a second aspect, the present invention also provides a method for data quality detection and automatic correction, including the following steps: Step S1: Select the data set to be detected; Step S2: Configure a number of work tasks for the data set to be detected. Each work task is used to process a type of data quality problem, and the configuration content of each work task includes fields to be inspected, inspection rules, correction methods, and priorities. Step S3: Select a scheduling method and schedule the configured work tasks according to the scheduling method; Step S4: Execute the work tasks according to the scheduling results, and inspect and correct the data set to be detected according to the configuration content of the work tasks to obtain a result data set, and generate a corresponding work report after the work tasks are completed; Step S5: Audit and correct the result data set and the work report; Step S6: After the auditing and correction are completed, output the result data set and the work report.
[0015] Compared with the prior art, the present invention has the following beneficial effects: Each module of the device for data quality detection and automatic correction enables users to automatically detect and correct data quality problems, reduce manual intervention and errors, and can detect various data quality problems existing in the dataset, such as non-standard formats, non-unique data, incomplete data, inaccurate data, etc.; at the same time, the present invention can automatically schedule work tasks, can quickly process the quality problems of a large amount of data, is beneficial to improving the efficiency of data processing and analysis; and, the present invention can be extended to multiple data sources, data types and problems, and through the audit module to assist manual audit, can perform customized inspections and corrections, has good scalability, and can further ensure the quality of the data after detection and correction. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Appendix Figure 1 is a schematic diagram of the device for data quality detection and automatic correction of the present invention.
[0017] Appendix Figure 2 is a flowchart of the method for data quality detection and automatic correction of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Next, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of the present application.
[0019] The following disclosure provides different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or reference letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed. The following provides a detailed description of a device and method for data quality detection and automatic correction provided by the present invention. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0020] Embodiment 1 Please refer to Appendix Figure 1 , this embodiment provides a device for data quality detection and automatic correction. The device can realize the cooperation of an input module, a configuration module, a scheduling module, an execution module, an audit module and an output module to perform quality detection and automatic correction on data, and output a work report and a result dataset after passing the audit.
[0021] A device for data quality detection and automatic correction provided in this embodiment specifically includes: Input module 11: used to input the dataset to be detected; Configuration module 12: used to configure a number of work tasks for the dataset to be detected. Each work task is used to process a type of data quality problem, and the configuration content of each work task includes fields to be checked, check rules, correction methods, and priorities; Scheduling module 13: used to set a scheduling method and schedule the configured work tasks according to the scheduling method; Execution module 14: used to execute work tasks according to the scheduling results of the scheduling module, and check and correct the dataset to be detected according to the configuration content of the work tasks to obtain a result dataset, and generate a corresponding work report after the work tasks are completed; Audit module 15: used to receive the result dataset and the work report, and audit and revise the result dataset and the work report; Output module 16: used to output the result dataset and the work report after the audit and revision by the audit module are completed.
[0022] Specifically, the configuration content of the work task includes fields to be checked, check rules, correction methods, and priorities. One work task is used to process a type of data quality problem. Specifically, the number of work tasks can be configured according to the judgment of business needs and data quality status. If there are multiple problems in the data, multiple work tasks can be configured. Among them, the priority can play a reference role when scheduling the execution order of work tasks later. Specifically, work tasks with higher priorities are executed first.
[0023] Specifically, the fields to be checked here refer to several fields included in the dataset to be detected, such as the "name" field or the "age" field. Since one work task is used to process a type of data quality problem, a single field or multiple fields can be selected when configuring each work task. When selecting fields, it can be selected from the perspectives of importance, severity of data quality problems, and field relevance based on the understanding of the dataset to be detected and the business and uses to which the dataset to be detected belongs.
[0024] The check rules and correction methods can be stored procedures implemented in SQL, code files implemented in various programming languages, model files, and algorithm code files.
[0025] Specifically, the check rules are used to check whether there are data quality problems in the data. The data quality problems mainly include, but are not limited to, consistency, uniqueness, integrity, accuracy, timeliness, relevance, etc.
[0026] Consistency refers to whether the data follows a unified specification, that is, whether the data set maintains a unified format, including cases where the data format is not standardized or not unified.
[0027] Uniqueness means that for a certain data item or a group of data, the values must be unique, such as exclusive data like ID types.
[0028] Integrity refers to whether there is a situation of missing data information. Data missing may be the missing of the entire data record or the missing of the record of a certain field information in the data.
[0029] Accuracy refers to whether there are abnormalities or errors in the information recorded in the data. Common data accuracy errors include garbled characters, abnormally large or small data values, etc.
[0030] Timeliness refers to the time interval from the generation of data to its being viewable, also called the latency duration of the data. If the time for data establishment and analysis is too long, it will cause the analysis results to become outdated.
[0031] Relevance refers to the association relationships between various data sets, such as functional relationships, correlation coefficients, primary-foreign key relationships, index relationships, etc. Ensuring the relevance between data helps to maintain the integrity and consistency of the data and improve the accuracy of data analysis and mining.
[0032] When configuring each work task, preset inspection rules can be selected or new inspection rules can be created according to the specifications of the fields to be inspected. The specifications of the fields mainly include aspects such as naming specifications, data type specifications, default value settings, description information, etc. For example, if the type of a field to be inspected is numeric, the inspection rule is to report an error and mark it if it is found that the type of a certain data in the field to be inspected is not numeric. Among them, the preset inspection rules can refer to existing common data inspection rules.
[0033] According to the differences in the inspection content, the inspection rules can be implemented in one or more ways of SQL, Python, other programming languages, and artificial intelligence algorithms.
[0034] The role of the correction method is to correct the incorrect data according to the data standards and specifications in one or more ways of SQL, Python, other programming languages, and artificial intelligence algorithms when incorrect data is detected. The artificial intelligence algorithms can be machine learning algorithms, deep learning algorithms, etc. In the inspection rules, the algorithm is trained to identify normal data patterns and regard data significantly different from these normal data as problem data. In the correction method, the algorithm is trained to generate data consistent with the normal data to replace the original problem data or null data.
[0035] When configuring each work task, according to the data standards and specifications of the fields to be checked, select a preset correction method or create a new correction method for the data with problems. Among them, the preset correction method can refer to the existing common data correction method. In this embodiment, data standards and specifications refer to a set of standards, rules and conventions for describing and defining data, which defines the structure, format, naming rules, data type, validity verification rules, etc. of the data to ensure the consistency, accuracy and availability of the data.
[0036] For different quality issues in the data, correction methods can be designed in a targeted manner. Specifically, the correction methods include: For inconsistent data formats, use SQL or other programming languages to convert the formats; For the uniqueness problem, a deduplication algorithm is used to remove duplicate data; For missing value problems, standard values are queried in the basic database based on the values of related fields, or data is generated using artificial intelligence algorithms to fill in the missing values. For accuracy issues, AI algorithms are used to identify and correct problematic data; For timeliness issues, timestamps are used to detect and handle them accordingly.
[0037] When a large amount of data needs to be processed, the number of work tasks may be large, which will consume a lot of computing resources and time, so it is necessary to schedule the work tasks according to certain rules. Specifically, different scheduling methods can be selected according to the system's resource status, data availability time, the urgency of data demand, and the logical order of data inspection. Among them, data availability time refers to the validity period or availability period of data. This is a concept commonly used in data management and information technology, mainly used to specify the time range for data storage, use, access or deletion.
[0038] Specifically, scheduling methods include determination method, priority method, and workload method; The deterministic method refers to performing work tasks according to specified cycles and time points.
[0039] Priority method refers to executing tasks according to the priority in the configuration content of the tasks, and the tasks with higher priority are executed first. For tasks with the same priority, the system decides whether to execute them simultaneously or randomly select some of the tasks to execute first based on the available resources. Priority method is generally used for tasks with urgent data or benchmark data that other data depends on.
[0040] The workload method is to evaluate the workload before executing the task, and the tasks with low workload are executed first. The advantage of the workload method is that most tasks can be completed in a shorter time.
[0041] The execution module specifically implements the execution of the work tasks selected by the scheduling module. After the execution of the work tasks is completed, corresponding work reports will also be generated. The work report is made for the work tasks. The work report includes inspection results and correction results, and its specific content includes: the forms and fields being inspected, the names and meanings of the inspection rules used, the quantity and proportion of unqualified records detected, the names and meanings of the correction methods used, and the quantity and proportion of the records being corrected.
[0042] The auditing module mainly audits the inspection results and correction results in the work report, and specifically can be audited and corrected manually or by artificial intelligence.
[0043] The auditing in the auditing module is divided into two levels: rule auditing and record auditing.
[0044] For each record being corrected, it is judged whether the inspection result and the correction result meet the expected values. If they meet the expected values, the corresponding inspection rules and correction methods are accepted; otherwise, the corresponding inspection rules and correction methods are rejected. Statistics are carried out on all the records being corrected that involve the same inspection rule. If the acceptance proportion is greater than the preset threshold, the inspection rule is accepted; otherwise, the inspection rule needs to be revised or a new inspection rule is used, and the work tasks are reconfigured and scheduled for execution.
[0045] Among them, in the case of accepting the inspection rule, the modification of some records can be retained as appropriate during the auditing process.
[0046] Specifically, the auditing and correction in the auditing module can also be carried out on the sampled fields and records. If the data volume is large and there are many inspection rules, manual auditing takes a long time and has a high cost. At this time, sampling is selected at all record and / or field levels, and auditing and correction are carried out on the sampled data, which can solve problems such as insufficient resources and long time consumption, and is beneficial to improving the auditing efficiency.
[0047] After the result data set and the work report pass the auditing, the output module outputs the result data set and the work report, saves the work report in file formats such as word and PDF according to the specified path, and saves the result data set in the corresponding table or data file according to the specified table name and path name.
[0048] Embodiment 2 As Figure 2 shown, the present invention also provides a method for data quality detection and automatic correction, including the following steps: Step S1: Select the data set to be detected; Step S2: Configure several work tasks for the dataset to be detected. Each work task is used to handle a type of data quality problem. The configuration content of each work task includes fields to be checked, checking rules, correction methods, and priorities. Step S3: Select a scheduling method and schedule the configured work tasks according to the scheduling method. Step S4: Execute the work tasks according to the scheduling results, and check and correct the dataset to be detected according to the configuration content of the work tasks to obtain a result dataset. After the work tasks are executed, generate corresponding work reports. Step S5: Review and revise the result dataset and the work reports. Step S6: After the review and revision are completed, output the result dataset and the work reports.
[0049] In summary, compared with the prior art, a device and method for data quality detection and automatic correction provided by the present invention have the following beneficial effects: 1. Automation: The present invention can realize the discovery and automatic correction of data quality problems, reducing manual intervention and errors.
[0050] 2. Multifunctionality: The present invention can discover various data quality problems existing in the dataset. For some problems, the system can automatically correct the problem data, such as non-standard formats, non-unique data, incomplete data, inaccurate data, etc.
[0051] 3. High efficiency: The present invention can automatically schedule work tasks, quickly process the quality problems of a large amount of data, and improve the efficiency of data processing and analysis.
[0052] 4. Scalability: The present invention can be extended to various data sources, data types, and problems, perform customized checks and corrections, and has good scalability.
[0053] The above has introduced in detail a device and method for data quality detection and automatic correction provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the technical solution and its core idea of the present application; those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A device for data quality detection and automatic correction, characterized in that, Including: Input module: used to input the dataset to be detected; Configuration module: used to configure several work tasks for the dataset to be detected, each work task is used to process a type of data quality problem, and the configuration content of each work task includes fields to be checked, check rules, correction methods, and priorities; Scheduling module: used to set the scheduling method and schedule the configured work tasks according to the scheduling method; Execution module: used to execute the work tasks according to the scheduling results of the scheduling module, and check and correct the dataset to be detected according to the configuration content of the work tasks to obtain the result dataset, and generate corresponding work reports after the work tasks are completed; Review module: used to receive the result dataset and the work report, and review and correct the result dataset and the work report; Output module: used to output the result dataset and the work report after the review and correction by the review module.
2. The device for data quality detection and automatic correction according to claim 1, wherein The configuration module includes: When configuring each work task, several of the fields to be checked are selected based on the understanding of the dataset to be detected and the business and uses to which the dataset to be detected belongs, and from the perspectives of importance, severity of data quality problems, and field relevance.
3. The device for data quality detection and automatic correction according to claim 1, characterized in that, The configuration module includes: When configuring each work task, the preset check rules are selected or new check rules are created according to the specifications of the fields to be checked.
4. The device for data quality detection and automatic correction according to claim 1, wherein The configuration module includes: When configuring each work task, preset correction methods or new correction methods are selected for the data in the fields to be checked where problems are detected according to the data standards and specifications of the fields to be checked.
5. The device for data quality detection and automatic correction according to claim 1, characterized in that, The correction methods specifically include: For problems with inconsistent data formats, SQL or other programming languages are used for format conversion; For uniqueness problems, duplicate data is deleted using a deduplication algorithm; For missing value problems, standard values are queried from the basic database based on the values of related fields for filling, or artificial intelligence algorithms are used to generate data for filling; For accuracy problems, artificial intelligence algorithms are used to identify and correct problem data; For timeliness problems, timestamps are used for detection and corresponding processing.
6. The device for data quality detection and automatic correction according to claim 1, wherein The scheduling module includes: Different scheduling methods are selected according to the resource status of the system, data availability time, urgency of data requirements, and logical order of data checks.
7. The device for data quality detection and automatic correction according to claim 1, characterized in that, The scheduling methods include the determination method, the priority method, and the workload method; The determination method means that the work tasks are executed according to the specified cycle and time points; The priority method means that the work tasks are executed according to the priorities in the configuration content of the work tasks. For work tasks with the same priority, it is decided whether to execute them simultaneously or randomly select some of the work tasks to execute first according to the available resource status of the system; The workload method means that the workload of the work tasks is evaluated before they are executed, and the work tasks with low workload are executed first.
8. The device for data quality detection and automatic correction according to claim 1, characterized in that, The work report includes inspection results and correction results, and its specific contents include: the forms and fields inspected, the names and meanings of the inspection rules used, the quantity and proportion of unqualified records detected, the names and meanings of the correction methods used, and the quantity and proportion of the records corrected.
9. The device for data quality detection and automatic correction according to claim 8, wherein The audit module includes the audit of the inspection results and the correction results, and includes: For each record with data corrected, determine whether the inspection results and the correction results meet the expected values. If they meet the expected values, accept the corresponding inspection rules and correction methods; otherwise, reject the corresponding inspection rules and correction methods; Statistically analyze all the corrected records related to the same inspection rule. If the acceptance proportion is greater than the preset threshold, accept the inspection rule; otherwise, it is necessary to revise the inspection rule or use a new inspection rule, reconfigure the work task and schedule its execution.
10. A method for data quality detection and automatic correction, characterized in that, It includes the following steps: Step S1: Select the dataset to be detected; Step S2: Configure several work tasks for the dataset to be detected. Each work task is used to handle a type of data quality problem, and the configuration content of each work task includes the fields to be inspected, inspection rules, correction methods, and priorities; Step S3: Select a scheduling method and schedule the configured work tasks according to the scheduling method; Step S4: Execute the work tasks according to the scheduling results, and perform inspections and corrections on the dataset to be detected according to the configuration content of the work tasks to obtain a result dataset, and generate a corresponding work report after the work tasks are completed; Step S5: Audit and revise the result dataset and the work report; Step S6: After the audit and revision are completed, output the result dataset and the work report.