Data quality and security processing system and method based on automatic matching rule
Through a data quality and security processing system based on automatic matching rules, the problems of low efficiency, poor rule scalability and threshold static in traditional data processing are solved, efficient and accurate data inspection and security processing are achieved, and the cost of manual intervention is reduced.
Patent Information
- Application Number
- CN202510402965.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, data quality inspection and safety processing are inefficient, rules are poorly scalable, verification efficiency is low, problem positioning is difficult, threshold value is static, and it is difficult to quickly respond to changes in business demand and data distribution.
The data quality and security processing system based on automatic matching rules is adopted, including the rule dynamic configuration module, the multi-source data acquisition module, the data feature identification module, the distributed computing engine and the closed-loop feedback module, and the rule threshold is optimized through JSON structured description, Spark parallel execution and history check.
It improves the efficiency and accuracy of data quality inspection, enhances data security, realizes flexible data processing, and reduces errors and hidden dangers caused by manual intervention.
Smart Images

Figure CN120336302A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data governance, and specifically relates to a data quality and security processing system and method based on automatic matching rules. Background Art
[0002] As data becomes an important asset of enterprises, the quality and security of data are crucial for enterprise decision-making and operation. However, in the existing data processing process, data quality inspection and security processing often rely on manually set rules and strategies, which are not only inefficient but also error-prone. At the same time, with the continuous increase in the amount of data, traditional data quality inspection methods have been difficult to meet the requirements of high efficiency and accuracy.
[0003] The following defects exist in the prior art:
[0004] Poor rule scalability: Traditional solutions require developing independent code for each rule, resulting in a bloated system and difficulty in quickly responding to changes in business requirements.
[0005] Low verification efficiency: The single-machine execution engine cannot handle the verification of massive data, and the timeliness is difficult to guarantee.
[0006] Difficult problem location: Lack of data lineage tracing ability, and it takes a long time to trace the source of abnormal data.
[0007] Threshold staticization: The rule threshold depends on manual experience setting and cannot adapt to changes in data distribution. Summary of the Invention
[0008] In view of the above problems, the present invention is proposed to provide a solution that overcomes or at least partially solves the above problems.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] A data quality and security processing system based on automatic matching rules, the system includes:
[0011] A rule dynamic configuration module that realizes the dynamic mapping of rule parameters and the front-end form through JSON structured description;
[0012] A multi-source data collection module for docking heterogeneous data sources;
[0013] A data feature recognition module that extracts data features through feature recognition technology and automatically matches and verifies rules;
[0014] A distributed computing engine module that realizes the parallel execution of data quality inspection and data security processing tasks based on Spark;
[0015] A closed-loop feedback module that optimizes the rule threshold through historical inspection results.
[0016] Optionally, the rule dynamic configuration module includes:
[0017] A rule description unit that abstracts rules into JSON structures, including input parameters and verification formulas;
[0018] A front-end rendering unit that dynamically generates form elements according to the JSON structure and supports rule classification;
[0019] A rule library unit for storing predefined rule templates.
[0020] Optionally, the predefined rule templates include null value check, field length check, two-table consistency check, ID number desensitization, and bank card number desensitization.
[0021] Optionally, the multi-source data collection module includes a connector unit that supports access to JDBC, HDFS, and Kafka data sources and is used to collect data to be processed from the data sources.
[0022] Optionally, the data sources include databases, file systems, and network interfaces.
[0023] Optionally, the data feature recognition module includes:
[0024] A data feature extraction unit that extracts data features from the data set through clustering algorithms and deep learning model techniques;
[0025] A rule matching engine unit that automatically matches rules according to the extracted data features.
[0026] Optionally, the distributed computing engine module includes:
[0027] A task scheduler unit that generates data processing tasks according to the automatically matched verification rules;
[0028] A Spark distributed execution engine unit that uses Spark to execute data processing tasks in parallel and adjusts the allocation of computing resources according to the dynamic resource allocation strategy.
[0029] Optionally, the closed-loop feedback module includes:
[0030] A result analysis unit that counts the rule trigger frequency and false alarm rate and generates a data quality report;
[0031] A threshold optimization unit that automatically adjusts the rule threshold based on the historical data distribution.
[0032] The present invention also provides a data quality and security processing method based on automatically matched rules. The method is based on the system described in any one of the foregoing, and the method includes the following steps:
[0033] S1. Define rules according to the requirements of data quality inspection and security processing, including the parameter input items of the rules, SQL definitions, and parameters related to result judgment;
[0034] S2. Achieve decoupling of rule configuration and the front-end interface through JSON structure description;
[0035] S3. Use clustering algorithms and deep learning models to extract key data features from the dataset, and the rule matching engine automatically matches verification rules according to the extracted data features;
[0036] S4. The distributed computing engine generates a data processing task according to the matched verification rules, automatically calculates the resources required for the current task and the number of concurrent tasks according to the data volume of the current data processing task and the type of verification rules, and decomposes the quality inspection task and data security processing task into multiple subtasks according to the calculation results;
[0037] S5. Call Spark to execute the subtasks in S4 in parallel, execute SQL to calculate statistical values, and perform formula verification with the comparison values. Among them, the data quality inspection task marks or corrects the data that does not conform to the rules;
[0038] S6. The closed-loop feedback module stores the processed data in the specified storage medium, generates a data quality report according to the calculation results, and triggers an alarm if there are abnormal results;
[0039] S7. According to the set sliding window calculation method, recalculate the threshold based on the historical data distribution of three months and automatically adjust the rule threshold.
[0040] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:
[0041] 1. The present invention improves the efficiency and accuracy of data quality inspection, realizes efficient inspection and accurate judgment of data through an automated method; enhances data security, and can timely discover and handle security risks in data by defining security processing rules.
[0042] 2. The present invention improves the flexibility of data processing. Users can customize rules according to actual needs to achieve personalized data quality inspection and security processing; reduces the cost of manual intervention and reduces data quality problems and security risks caused by human errors. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic flowchart of a data quality and security processing system based on automatic matching rules provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without any creative work fall within the scope of protection of the present invention.
[0045] Embodiment 1:
[0046] Please refer to Figure 1 , this embodiment provides a data quality and security processing system based on automatic matching rules. The system includes:
[0047] A rule dynamic configuration module that realizes the dynamic mapping between rule parameters and the front-end form through JSON structured description.
[0048] The rule dynamic configuration module includes:
[0049] A rule description unit that abstracts the rule into a JSON structure, including input parameters (data source, statistical value calculation logic, comparison value type) and verification formulas (comparison operator, threshold).
[0050] A front-end rendering unit that dynamically generates form elements according to the JSON structure and supports rule classification (accuracy / completeness / timeliness / normativity / consistency).
[0051] A rule library unit for storing predefined rule templates. The predefined rule templates include null value check, field length check, two-table consistency check, ID number desensitization, and bank card number desensitization.
[0052] A multi-source data collection module for connecting to heterogeneous data sources.
[0053] The multi-source data collection module includes a connector unit. The connector unit supports the access of data sources such as JDBC, HDFS, and Kafka, and is used to collect the data to be processed from the data source. The data sources include databases, file systems, and network interfaces.
[0054] A data feature recognition module that extracts data features through feature recognition technology and automatically matches and verifies rules.
[0055] The data feature recognition module includes:
[0056] A data feature extraction unit that extracts data features from the data set through clustering algorithms and deep learning model technologies.
[0057] A rule matching engine unit that automatically matches rules according to the extracted data features.
[0058] The distributed computing engine module implements the parallel execution of data quality inspection and data security processing tasks based on Spark.
[0059] The distributed computing engine module includes:
[0060] A task scheduler unit that generates data processing tasks according to automatically matched verification rules.
[0061] A Spark distributed execution engine unit that uses Spark to execute data processing tasks in parallel and adjusts the allocation of computing resources according to the dynamic resource allocation strategy.
[0062] A closed-loop feedback module that optimizes the rule threshold through historical inspection results.
[0063] The closed-loop feedback module includes:
[0064] A result analysis unit that counts the rule trigger frequency and false alarm rate and generates a data quality report;
[0065] A threshold optimization unit that automatically adjusts the rule threshold based on historical data distribution (such as the sliding window algorithm).
[0066] This embodiment improves the efficiency and accuracy of data quality inspection, realizes the efficient inspection and accurate judgment of data in an automated manner; enhances data security, and can timely detect and handle security risks in data by defining security processing rules.
[0067] This embodiment improves the flexibility of data processing. Users can customize rules according to actual needs to achieve personalized data quality inspection and security processing; reduces the cost of manual intervention and reduces data quality problems and security risks caused by human errors.
[0068] Embodiment 2:
[0069] Embodiment 2 discloses a data quality and security processing method based on automatically matched rules. This method is based on the above-mentioned data quality and security processing system based on automatically matched rules. The method includes the following steps:
[0070] S1. Define rules according to the requirements of data quality inspection and security processing, including parameter input items of the rules, SQL definition, and parameters related to result judgment.
[0071] S2. Achieve the decoupling of rule configuration and the front-end interface through JSON structure description. When adding new rules, there is no need to modify the code, while ensuring the simplicity of the front-end code and the user experience.
[0072] S3. Use data feature extraction technologies such as clustering algorithms and deep learning models to extract key data features from the dataset, and the rule matching engine automatically matches the verification rules according to the extracted data features.
[0073] S4. The distributed computing engine generates a data processing task according to the matched verification rules. According to the data volume of the current data processing task and the type of verification rules, it automatically calculates the resources required for the current task and the number of concurrent tasks, and decomposes the quality inspection task and the data security processing task into multiple subtasks according to the calculation results, including the execution method of the task (single execution or periodic execution), the task trigger condition, and the environment and resources required for task execution.
[0074] S5. Call Spark to execute the subtasks in step S4 in parallel, execute SQL to calculate statistical values (such as the number of null rows), and perform formula verification with the comparison values (such as: statistical value ≤ threshold). Among them, the data quality inspection task marks or corrects the data that does not conform to the rules.
[0075] For example, if a certain field is a required field but is empty in the data, mark that there is a quality problem with this piece of data. If the data format does not meet the requirements, try to perform format conversion or prompt an error. The data security verification task judges whether the data contains sensitive information according to the sensitive data identification rules. If it contains sensitive information, perform corresponding processing according to the access permission rules, such as encryption and desensitization.
[0076] S6. The closed-loop feedback module stores the processed data in the specified storage medium, generates a data quality report according to the calculation results, and triggers an alarm if there are abnormal results.
[0077] S7. According to the set sliding window calculation method, recalculate the threshold based on the historical data distribution of the past three months and automatically adjust the rule threshold.
[0078] The following combines specific application cases to further describe the technical solution in detail:
[0079] According to the requirements of data quality inspection and security processing, a series of rules are defined. For example, for data quality inspection, a rule is defined to check whether the "user name" field is empty. The parameter input items of the rule include the field name "user name", and the SQL definition is: SELECT COUNT(*) FROM table name WHERE user name IS NULL. The parameter related to the result judgment is "if the count is greater than 0, there are null values".
[0080] Step 1: Rule configuration is decoupled from the front end. JSON Schema is used to define rule parameters and SQL templates, and a cascading selector is supported to achieve three-level linkage of cluster → library → table. When adding a new rule, only a new JSON configuration needs to be added without modifying the code. Each rule consists of two parts: parameter input items and SQL logic definition, forming a complete quality inspection closed loop.
[0081] Details of the parameter input items are shown in the table:
[0082]
[0083]
[0084] Table 1 Parameter input items In this application case, the above rules can be described in the following JSON format:
[0085]
[0086]
[0087] Step 2: Data feature extraction and rule matching. Advanced data feature extraction technologies such as clustering algorithms and deep learning models are used to identify key features from the original dataset. This process is crucial for automatically matching appropriate verification rules. By analyzing the patterns and distributions in the dataset, corresponding rules can be more accurately applied for data quality inspection or security processing. For example, in this instance, based on the DBSCAN clustering algorithm, abnormal fields are identified, combined with metadata analysis (field type, constraint conditions)
[0088] It is found that there are a large number of null values in the "user name" field, and the rule matching engine automatically matches the above-defined rules.
[0089] Step 4: Distributed computing and task decomposition. The distributed computing engine generates data processing tasks according to the matched rules. The system automatically calculates the required resources and the number of concurrent tasks according to the task data volume and rule type. For example, in this instance, the quality inspection task is decomposed into multiple subtasks, and each subtask is responsible for checking a part of the data.
[0090] Step 5: Subtask execution and data verification. The distributed computing engine calls the Spark cluster to execute subtasks in parallel, executes SQL to calculate statistical values, such as the number of null rows. For data that does not conform to the rules, it is marked or corrected. For example, when executing the task in this instance, data entries with a null value in the "user name" field are found and marked.
[0091] Step 6: Closed-loop Feedback and Report Generation. Based on the processing results, the system automatically generates a detailed data quality report, which not only includes the overall quality status but also lists specific anomalies and their locations. For example, the report shows the number of entries with null values in the "User Name" field and triggers an alert to notify the administrator.
[0092] Step 7: Automatic Threshold Adjustment. To adapt to the ever-changing data environment, the system also recalculates the thresholds based on historical data regularly (such as every three months) and automatically adjusts the rule thresholds to maintain the optimal data quality and security levels. For example, in this instance, based on the data distribution in the past three months, the system finds that the proportion of null values in the "User Name" field is gradually increasing, so it automatically adjusts the threshold to monitor this field more strictly.
[0093] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A data quality and security processing system based on automatic matching rules, characterized in that The system includes: A rule dynamic configuration module that realizes the dynamic mapping between rule parameters and the front-end form through JSON structured description; A multi-source data collection module for connecting to heterogeneous data sources; A data feature recognition module that extracts data features through feature recognition technology and automatically matches and verifies rules; A distributed computing engine module that realizes the parallel execution of data quality inspection and data security processing tasks based on Spark; A closed-loop feedback module that optimizes the rule threshold through historical inspection results.
2. The data quality and security processing system based on an automatic matching rule as described in claim 1, wherein The rule dynamic configuration module includes: A rule description unit that abstracts rules into JSON structures, including input parameters and verification formulas; A front-end rendering unit that dynamically generates form elements according to the JSON structure and supports rule classification; A rule library unit for storing predefined rule templates.
3. A data quality and security processing system based on an automatic matching rule as claimed in claim 2, wherein The predefined rule templates include null value check, field length check, two-table consistency check, ID card number desensitization, and bank card number desensitization.
4. A data quality and security processing system based on an automatic matching rule as described in claim 1, characterized in that, The multi-source data collection module includes a connector unit that supports the access of JDBC, HDFS, and Kafka data sources and is used to collect data to be processed from the data sources.
5. A data quality and security processing system based on an automatic matching rule as claimed in claim 4, characterized in that, The data sources include databases, file systems, and network interfaces.
6. A data quality and security processing system based on an automatic matching rule as claimed in claim 1, wherein The data feature recognition module includes: A data feature extraction unit that extracts data features from the data set through clustering algorithms and deep learning model technologies; A rule matching engine unit that automatically matches rules according to the extracted data features.
7. A data quality and security processing system based on an automatic matching rule as described in claim 1, wherein The distributed computing engine module includes: A task scheduler unit that generates data processing tasks according to the automatically matched verification rules; A Spark distributed execution engine unit that uses Spark to execute data processing tasks in parallel and adjusts the allocation of computing resources according to the dynamic resource allocation strategy.
8. A data quality and security processing system based on an automatic matching rule as claimed in claim 1, wherein, The closed-loop feedback module includes: A result analysis unit that counts the rule trigger frequency and false alarm rate and generates a data quality report; A threshold optimization unit that automatically adjusts the rule threshold based on the historical data distribution.
9. A method for data quality and security processing based on an automatic matching rule, which is based on the data quality and security processing system according to any one of claims 1-8, and is characterized in that, The method includes the following steps: S1. Define rules according to the requirements of data quality inspection and security processing, including parameter input items of the rules, SQL definitions, and parameters related to result judgment; S2. Realize the decoupling of rule configuration and the front-end interface through JSON structure description; S3. Use clustering algorithms and deep learning models to extract key data features from the data set, and the rule matching engine automatically matches verification rules according to the extracted data features; S4. The distributed computing engine generates a data processing task according to the matched verification rules, automatically calculates the resources required for the current task and the number of concurrent tasks according to the data volume of the current data processing task and the type of verification rules, and decomposes the quality inspection task and the data security processing task into multiple subtasks according to the calculation results; S5. Call Spark to execute the subtasks in step S4 in parallel, execute SQL to calculate statistical values, and perform formula verification with comparison values. Among them, the data quality inspection task marks or corrects data that does not conform to the rules; S6. The closed-loop feedback module stores the processed data in the specified storage medium, generates a data quality report according to the calculation results, and triggers an alarm if there are abnormal results; S7. Based on the set sliding window calculation method, recalculate the threshold based on the historical data distribution of the past three months and automatically adjust the rule threshold.
Citation Information
Patent Citations
Data verification system and method under front-end and rear-end separation architecture
CN110569036A
Data quality inspection method and system
CN111400288A
Data quality checking system, method, computer equipment and storage medium
CN113886370A
Dynamic strategy generation method based on JSON
CN118860346A
Automatic data quality rule generation method and system based on data standard
CN118861013A
Cited By
Data processing method and device, electronic equipment and storage medium
CN121217838A