Method for automatically generating quality check rule based on data standard

By automatically matching database fields and data standards, quality inspection rules are generated, and the problem of separation between data standards and quality inspection rules definitions in the existing technology is solved, and efficient and accurate data quality management and real-time update capabilities are achieved.

CN120162322APending Publication Date: 2025-06-17INSPUR GENERSOFT CO LTD

Patent Information

Application Number
CN202510305868.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, the definition of data standards and the definition of quality inspection rules are separated, and there is a lack of automation connection, resulting in inefficiency, increasing workload and error rate, and there may be inconsistencies between the data standards and the inspection rules, affecting the accuracy of data quality assessment.

Method used

By defining data standards, matching database fields with data standards, determining the verification scope, generating mapping relationships, and automatically generating quality verification rules based on predefined verification rules templates, data standards and mapping relationships.

Benefits of technology

It realizes the automated connection between data standards and quality inspection rules, improves the efficiency of rule definition, reduces manual operation errors, ensures the consistency between data standards and inspection rules, improves the accuracy of data quality assessment, and supports real-time updates and dynamic management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162322A_ABST
    Figure CN120162322A_ABST
Patent Text Reader

Abstract

The invention provides a method for automatically generating a quality check rule based on a data standard, and relates to the technical field of data governance, and the method comprises the steps: defining the data standard, matching a database field with the data standard, and determining a check range to generate a mapping relation; based on a predefined checking rule template, the data standard and the mapping relation, generating a quality checking rule; and triggering and executing the quality checking rule according to a preset period or event.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data governance, and particularly relates to a method for automatically generating quality inspection rules based on data standards. Background Art

[0002] As the importance of data assets in enterprise operations increases day by day, data quality management has become the core link of data governance. Data quality inspection rules are the key tools to ensure data accuracy and consistency. Traditional data quality management relies on manually defining inspection rules, but with the expansion of data scale and the complexity of data standards, the inefficiency and high error rate of manual operations have gradually emerged. In recent years, data middle platforms and data governance platforms have gradually become popular, but there are still obvious deficiencies in the automated generation of data standards and quality inspection rules.

[0003] Currently, the mainstream data quality management process is usually divided into two independent steps: defining the business attributes, technical attributes, and quality attributes of data through a data standard module. Manually defining inspection rules through a data quality module, or referring to some attributes in the data standard. The typical process of the prior art is as described in Patent CN 110119395 A. The definition of data standards and the definition of quality inspection rules are separated and need to be completed manually respectively. For example, when defining data standards, although the quality requirements of data (such as value range, format specification, etc.) are clarified, when defining quality inspection rules, these standard attributes still need to be manually referred to, and automated generation cannot be achieved.

[0004] Therefore, the inventors found that the definition of data standards and the definition of quality inspection rules are two independent steps, lacking automated connection, resulting in low efficiency. It is necessary to manually define inspection rules or manually refer to the attributes in the data standard, which increases the workload and error rate. Due to manual operations, there may be inconsistencies between data standards and inspection rules, affecting the accuracy of data quality assessment. When data standards or mapping relationships change, it is necessary to manually redefine inspection rules and real-time synchronization cannot be achieved. Summary of the Invention

[0005] This application provides a method for automatically generating quality inspection rules based on data standards to solve one of the above technical problems.

[0006] The technical solution adopted by this application is as follows: An embodiment of this application provides a method for automatically generating quality inspection rules based on data standards, including: Defining data standards, matching database fields with the data standards, and determining the inspection scope to generate a mapping relationship; Generating quality inspection rules based on a predefined inspection rule template, the data standards, and the mapping relationship; Execute the quality inspection rule according to a preset period or event trigger.

[0007] According to an embodiment of the present application, for the defined data standard, match the database fields with the data standard, determine the inspection scope, and generate a mapping relationship, specifically: Divide the attributes of the data standard into business attributes, technical attributes, management attributes, and quality attributes; For each of the data standards, clarify the attributes related to quality; Associate the database table fields with the corresponding data standards one-to-one or many-to-one; Clarify the inspection scope, that is, the mapped fields are used as the target fields for quality inspection.

[0008] According to an embodiment of the present application, the predefined inspection rule simulation is specifically: Check whether the field is empty; Check whether the field value is within the enumerated range; Check whether the numerical value is within the specified interval; Verify the format through regular expressions; Verify whether the field value exists in the reference data table.

[0009] According to an embodiment of the present application, for generating the quality inspection rule based on the predefined inspection rule template, the data standard, and the mapping relationship, specifically: Extract the attributes related to quality from the data standard; According to the mapping relationship, clarify the database tables and fields to be inspected; Match the attributes of the data standard with the predefined inspection rule template, and bind them with the mapped fields, inspection logic, and inspection parameters to generate a quality inspection rule.

[0010] According to an embodiment of the present application, the matching logic includes: null value rule, value range rule, threshold rule, format specification rule, and data correlation rule.

[0011] According to an embodiment of the present application, after executing the quality inspection rule according to the preset period or event trigger, it further includes: Run the quality inspection rule, identify the data that does not meet the quality requirements as abnormal data; Record the abnormal data in the quality report and trigger an alarm notification; Support manual or automatic repair processes.

[0012] According to an embodiment of the present application, after executing the quality inspection rule according to the preset period or event trigger, it further includes: Monitoring update events of the data standard library and mapping relationships; Re-match the template according to the changed content and generate new rules, overwriting the old rules; Keep historical rule versions and support rollback and auditing.

[0013] A computer program product containing instructions, when running on a device, is characterized in that it enables the device to execute the steps in a method for automatically generating quality check rules based on data standards.

[0014] A computer-readable storage medium stores a program, which, when executed by a processor, implements the steps in a method for automatically generating quality check rules based on data standards.

[0015] An electronic device comprises a memory, a processor and a program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method of automatically generating quality inspection rules based on data standards are implemented.

[0016] Due to the adoption of the above technical solution, the beneficial effects achieved by this application are as follows: This application clarifies the scope of verification by automatically matching database fields with data standards, ensuring the pertinence and accuracy of verification rules and reducing manual intervention. Using predefined templates and data standard attributes, verification rules are automatically generated, significantly improving the efficiency of rule definition and avoiding manual operation errors. : Through timing or event triggering mechanisms, verification rules are automatically executed to ensure real-time monitoring and dynamic updating of data quality.

[0017] This application connects the definition of data standards with the generation of quality verification rules to form a coherent automated process, significantly improving the efficiency of data governance. By automatically generating rules, the consistency between data standards and verification rules is ensured, and the accuracy of data quality assessment is improved. : When data standards or mapping relationships change, the verification rules are automatically updated to ensure the real-time and effectiveness of the rules. This greatly reduces manual operations, reduces workload and error rates, and improves the intelligence level of data governance. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A flowchart of a method for automatically generating quality check rules based on data standards provided in an embodiment of the present application. DETAILED DESCRIPTION

[0019] To more clearly illustrate the overall concept of this application, the following will provide a detailed description by way of examples in conjunction with the accompanying drawings of the specification.

[0020] In the following description, numerous specific details are set forth in order to provide a thorough understanding of this application. However, this application may be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited by the specific embodiments disclosed below. It should be noted that, without conflict, the embodiments of this application and the features in each embodiment may be combined with each other.

[0021] In this application, unless otherwise clearly specified and defined, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0022] Embodiment 1 As Figure 1 shown, a method for automatically generating quality inspection rules based on data standards includes: Defining data standards, matching database fields with the data standards, and determining the inspection scope to generate a mapping relationship.

[0023] Specifically, the attribute classification of data standards: First, it is necessary to divide the attributes of data standards into different categories, including business attributes, technical attributes, management attributes, and quality attributes. Each attribute category has its specific content and uses.

[0024] Clarifying the attributes related to quality: For each data standard, clarify which attributes are directly related to data quality. For example, "nullable", "format specification", "value range constraint", etc. are all technical attributes that directly affect data quality.

[0025] The association between fields and data standards: Next, associate the database table fields with the corresponding data standards one-to-one or many-to-one. This means that one or more database fields can be mapped to the same data standard to ensure data consistency and accuracy.

[0026] Determine the verification scope: After completing the above associations, clarify the verification scope. Specifically, identify which mapped fields will serve as the target fields for quality verification. These fields will become the key objects for subsequent data quality assessment.

[0027] Generate mapping relationships: The last step is to generate mapping relationships based on the above analysis results. This not only involves matching database fields with data standards but also includes generating specific quality verification rules according to the quality-related attributes extracted from the data standards and combining predefined verification rule templates.

[0028] For example, the definition of data standards First, a series of data standards need to be defined, which include business attributes, technical attributes, management attributes, and quality attributes. As shown in Table 1, the following are standard examples of several fields in the student information table:

[0029] Table 1 Matching of database fields with data standards Next, match the fields in the actual database with the data standards defined above: ID number: According to its standard number Stu_ID, it is known that this field should be of character type, with a length of 18, not nullable, and should conform to the ID number format.

[0030] Gender: The standard number is Stu_Sex, which stipulates that it can only be "male" or "female", and is also not nullable.

[0031] Age: The standard number Stu_Age indicates that this field should be an integer, and the value should be within [6, 13].

[0032] Address: The standard number Stu_Home allows it to be nullable, but if there is a value, it should be accurate to the township / street level and refer to the administrative division dictionary.

[0033] Determine the verification scope and generate mapping relationships After clarifying the data standards corresponding to each field, the specific verification scope and rules can be determined. For example: For the ID number field, null value check rules and format specification rules will be automatically generated.

[0034] For the gender field, null value check rules and value range rules will be generated to ensure that the input is one of "male" or "female".

[0035] For the age field, null value check rules and threshold rules will be generated to ensure that the value is within the range of [6, 13].

[0036] The address field generates data association rules and verifies whether it exists in the specified administrative division dictionary.

[0037] Furthermore, when defining data standards, business attributes can be described in more detail. For example, for each data standard, in addition to the basic Chinese name and English name, information such as usage scenarios and update frequencies can be added. This helps to more accurately understand the purpose and maintenance requirements of each data item. In addition to the basic data types, lengths, and precisions, index information, encryption requirements, etc. can also be considered. For example, for fields involving sensitive information (such as ID numbers), it is clear whether encryption storage is required and which encryption algorithm to use. In the management attributes, in addition to the management department, information such as approval processes and change records can be added. In this way, when the data standard changes, it is possible to track who initiated the change, the reason for the change, and the specific change content. For quality attributes, in addition to the existing value range constraints and format specifications, data integrity rules, consistency check rules, etc. can be introduced. For example, some fields may need to be consistent with fields in other tables, and corresponding data association rules can be defined at this time.

[0038] Using machine learning algorithms, based on the existing data standards and the historical matching situations of fields, intelligent matching suggestions are provided for new fields. This not only improves the matching efficiency but also reduces human errors. Considering the changes in business requirements and technical environments, implement a mechanism that can automatically adjust the matching relationships according to new business requirements or technical updates. For example, when the data type of a certain field changes from character type to integer type, the system can automatically identify and adjust the relevant matching relationships and verification rules.

[0039] Allow the definition of different levels of verification scopes, such as at the single field level, row level, table level, or even cross-table level. This enables flexible setting of verification granularity according to different business requirements. Add a time dimension when defining the verification scope, that is, not only can the verification rules in the current state be defined, but also the verification rules at a certain future point in time or within a certain past time period can be set. This is particularly useful for processing historical data.

[0040] Based on the predefined verification rule templates, the data standards, and the mapping relationships, quality verification rules are generated.

[0041] Specifically, the predefined verification rule template The predefined verification rule template refers to a series of pre-set logics and parameters for checking whether data meets specific conditions. These templates usually include but are not limited to the following types: Null value check: Verify whether the field allows null values.

[0042] Range check: Verify whether the field value is within the specified enumeration range or numerical interval.

[0043] Format specification: Verify whether the field value conforms to the expected format (such as date format, phone number format, etc.) through regular expressions or other methods.

[0044] Data correlation check: Ensure that the field value exists in the reference data table.

[0045] Application of data standards Data standards here refer to the specific requirements established to ensure data quality, covering aspects such as business attributes, technical attributes, management attributes, and quality attributes. For example: Business attributes: Include standard number, Chinese name, English name, and business rules, etc.

[0046] Technical attributes: Include data type, length, precision, nullability, and format specification, etc.

[0047] Management attributes: Involve information such as the management department responsible for data standard maintenance.

[0048] Quality attributes: Specifically refer to requirements such as value range constraints, reference data, and value range that directly affect data quality.

[0049] Mapping relationship The mapping relationship refers to the correspondence between the actual fields in the database and the data standards. This step clarifies which database fields should follow which data standards and determines the quality verification rules that need to be carried out based on these standards.

[0050] Generate quality verification rules Combining the above three aspects, the process of generating quality verification rules can be broken down into the following steps: Extract quality-related attributes: Extract quality-related attributes from the data standards, such as nullability, format requirements, value range, etc.

[0051] Define verification objects: Based on the mapping relationship, clarify which database tables and fields need to undergo quality verification.

[0052] Match and generate rules: Match the quality attributes in the data standards with the predefined verification rule templates, and at the same time combine specific mapping fields, verification logics, and parameters to finally generate quality verification rules applicable to this field.

[0053] For example, in a student information management system, if a certain field represents the student's ID number, its data standard may stipulate that this field cannot be null and must conform to a specific format. Based on this, the system will automatically generate corresponding null value check rules and format specification rules to ensure data quality.

[0054] In addition, when changes are detected in the data standard library or mapping relationships, the system can automatically update the relevant quality inspection rules to ensure the timeliness and effectiveness of the rules. This mechanism not only reduces the need for manual intervention but also improves the overall efficiency of data quality management.

[0055] For example, as shown in Table 1, assume there is a student information table (Table 1) containing the following fields: ID number, gender, age, and address. The data standards corresponding to each field are as follows in the predefined inspection rule templates.

[0056] The predefined inspection rule templates include but are not limited to: Null value check: Verify whether the field allows null values.

[0057] Value range check: Confirm whether the field value is within the specified enumeration range or numerical interval.

[0058] Format specification: Verify whether the field value conforms to the expected format through regular expressions or other methods.

[0059] Data correlation check: Ensure that the field value exists in the reference data table.

[0060] Generate quality inspection rules Based on the above data standards and predefined inspection rule templates, we can generate specific quality inspection rules for each field: ID number field Data standard: Cannot be null and must conform to the ID number format.

[0061] Inspection rules: Null value rule: Ensure that the ID number field is not null.

[0062] Format specification rule: Use regular expressions to verify whether the ID number format is correct (e.g., ^\d{17}[\dXx]$).

[0063] Gender field Data standard: Cannot be null, and the value range is "male" or "female".

[0064] Inspection rules: Null value rule: Ensure that the gender field is not null.

[0065] Value range rule: Verify whether the value of the gender field is within the range of "male" and "female".

[0066] Age field Data standard: Cannot be null, and the value range is [6, 13].

[0067] Inspection rules: Null value rule: Ensure that the age field is not null.

[0068] Threshold rule: Verify whether the value of the age field is within the range of [6, 13].

[0069] Residential address field Data standard: It can be null, and the value range is restricted to reference data (administrative division dictionary).

[0070] Verification rule: Data relevance rule: If the residential address field has a value, verify whether the value exists in the administrative division dictionary.

[0071] Example process Suppose we want to generate quality verification rules for this student information form. The specific steps are as follows: Extract quality-related attributes from the data standard: For each field, identify its quality-related attributes such as whether it can be null, format requirements, value range, etc.

[0072] Identify the database tables and fields to be verified according to the mapping relationship: For example, the "ID number" field corresponds to a certain column in the database.

[0073] Match the attributes of the data standard with the predefined verification rule templates: For example, for the "ID number" field, since it cannot be null and has specific format requirements, the null value rule and format specification rule are matched.

[0074] Bind with the mapping fields, verification logic, and verification parameters to generate quality verification rules: Finally, generate specific verification rules for each field. For example, the rule generated for the ID number field is to ensure that it is not null and conforms to the regular expression of the ID number format.

[0075] In this way, by combining the data standard, predefined verification rule templates, and the mapping relationship between fields and standards, the system can automatically generate verification rules for ensuring data quality, thus realizing automated data quality management.

[0076] Furthermore, in addition to the existing null value check, value range check, format specification, and data relevance check, more domain-specific verification rules can be introduced. For example, for financial data, a credit score range check can be added; for medical data, a normal range check for health indicators can be added, etc. Allow users to customize the parameters of the verification rules according to specific business requirements. For example, in the numerical interval check, users can flexibly set the minimum and maximum values according to the actual situation. Set priorities for different verification rules to ensure that they are processed in order of importance during verification. For example, the data integrity of some key fields may be more important than format specification.

[0077] When business needs or external environment change, data standards can be dynamically updated and corresponding quality verification rules can be automatically regenerated. For example, when laws and regulations change requirements for certain sensitive information (such as personal privacy), the system can adjust relevant standards and update verification rules in a timely manner. Support and management of different versions of data standards are implemented to facilitate the tracing of quality requirements of historical data. This helps to accurately trace the data quality status within a specific time period during the audit process.

[0078] Use machine learning algorithms to analyze historical data and mapping relationships to provide the best matching suggestions for new fields. This will not only improve work efficiency, but also reduce the possibility of human error. Enhance the system's ability to identify and process field mappings across tables and even across databases. This is especially important for scenarios involving multiple subsystems in large enterprise applications. Establish a mechanism to monitor changes in database structure in real time and automatically update related mapping relationships and their derived quality inspection rules. Ensure that any structural changes can be quickly reflected in the data quality management process.

[0079] The quality check rules are executed according to a preset period or event trigger.

[0080] Specifically, the preset cycle triggers Set periodic tasks: Users can set the time interval for regular quality check tasks according to business needs, such as daily, weekly or monthly. This periodic check helps ensure that data remains of high quality over a long period of time.

[0081] Scheduled scheduling system: Use a scheduled scheduling system (such as cron jobs) to schedule these periodic quality inspection tasks. The system will automatically start the inspection process at the scheduled time without manual intervention.

[0082] Comprehensive coverage and focused monitoring: You can set up periodic verification tasks for the entire database or selected key tables. For particularly important data sets, you can also increase the frequency of verification to strengthen monitoring efforts.

[0083] Event triggering Define trigger events: Determine which events will trigger the execution of quality check rules. Common trigger events include but are not limited to new data import, data update operation completion, data migration completion, etc.

[0084] Real-time response mechanism: Once a specified event is detected, the system will automatically start the corresponding quality check rules for inspection. This method is particularly suitable for scenarios where data integrity and accuracy need to be verified immediately.

[0085] Custom Event Listeners: Develop specialized event listeners to capture the occurrence of specific events and trigger subsequent quality inspection processes. This approach offers a high degree of flexibility, allowing the monitoring logic to be customized according to actual requirements.

[0086] Handling after Executing Quality Inspection Rules Identifying Abnormal Data: After running the quality inspection rules, the system will identify all data that does not meet the predefined standards as abnormal data.

[0087] Recording in the Quality Report: Record all discovered abnormal situations in detail in the quality report for subsequent analysis and tracking.

[0088] Alarm Notification: Trigger alarm notifications for serious quality issues, and relevant personnel can be informed to take actions via email, SMS, or other communication tools.

[0089] Supporting the Repair Process: Provide manual or automatic ways to correct the detected problem data to ensure data accuracy and consistency.

[0090] For example, triggering by preset cycle Setting Periodic Tasks Background: Before the start of each semester in school, a comprehensive data quality check needs to be carried out on the information of all students.

[0091] Implementation Method: Set a periodic task to automatically run the quality inspection rules one week before the start of each semester. For example, a scheduling tool similar to cron can be used to set it to execute at 8:00 am on the first Monday of February and September each year.

[0092] Timed Scheduling System Specific Operations: Configure a timed task in the school's IT system, and this task will call the previously defined quality inspection rules. These rules include but are not limited to: Check whether the ID number field conforms to the format specification (such as ^\d{17}[\dXx]$).

[0093] Confirm whether the value of the gender field is one of "male" or "female".

[0094] Verify whether the value of the age field is within the range of [6, 13].

[0095] Verify whether the address field exists in the administrative division dictionary.

[0096] Comprehensive Coverage and Key Monitoring Application Scenario: For key fields in the student information table, such as ID number, name, etc., the inspection frequency can be increased, for example, an additional quality check is carried out monthly to ensure the accuracy of these important information.

[0097] Event Trigger Define Trigger Event Example: Whenever a new student registers or the personal information of an existing student changes, the data quality verification rules are immediately triggered.

[0098] Implementation Details: After an INSERT or UPDATE operation is detected through a database trigger or application logic, the corresponding quality verification process can be automatically started.

[0099] Real-time Response Mechanism Immediate Verification: Once new or modified student information is detected, the system immediately executes all relevant quality verification rules for that record. For example, if the address information of a student is updated, it is necessary to immediately check whether the address meets the format requirements and exists in the reference data table.

[0100] Custom Event Listener Development Example: Write an event listener that captures new INSERT or UPDATE requests when the database receives them and calls a pre-defined quality verification function. This function applies all quality verification rules to each affected record one by one.

[0101] Handling after Executing Quality Verification Rules Identify Abnormal Data: If any data that does not meet expectations is found, such as incorrect ID number format, age out of range, etc., these records will be marked as abnormal data.

[0102] Record in Quality Report: All abnormal situations will be recorded in a dedicated quality report, including detailed error descriptions and location information for subsequent review.

[0103] Alarm Notification: For serious problems, the system will send an alarm notification to the relevant person in charge, possibly informing the specific situation of the problem via email or text message.

[0104] Support Repair Process: Provide manual or automated tools to help correct these problems and ensure data consistency and accuracy.

[0105] Furthermore, it supports the setting of various time expressions, not limited to fixed time intervals (such as daily, weekly), but also more complex scheduling patterns can be set, such as the first working day of each month or the last week of each quarter. Provide a visual time scheduling configuration interface to enable users to intuitively define and adjust the time arrangement of task execution. Automatically adjust the task execution frequency according to the change of data volume or system load. For example, increase the inspection frequency during the period of frequent data updates, and reduce the number of inspections when the data is relatively stable to save resources. Introduce a multi-level periodic inspection mechanism, including short cycles (such as every hour) for quickly responding to immediate problems, medium cycles (such as every day) for comprehensive inspections, and long cycles (such as every month) for in-depth analysis and trend prediction.

[0106] Not limited to database operations (such as INSERT, UPDATE), it should also support other types of event triggers, such as API calls, file upload completion, third-party service callbacks, etc. Implement the support for complex event processing, such as triggering the verification process only when multiple conditions are met simultaneously. Use machine learning algorithms to automatically identify and classify different types of events to improve the accuracy and efficiency of event triggers. Automatically learn and adapt to changes in business processes, and dynamically adjust which events should trigger the quality verification rules. Support the concept of event chains, that is, after an event is triggered, a series of related subsequent events can be chained to form a complete event processing chain to ensure that all related data can be verified in a timely manner.

[0107] For different types or priorities of tasks, adopt differentiated execution strategies. For example, high-priority tasks are executed immediately, while low-priority tasks can be executed during system idle periods. Support the pause, resume, and retry mechanisms of tasks to ensure that the overall verification process is not affected in case of temporary failures. After each execution of the quality verification rules, provide a detailed execution report, including information such as execution time, the number of detected problems and their severity. Give optimization suggestions based on historical execution results to help improve future verification plans and resource allocation.

[0108] According to an embodiment of the present application, the defining of data standards, matching the database fields with the data standards, determining the verification scope to generate a mapping relationship, specifically: Divide the attributes of the data standards into business attributes, technical attributes, management attributes, and quality attributes; For each of the data standards, clarify the attributes related to quality; Associate the database table fields with the corresponding data standards one-to-one or many-to-one; Clarify the verification scope, that is, the mapped fields are used as the target fields for quality verification.

[0109] Specifically, the attribute classification of data standards First, it is necessary to classify the attributes of data standards. The data standards here are divided into the following four main categories: Business attributes: Include business-related identification information, such as standard number, Chinese name, and English name, etc. These attributes help to identify and understand the role of specific data items in the business process.

[0110] Technical attributes: Cover the technical details of data, such as data type (character type, integer type, etc.), data length, precision, and whether it can be null, etc. These attributes ensure the consistency and correctness of data storage and processing.

[0111] Management attributes: Involve information related to management and maintenance, such as which department is responsible for the update and maintenance of this data standard. This helps to track the responsibility attribution and change history.

[0112] Quality attributes: Focus on the requirements for ensuring data quality, such as value range constraints (value range or reference data), format specifications, etc. These attributes are directly related to the validity and accuracy of data.

[0113] Clarify the attributes related to quality Next, for each data standard, it is necessary to clarify which attributes are related to data quality. For example, in a student information table, the quality attributes of the "ID number" field may include not being null and conforming to a specific format (such as regular expression verification). In this way, the quality requirements that each field should meet can be specifically pointed out.

[0114] Association between fields and data standards Then, it is necessary to associate the fields in the database table with the data standards defined above one-to-one or many-to-one. This means: One data standard can correspond to multiple database fields (many-to-one); Or each database field corresponds to only one data standard (one-to-one).

[0115] For example, the "gender" field may only correspond to one data standard containing "male" and "female"; while the "address" field may need to meet the requirements of format specification and existence in a specific administrative division dictionary at the same time, which involves the association of multiple data standards.

[0116] Determine the verification scope The last step is to clarify the verification scope. At this stage, the mapping relationship established in the previous steps is used to determine which fields should be the target fields for quality verification. That is to say, once the association between the fields and the data standards is completed, specific verification rules can be defined according to these standards, and then quality checks can be performed on the corresponding fields.

[0117] For example, if a certain field is associated with a data standard that specifies non - null and a specific format, then this field becomes an object of quality verification, and the system will automatically generate corresponding verification rules to verify whether the field meets these quality requirements.

[0118] Through this systematic method, not only can data quality and consistency be ensured, but also the data quality management process can be simplified, the need for manual operations reduced, and work efficiency improved.

[0119] According to an embodiment of the present application, the predefined verification rule simulation is specifically as follows: Check whether the field is empty; Check whether the field value is within the enumeration range; Check whether the numerical value is within the specified interval; Verify the format through regular expressions; Verify whether the field value exists in the reference data table.

[0120] Specifically, the predefined verification rule simulation The predefined verification rules refer to a series of check logics preset according to the requirements of data quality, used to verify the integrity and accuracy of data. Specifically, it includes the following aspects: Check whether the field is empty Purpose: Ensure that certain key fields are not empty to maintain data integrity.

[0121] Implementation method: Detect whether a specific field has a value through SQL query or program logic. For example, for the "ID number" field in a student information table, if it is specified as non - null, any record with an empty value in this field will be regarded as not meeting the requirements.

[0122] Check whether the field value is within the enumeration range Purpose: Ensure that the field value belongs to a set of predefined valid values.

[0123] Implementation method: For fields with a fixed value range (such as the gender field), an enumeration list (for example, "male", "female") can be used to limit its possible values. The system will check whether the field value of each record is included in these predefined values.

[0124] Check whether the numerical value is within the specified interval Purpose: Ensure that the value of the numerical field falls within a reasonable interval.

[0125] Implementation method: For fields with clear numerical range limitations (such as the age field), set the minimum and maximum values. For example, if the age of a student should be between 6 and 18 years old, then the system needs to verify whether the age field of all records meets this condition.

[0126] Verify the format through regular expressions Purpose: Ensure that text fields follow specific format specifications.

[0127] Implementation method: Use regular expressions to verify the format of fields. For example, the ID number should conform to the standard format of Chinese ID numbers (18 digits or the last digit is X). Corresponding regular expressions can be written to match these format requirements and applied to the inspection of relevant fields.

[0128] Verify whether the field value exists in the reference data table Purpose: Confirm that the field value is consistent with an external authoritative data source.

[0129] Implementation method: For fields that depend on external data sources (such as the address field), it is necessary to verify whether its value exists in the specified reference data table. For example, the address field should be able to correspond to a valid administrative division dictionary. This usually involves cross-table queries or calling API interfaces to complete the verification.

[0130] Example illustration Suppose there is a student information table in a primary school, which contains the following fields and corresponding requirements: ID number (cannot be empty and conform to the format) Gender (can only be "male" or "female") Age (must be between 6 and 13 years old) Address (needs to exist in the administrative division dictionary) Based on the above predefined verification rules, the following operations can be performed: For the ID number field, first check whether it is empty, and then use regular expressions to verify whether its format is correct.

[0131] For the gender field, check whether its value is within the predefined enumeration range (i.e., "male" or "female").

[0132] For the age field, check whether its value is within the range of [6, 13].

[0133] For the address field, verify whether its value exists in the administrative division dictionary.

[0134] According to an embodiment of the present application, the quality verification rules are generated based on the predefined verification rule template, the data standard, and the mapping relationship, specifically: Extract the quality-related attributes from the data standard; According to the mapping relationship, clarify the database tables and fields to be verified; Match the attributes of the data standard with the predefined verification rule template, and bind them with mapping fields, verification logic, and verification parameters to generate quality verification rules.

[0135] Specifically, extract the quality-related attributes from the data standard First, it is necessary to extract those attributes directly related to data quality from the already defined data standard. These attributes typically include, but are not limited to: Value range constraint: For example, the valid value range or enumeration list of a certain field.

[0136] Format specification: Specific format requirements such as date format, phone number format, etc.

[0137] Nullability: Whether the field allows null values.

[0138] For example, in a student information table, the quality attributes of the "ID number" field may include the requirement of not allowing null values and conforming to a specific format (such as regular expression verification); while the quality attributes of the "gender" field may be that it can only take "male" or "female".

[0139] Based on the mapping relationship, clarify the database tables and fields to be verified Next, according to the pre-established mapping relationship, determine which database tables and fields need to be subject to quality verification. The mapping relationship refers to the process of associating the actual fields in the database with the corresponding data standards. This step clarifies: Which tables need to be verified.

[0140] Among these tables, which specific fields are the objects of verification.

[0141] For example, in the above example of the student information table, assuming that the four fields of "ID number", "gender", "age", and "address" respectively correspond to the corresponding data standards, then these fields are the target fields for which we need to conduct quality verification.

[0142] Match the attributes of the data standard with the predefined verification rule template, and bind them with mapping fields, verification logic, and verification parameters to generate quality verification rules In this step, combine the quality-related attributes in the data standard with the predefined verification rule template to generate quality verification rules for specific fields. The specific process is as follows: Select a suitable verification rule template: Select an appropriate verification rule template according to the quality attributes in the data standard. For example, if a certain field has a value range constraint, select a value range check template; if there is a format specification requirement, select a format verification template.

[0143] Bind mapping fields, verification logic, and parameters: For each selected field, bind it to the corresponding verification logic and parameters. This means specifying specific verification conditions for each field. For example: For the "ID number" field, use null value check and regular expression to verify the format.

[0144] For the "gender" field, use range check to ensure that its value is "male" or "female".

[0145] For the "age" field, use interval check to ensure that its value is between [6, 13].

[0146] For the "address" field, use data correlation check to ensure that its value exists in the administrative division dictionary.

[0147] Generate the final quality verification rules: Through the above steps, generate specific quality verification rules applicable to each field. These rules can be automatically applied to the corresponding fields in the database to ensure that the data meets the predetermined quality standards.

[0148] Example Suppose there is a student information table, which contains the following fields and their corresponding verification rules: ID number: Cannot be null, and it should conform to a specific format (such as ^\d{17}[\dXx]$) Gender: Cannot be null, and the value range is "male" or "female" Age: Cannot be null, and the value should be in the range of [6, 13] Address: Allows null, but if not null, it should exist in the administrative division dictionary Based on the above process, the system will automatically generate specific verification rules for each field and apply these rules during execution to ensure the quality of the data.

[0149] According to an embodiment of the present application, the matching logic includes: null value rule, value range rule, threshold rule, format specification rule, and data correlation rule.

[0150] Specifically, the null value rule (Null Value Rule) Definition: Used to check whether a field allows null. If a field is set to not allow null, any record containing a null value will be considered not meeting the quality requirements.

[0151] Application scenario: For example, in a student information table, the "ID number" field usually does not allow null because it is one of the unique identifiers for identifying students.

[0152] Value range rule (Domain Value Rule) Definition: Ensure that the field value belongs to a predefined set of valid values or an enumerated list. This can be a fixed list of values or defined through an external reference data table.

[0153] Application scenario: For example, if the "gender" field can only take values of "male" or "female", then the value range rule is needed to restrict the input values to within these two options.

[0154] Threshold Rule Definition: Applies to numeric fields and is used to check whether the field value falls between the specified minimum and maximum values.

[0155] Application scenario: For example, the "age" field should satisfy the range between 6 and 13 years old, that is, the minimum value is 6 and the maximum value is 13. All records outside this range will be regarded as anomalies.

[0156] Format Specification Rule Definition: Used to verify whether a text field conforms to specific format requirements. These format requirements can be defined through regular expressions.

[0157] Application scenario: For example, the "ID number" field needs to conform to the standard format of Chinese ID numbers (18 digits or the last digit is X), and the regular expression ^\d{17}[\dXx]$ can be used for verification.

[0158] Data Referential Integrity Rule Definition: Ensure that the field value exists in another reference data table to maintain data consistency and integrity.

[0159] Application scenario: Assume that the "address" field should correspond to a valid administrative division dictionary. This means that each value in this field must be able to find a corresponding entry in the administrative division dictionary.

[0160] Example Suppose there is a student information table in a primary school, and its fields and corresponding verification rules are as follows: ID number: Apply the null value rule and the format specification rule (such as ^\d{17}[\dXx]$).

[0161] Gender: Apply the null value rule and the value range rule (value range is "male", "female").

[0162] Age: Apply the null value rule and the threshold rule (interval is [6, 13]).

[0163] Address: Apply data correlation rules (the address must exist in the administrative division dictionary).

[0164] Through the above different types of verification rules, the system can comprehensively evaluate the quality of data in the database, ensuring that the data is both complete and accurate, thereby improving the overall quality and reliability of the data. This customized setting method of verification rules based on specific business requirements makes data quality management more flexible and efficient.

[0165] According to an embodiment of the present application, after executing the quality verification rules according to a preset cycle or event trigger, it further includes: Run the quality verification rules to identify data that does not meet the quality requirements as abnormal data; Record the abnormal data in the quality report and trigger an alarm notification; Support manual or automatic repair processes.

[0166] Specifically, run the quality verification rules to identify data that does not meet the quality requirements as abnormal data. Once the quality verification rules are triggered according to a preset cycle or specific event, the system will start to execute these rules to check whether the data in the database meets the predetermined quality standards. The specific steps are as follows: Execute verification: The system will check relevant fields one by one according to the preset verification rules (such as null value rules, value range rules, threshold rules, format specification rules, and data correlation rules).

[0167] Identify abnormalities: During the inspection process, if any record does not meet the corresponding quality standard, that record will be marked as abnormal data. For example, the ID number field is empty or has an incorrect format, the age exceeds the specified range, etc.

[0168] Record the abnormal data in the quality report and trigger an alarm notification After identifying the abnormal data, the next step is to record this information and notify the relevant personnel: Record in the quality report: All data marked as abnormal will be detailedly recorded in a quality report. This report usually contains information such as the specific location of the abnormal data (table name, field name), error description, and possible cause analysis.

[0169] Trigger an alarm notification: For quality problems with a higher severity level, the system will automatically send an alarm notification to the relevant management personnel or responsible team. The notification method can be diversified, such as sending warning messages via email, text message, or other instant messaging tools to ensure that the relevant personnel can be informed of the problem in a timely manner and take corresponding measures.

[0170] Support manual or automatic repair processes The last step is to fix the detected problems, and the system should provide flexible repair options: Manual repair: Allows users to directly view and edit abnormal data through the interface. This method is suitable for situations that require manual judgment or complex corrections. For example, some data errors may be caused by changes in business logic and require manual intervention to adjust.

[0171] Automatic repair: For some simple and clear errors, batch corrections can be made through automated scripts or programs. For example, if the format of a certain field does not meet the expectations but can be corrected through simple regular expression conversion, an automatic repair process can be set up to complete this task.

[0172] According to an embodiment of the present application, after executing the quality inspection rule according to a preset period or event trigger, it further includes: Monitoring update events of the data standard library and mapping relationships; Re-matching the template according to the change content and generating new rules to overwrite the old rules; Retaining historical rule versions to support rollback and auditing.

[0173] Specifically, monitoring update events of the data standard library and mapping relationships Real-time monitoring mechanism: The system needs to have a monitoring module for real-time detection of changes in the data standard library and the mapping relationship between fields and data standards. This can be achieved by listening to specific tables in the database or using an event-driven architecture.

[0174] Event trigger: When the data standard changes (such as adding new value range constraints, modifying format specifications) or the mapping relationship between fields and data standards changes, the event will be triggered and relevant processing logic will be notified.

[0175] Re-matching the template according to the change content and generating new rules to overwrite the old rules Change analysis: Once an update event is detected, the system will first analyze the specific change content. For example, the data type of a certain field changes from character type to integer type, or a not-null constraint is added to a certain field, etc.

[0176] Re-matching the template: Based on the analysis results, the system will re-select a suitable predefined inspection rule template according to the latest data standards and mapping relationships. For example, if a certain field now requires compliance with a specific format, the corresponding format verification template needs to be selected.

[0177] Generating new rules: Using the latest matched template, combined with the changed data standards and mapping relationships, automatically generate new quality inspection rules. These new rules will overwrite the original old rules to reflect the latest data quality management requirements.

[0178] Rule Deployment: The newly generated quality inspection rules will be deployed into the system and take effect immediately, ensuring that all subsequent data checks are carried out according to the latest rules.

[0179] Retain historical rule versions, support rollback and auditing Version Control: To ensure the stability and traceability of the system, each change to the quality inspection rules creates a new version. This means that the system retains all historical versions of the rules, not just the currently used version.

[0180] Rollback Function: If the new rules cause some unforeseen problems, the system should provide a convenient rollback mechanism that allows users to quickly switch back to any previous historical version. This helps to quickly restore the normal operation of the system and reduce the risks brought by rule changes.

[0181] Auditing Support: In addition to rollback, retaining historical versions also helps with auditing. By viewing the rule versions at different time points, the historical changes of data standards and mapping relationships can be traced, which is very useful for compliance reviews and problem troubleshooting. For example, when data quality problems occur, the root cause can be found by comparing the rule versions at different times.

[0182] A computer program product containing instructions, when running on a device, is characterized in that it causes the device to execute the steps in the method for automatically generating quality inspection rules based on data standards.

[0183] A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the method for automatically generating quality inspection rules based on data standards.

[0184] An electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the steps in the method for automatically generating quality inspection rules based on data standards.

[0185] What is not described in this application can be implemented by adopting or referring to the existing technologies.

[0186] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are to illustrate the differences from other embodiments.

[0187] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for automatically generating quality inspection rules based on data standards, characterized in that: include: Define data standards, match database fields with the data standards, determine the scope of verification, and generate mapping relationships; Generating quality check rules based on the predefined check rule template, the data standard and the mapping relationship; The quality check rules are executed according to a preset period or event trigger.

2. The method according to claim 1, characterized in that: The data standard is defined, the database fields are matched with the data standard, and the verification scope is determined to generate a mapping relationship, specifically: Dividing the attributes of the data standard into business attributes, technical attributes, management attributes and quality attributes; For each of the data standards, identify the quality-related attributes; Associating the database table fields with the corresponding data standards one-to-one or many-to-one; The verification scope is clarified, that is, the mapped fields are used as target fields for quality verification.

3. The method according to claim 1, characterized in that: The predefined verification rule simulation is specifically as follows: Check if the field is empty; Check if the field value is within the enumeration range; Check whether the value is within the specified range; Validate the format via regular expressions; The field value under validation exists in the reference table.

4. The method according to claim 1, characterized in that: The quality check rule is generated based on the predefined check rule template, the data standard and the mapping relationship, specifically: extracting quality-related attributes from the data standard; According to the mapping relationship, the database tables and fields that need to be checked are clearly identified; The attributes of the data standard are matched with the predefined check rule template, and are bound with the mapping fields, check logic, and check parameters to generate quality check rules.

5. The method according to claim 4, characterized in that The matching logic includes: null value rules, value range rules, threshold rules, format specification rules and data relevance rules.

6. The method according to claim 1, characterized in that After executing the quality check rule according to a preset period or event trigger, the method further includes: Running the quality check rules to identify data that does not meet the quality requirements as abnormal data; Record the abnormal data in the quality report and trigger an alarm notification; Supports manual or automated repair processes.

7. The method according to claim 1, characterized in that After executing the quality check rule according to a preset period or event trigger, the method further includes: Monitoring update events of the data standard library and mapping relationships; Re-match the template according to the changed content and generate new rules, overwriting the old rules; Keep historical rule versions and support rollback and auditing.

8. A computer program product comprising instructions, which, when executed on a device, is characterized in that: The device is enabled to execute the steps in the method for automatically generating quality check rules based on data standards as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the method for automatically generating quality check rules based on data standards as described in any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for automatically generating quality check rules based on data standards as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method for realizing data standard and data quality association processing based on metadata in big data governance

    CN110119395A

Cited By

  • Data quality detection method and device based on data standard

    CN121188028A

  • Metadata-driven cross-platform data access and fusion sharing method and system

    CN121705271A