Data quality control method and device, computer equipment and storage medium

Through timed scheduling tools and quality assessment rules, data quality assessment and repair tasks are automatically performed, solving the problem of high manual participation in existing technologies, realizing intelligent and automated management and control of data quality, and improving efficiency and accuracy.

CN120653636APending Publication Date: 2025-09-16CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510687406.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing data quality control solutions require a lot of manual participation, resulting in low data quality control efficiency and difficulty in achieving intelligence and automation.

Method used

Through timed scheduling tools and quality assessment rules, data quality assessment tasks and data repair tasks are automatically executed. Data repair strategies are intelligently determined based on quality assessment results, and data repair is performed under preset conditions, achieving intelligent and automated data quality control.

Benefits of technology

It improves the efficiency and accuracy of data quality control, realizes intelligent and automated evaluation and repair of data quality, and reduces manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653636A_ABST
    Figure CN120653636A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and particularly discloses a data quality control method and device, computer equipment and a storage medium. According to the method, the quality evaluation task and the data recovery task are automatically executed through the timed scheduling tool, and the data recovery strategy is intelligently determined according to the quality evaluation result, so that quality evaluation and data recovery are performed according to the corresponding target quality evaluation rule and the data recovery strategy; intelligent and automatic data quality evaluation and data restoration are realized, and then the data management and control efficiency is improved. When the method is applied to the data quality management and control service of the financial system, the intelligence and automation level of data quality management and control can be improved, and then the data quality management and control efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data quality control method, apparatus, computer equipment, and storage medium. Background Art

[0002] The rapid development of the financial industry and the widespread application of financial technology are driving the need to process and analyze massive amounts of data. Data from various financial businesses is characterized by large volumes, diverse sources, and complex data types. For example, data sources in the insurance industry include customer information systems, claims systems, and third-party data systems, and data types include structured, semi-structured, and unstructured data. However, existing data quality control solutions require significant manual effort during data preprocessing and data quality assessment. For example, the generation and triggering of preprocessing and quality assessment tasks require manual processing, reducing data quality control efficiency. Therefore, implementing intelligent and automated data quality control to improve data quality control efficiency has become a pressing issue. Summary of the Invention

[0003] The present application provides a data quality control method, apparatus, computer equipment and storage medium to achieve intelligent and automated data quality control and improve data quality control efficiency.

[0004] In a first aspect, the present application provides a data quality control method, the method comprising:

[0005] Obtaining the data to be controlled and the target quality assessment rules corresponding to the data to be controlled;

[0006] Based on a preset timing scheduling tool and the target quality assessment rule, a quality assessment task is executed regularly to obtain a quality assessment result of the data to be managed, wherein the quality assessment result includes a quality verification result;

[0007] When the quality check result is failure, determining a data repair strategy according to the quality assessment result, and generating a data repair task based on the data repair strategy;

[0008] When a preset data repair trigger condition is reached, the data repair task is executed regularly based on the timing scheduling tool and the data repair strategy, and the repaired data is stored in a preset data warehouse.

[0009] In a second aspect, the present application further provides a data quality control device, comprising:

[0010] A data acquisition module is used to acquire the data to be controlled and the target quality assessment rules corresponding to the data to be controlled;

[0011] A quality assessment module, configured to periodically execute quality assessment tasks based on a preset time scheduling tool and the target quality assessment rules, and obtain quality assessment results of the data to be managed, wherein the quality assessment results include quality verification results;

[0012] a strategy determination module, configured to determine a data repair strategy according to the quality assessment result when the quality check result is a check failure, and generate a data repair task based on the data repair strategy;

[0013] The data repair module is used to execute the data repair task regularly based on the timing scheduling tool and the data repair strategy when the preset data repair trigger condition is met, and store the repaired data in a preset data warehouse.

[0014] In a third aspect, the present application also provides a computer device, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the data quality control method as described above when executing the computer program.

[0015] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the data quality control method as described above.

[0016] The present application discloses a data quality control method, apparatus, computer equipment and storage medium, which obtains the data to be controlled and the target quality assessment rules corresponding to the data to be controlled; based on a preset timing scheduling tool and the target quality assessment rules, regularly executes the quality assessment task to obtain the quality assessment result of the data to be controlled, wherein the quality assessment result includes a quality verification result; when the quality verification result is a verification failure, determines a data repair strategy based on the quality assessment result, and generates a data repair task based on the data repair strategy; when the preset data repair trigger condition is reached, regularly executes the data repair task based on the timing scheduling tool and the data repair strategy, and stores the repaired data in a preset data warehouse. The present application automatically executes quality assessment tasks and data repair tasks through a timing scheduling tool, and intelligently determines a data repair strategy based on the quality assessment results, so as to perform quality assessment and data repair in accordance with the corresponding target quality assessment rules and data repair strategy, thereby realizing intelligent and automated data quality assessment and data repair, thereby improving data control efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 This is a schematic flow chart of a data quality control method provided in the first embodiment of the present application;

[0019] Figure 2 is a schematic flow chart of a data quality control method provided in the second embodiment of the present application;

[0020] Figure 3 This is a schematic flow chart of a data quality control method provided in the third embodiment of the present application;

[0021] Figure 4 A schematic block diagram of a data quality control device provided in an embodiment of the present application;

[0022] Figure 5 A schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0024] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0025] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0026] It should be further understood that the term “and / or” used in this specification and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0027] The embodiments of the present application provide a data quality control method, apparatus, computer device, and storage medium. The data quality control method can be applied to a server, which can be a standalone server or a server cluster.

[0028] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0029] See also Figure 1 , Figure 1 This is a schematic flow chart of a data quality control method provided in an embodiment of the present application. This data quality control method can be applied to a server to automatically execute quality assessment tasks and data repair tasks through a timed scheduling tool, and intelligently determine a data repair strategy based on the quality assessment results. Quality assessment and data repair are performed according to the corresponding target quality assessment rules and data repair strategy, thereby achieving intelligent and automated data quality assessment and data repair, thereby improving data control efficiency.

[0030] like Figure 1 As shown, the data quality control method specifically includes steps S101 to S104.

[0031] S101: Obtaining data to be controlled and target quality assessment rules corresponding to the data to be controlled;

[0032] In one embodiment, the data to be managed can be data that is being entered or acquired for the first time, such as newly generated and entered business data after completing an insurance claim. When data is first entered or acquired, it needs to be preprocessed, including steps such as data cleaning, data conversion, and data standardization. This preprocessed data is then used as the data to be managed.

[0033] In another embodiment, the data to be managed may also be data stored in a data warehouse.

[0034] Furthermore, the obtaining of the data to be controlled includes: collecting data in real time based on a preset data capture and transmission tool; and preprocessing the data based on a preset data preprocessing tool to obtain the data to be controlled.

[0035] In one embodiment, data preprocessing efficiency is low due to the diversity of data sources and data structures. For example, data in the insurance industry comes from a wide range of sources, including customer information, claims records, and third-party data. The data categories are complex, involving structured, semi-structured, and unstructured data. Data is scattered across different systems and channels, and data integration and consolidation face challenges with heterogeneous data sources. Existing technologies struggle to achieve efficient integration and unified management when processing multi-source heterogeneous data, making it difficult to ensure data consistency. Furthermore, they lack intelligent and automated tools when processing unstructured data, resulting in low cleaning and standardization efficiency.

[0036] In order to improve the efficiency of data preprocessing, the embodiment of the present application adopts the distributed computing capabilities of the distributed computing framework to achieve efficient preprocessing of massive data. Data preprocessing includes steps such as data cleaning, data conversion and data standardization.

[0037] Specifically, data capture and transmission tools are used to collect structured or unstructured data from various sources into real-time stream processing tools for preprocessing. The preprocessed data is then stored in a pre-set data warehouse as data to be managed.

[0038] Furthermore, obtaining the target quality assessment rules corresponding to the data to be controlled includes: monitoring data changes of the data to be controlled, and obtaining a data quality list based on the data changes; adjusting the preset quality assessment rules based on the data quality list to obtain target quality assessment rules.

[0039] In one embodiment, due to the dynamic changes in the data environment, business needs are constantly updated. Existing technologies lack flexibility in dynamically adjusting data quality assessment rules and models, making it difficult to monitor and quickly respond to data quality issues in real time. Therefore, this embodiment improves the efficiency and accuracy of data quality control by monitoring data changes and dynamically adjusting quality assessment rules. Among them, the dynamic changes in the data environment refer to the diversity of data sources. With the advancement of the times, insurance industry data is exposed to more non-traditional data such as unstructured data such as social media; sudden events have caused a large number of insurance claims, and the amount of data has increased dramatically.

[0040] The data quality checklist includes quality assessment dimensions and standard indicators corresponding to each quality assessment dimension.

[0041] Specifically, real-time stream processing tools are used to read metadata information in the data stream, and combined with the data lineage relationship map, changes in the data to be controlled are analyzed in real time. Based on the details of the data changes, the quality assessment dimensions of the current data to be controlled and the specified indicators for each dimension are analyzed. The obtained quality assessment dimensions and indicators are compared with the preset quality assessment rules, and the quality assessment dimensions and indicators in the preset quality assessment rules are adjusted to the currently obtained quality assessment dimensions and standard indicators to obtain the target quality assessment rules. For example, in the insurance claims business, claims data is monitored in real time. When there is an increase in data, the quality assessment rules need to be adjusted according to the specific circumstances of the newly added data to obtain the target quality assessment rules.

[0042] In the above embodiment, by combining metadata management and data lineage tracking, a data quality list can be provided for the dynamic adjustment of quality assessment rules, and the quality assessment rules can be dynamically adjusted according to the data quality list, thereby improving the efficiency and accuracy of data quality control.

[0043] S102: Based on a preset time scheduling tool and the target quality assessment rule, regularly execute a quality assessment task to obtain a quality assessment result of the data to be managed, wherein the quality assessment result includes a quality verification result;

[0044] In one embodiment, a scheduled scheduling tool (such as the Linkdo scheduling tool) can trigger data quality assessment tasks at preset time intervals (e.g., 2:00 AM daily) or according to the triggering conditions of other data quality assessment tasks. This ensures that data quality verification work is carried out in accordance with specified requirements and that no important data quality inspection time points are missed.

[0045] In one embodiment, the data quality assessment task includes quality verification, and when the quality verification result is verification failure, it also includes anomaly detection.

[0046] The data quality assessment task is triggered and executed through the timing scheduling tool, and the specified data to be controlled is verified according to the target quality assessment rules. If the data to be controlled passes the verification, the quality assessment result is directly obtained; if the data to be controlled fails the verification, anomaly detection is performed on the abnormal data to obtain anomaly detection results.

[0047] It can be understood that when the quality verification result is passed, the quality assessment result only includes the quality verification result, that is, passed; when the quality verification result is failed, the quality assessment result also includes the abnormality detection result, that is, failed, abnormal problem, and abnormal level.

[0048] S103: When the quality check result is failure, determining a data repair strategy according to the quality assessment result, and generating a data repair task based on the data repair strategy;

[0049] In one embodiment, when the verification fails, the quality assessment result includes abnormal problems and abnormal levels, and a data repair strategy is determined based on the abnormal problems and abnormal levels.

[0050] Specifically, the repair method corresponding to the abnormal problem and the repair priority corresponding to the abnormal level can be determined by looking up a table (including a list of abnormal problem repair methods and a list of abnormal level repair priorities). The abnormal problem repair method list and the abnormal level repair priority list can be freely set by the user according to actual conditions.

[0051] In one embodiment, specific data repair tasks are created based on the data repair strategy, clearly defining the repair objectives, scope, steps, and repair priority of the data repair task. The execution order of the data repair tasks is arranged based on the repair priority. High-priority tasks are executed first, while low-priority repair tasks can be executed in batches at a preset time.

[0052] S104: When a preset data repair trigger condition is met, the data repair task is executed regularly based on the timing scheduling tool and the data repair strategy, and the repaired data is stored in a preset data warehouse.

[0053] In one embodiment, a timer scheduling tool is used to execute data repair tasks when preset data repair trigger conditions are met. Data repair trigger conditions can include reaching a preset time, such as executing the repair task at 2:00 a.m. every day to avoid affecting system performance during peak business hours; triggering the repair task and executing it in batches when the accumulated data to be repaired reaches a certain amount (such as 1,000 records); and requiring immediate execution of the data repair task, such as repairing abnormalities in key field data.

[0054] For automated repair tasks, the repair tool cleans, transforms, and updates data according to pre-set rules. For example, it can automatically fill in missing values ​​or correct incorrect date formats. For tasks requiring manual intervention, the relevant personnel are notified for review and correction.

[0055] In one embodiment, repaired data is stored in a data warehouse, where the data's repair history and lineage are recorded. Information such as the repair operations performed on the data, who repaired it, and when it was repaired is noted, facilitating data traceability and auditing. For example, the data warehouse can include one built on Hive, supporting the storage and query of structured and semi-structured data, enhancing data integration and management capabilities.

[0056] The above embodiment provides a data quality control method, apparatus, computer equipment and storage medium, which obtains the data to be controlled and the target quality assessment rules corresponding to the data to be controlled; based on the preset timing scheduling tool and the target quality assessment rules, the quality assessment task is executed regularly to obtain the quality assessment result of the data to be controlled, wherein the quality assessment result includes the quality verification result; when the quality verification result is that the verification fails, the data repair strategy is determined according to the quality assessment result, and a data repair task is generated based on the data repair strategy; when the preset data repair trigger condition is reached, the data repair task is executed regularly based on the timing scheduling tool and the data repair strategy, and the repaired data is stored in a preset data warehouse. The present application automatically executes quality assessment tasks and data repair tasks through a timing scheduling tool, and intelligently determines the data repair strategy based on the quality assessment results, so as to perform quality assessment and data repair in accordance with the corresponding target quality assessment rules and data repair strategy, thereby realizing intelligent and automated data quality assessment and data repair, thereby improving data control efficiency.

[0057] See also Figure 2 , Figure 2 This is a schematic flow chart of a data quality control method provided in an embodiment of the present application. This data quality control method can be applied to a server and uses technologies and methods such as timed scheduling tools, quality assessment rules, and anomaly detection models to achieve automated comprehensive monitoring and assessment of data quality.

[0058] like Figure 2 As shown, step S102 of the data quality control method specifically includes steps S201 to S203.

[0059] S201: Performing quality verification on the data to be controlled based on the timing scheduling tool and the target quality assessment rule to obtain the quality verification result;

[0060] In one embodiment, a scheduled scheduling tool (such as the Linkdo scheduling tool) can trigger data quality assessment tasks at preset time intervals (e.g., 2:00 AM daily) or according to the triggering conditions of other data quality assessment tasks. This ensures that data quality verification work is carried out in accordance with specified requirements and that no important data quality inspection time points are missed.

[0061] In one embodiment, the data quality assessment task includes quality verification, and when the quality verification result is verification failure, it also includes anomaly detection.

[0062] In one embodiment, quality assessment rules include but are not limited to data integrity verification rules (such as whether key fields are missing), data accuracy verification rules (such as whether the data is within a reasonable range), data consistency verification rules (such as whether the data in different systems are consistent), etc.

[0063] In one embodiment, a data quality assessment task is triggered and executed through a time scheduling tool, and a verification operation is performed on the designated data to be managed according to target quality assessment rules.

[0064] Each controlled data is verified against the target quality assessment rules. For example, the "Claim Amount" field in insurance claim data is checked to see if it falls within the policy's coverage range, or the "ID Number" field in customer information is checked to see if it complies with the ID number format requirements. During the quality verification process, the verification results for each record are recorded, indicating whether the record has passed the quality verification.

[0065] After completing the verification of all data to be controlled, the verification results are counted and summarized. The statistical content includes but is not limited to the number of data records that passed the verification, the number of data records that failed the verification, and the distribution of various data quality issues.

[0066] The quality verification results are stored in a designated storage location (such as a database or file system) for subsequent query and analysis. The quality verification results can also be displayed through the data quality control platform's visual interface, allowing business and system managers to intuitively understand the data quality status. Display content includes but is not limited to data quality trend charts and pie charts of the proportion of various data quality issues.

[0067] S202: When the quality check result is failure, the data to be controlled is marked as data to be repaired, and a preset anomaly detection model is called to perform anomaly detection on the data to be repaired to obtain anomaly problems and anomaly levels of the data to be repaired;

[0068] In one embodiment, when the quality check result shows that the data has failed the check, the data record that has failed the check is automatically marked as "data to be repaired" to distinguish the abnormal data from other normal data, facilitating subsequent repair processing.

[0069] Separate the data to be repaired from the original data set and store it in a temporary storage area. This temporary storage area can be an independent database table, folder, or data partition, which is used to centrally store all data records to be repaired. This prevents the repaired data from interfering with normal data and facilitates centralized management and processing of the repaired data.

[0070] In one embodiment, an anomaly detection model (such as a clustering algorithm) can be trained based on historical data to identify anomalies in the data.

[0071] The data to be repaired is transferred to the anomaly detection model for analysis. Based on its internal algorithms and rules, the anomaly detection model conducts in-depth anomaly detection on the data. During this process, the model may detect outliers, calculate feature vectors of data records, and compare them with normal data patterns to identify potential anomalies. For example, the model may discover that a claim amount is too high or that the time of the claim is unusual.

[0072] The anomaly detection model will output specific anomalies in the data to be repaired, including but not limited to data accuracy, completeness, consistency, and other aspects. For example, the model may identify anomalies such as "the claim amount exceeds the policy limit" or "the name in the customer information does not match the ID number."

[0073] In one embodiment, the abnormality level of the abnormal data is determined according to a preset abnormality level determination rule. Exemplarily, the abnormality level of each abnormal problem is determined based on the severity of the abnormal problem and the impact on the business. The abnormality level can be divided into three levels: low, medium, and high. For example, "the claim amount exceeds the policy limit" and the excess amount is greater than the threshold, which is determined to be a high-level abnormality; "the contact number format in the customer information is incorrect" is determined to be a low-level abnormality; when a certain data has more than a preset number of low-level abnormalities, it is also determined to be a high-level abnormality. Among them, the abnormality level determination rules can be freely set by the user according to actual needs.

[0074] In another embodiment, the abnormality level determination rule may further include performing trend analysis on the data based on multiple consecutive quality assessment results to display the changes in data quality over time. If the data abnormality trend increases, the data level is determined to be a high-level abnormality.

[0075] S203: Generate the quality assessment result based on the quality verification result, the abnormal problem and the abnormality level.

[0076] In one embodiment, relevant information such as quality verification results, anomaly issues, and anomaly levels are collected. Quality verification results provide an overview of the overall data quality, including the number of data records that passed and failed verification. Anomaly issues detail specific quality issues within the data, while anomaly levels reflect the severity of the data anomaly. Anomaly issues and anomaly levels are collated and summarized to generate a comprehensive quality assessment result.

[0077] It can be understood that when the quality verification result is passed, the quality assessment result only includes the quality verification result, that is, passed; when the quality verification result is failed, the quality assessment result also includes the abnormality detection result, that is, failed, abnormal problem, and abnormal level.

[0078] In one embodiment, the quality assessment report can be in a variety of formats, such as text reports, table reports, graphical reports, etc., so that the report content is concise and clear, and the key points are highlighted, which makes it easier for business personnel and responsible personnel to quickly understand the data quality status and take corresponding measures.

[0079] In the above embodiments, automated comprehensive monitoring and evaluation of data quality are achieved by utilizing technologies and means such as timing scheduling tools, quality assessment rules, and anomaly detection models.

[0080] See also Figure 3 , Figure 3 This is a schematic flow chart of a data quality control method provided in an embodiment of the present application. This data quality control method can be applied to a server to enable users to flexibly access data through a data quality control platform, achieve cross-departmental collaboration, promote data sharing and process optimization, and improve the utilization of data resources.

[0081] like Figure 3 As shown, the data quality control method specifically includes steps S301 to S304.

[0082] S301. Based on the preset data quality control platform, obtain the data responsibility list and receive user access requests;

[0083] In one embodiment, the data quality control platform supports cross-departmental collaboration and is used to manage various data quality-related tasks, including data responsibility list management, user access control, and data lineage tracking. It provides a unified operating interface for data managers and users to ensure data security, accuracy, and reliability.

[0084] By adding data lineage relationships through the form of database and table responsibility, the responsibility of the corresponding table is consolidated to the relevant business part, and data query permissions are opened to ensure that business problems of one party can be traced back to the relevant data information. For example, the insurance policy underwriting location belongs to the underwriting module, but when a customer makes a claim, it can be traced back to the insurance policy underwriting location, ensuring data sharing but controllable permissions.

[0085] In one embodiment, the data responsibility list clearly defines the responsible departments and individuals for each data table and data field. This list can be read from a backend database or configuration file, so that when a user requests data, they can clearly understand who is responsible for each piece of data. For example, in the insurance industry, the data responsibility list might record information such as "The customer information table is under the responsibility of the customer service department, and the responsible individual is Zhang San."

[0086] S302: Parse the user access request to obtain user information and request data;

[0087] In one embodiment, after the platform receives a user access request, it uses parsing technology to decompose and extract the request content to obtain user information and request data requested by the user.

[0088] User information typically includes an identifier (such as username, employee ID), department, and role, which can be used to determine subsequent permissions. For example, the user's identity information can be parsed from the authentication token in the request header to determine whether the user is an internal employee or external partner of the insurance company, as well as their position and department within the company. The request data specifies the specific data resources the user wishes to access.

[0089] In another embodiment, the user access request may further include the user's intention, such as whether it is a read-only query or data modification, so as to accurately verify whether the user's intention is permitted during permission verification.

[0090] S303: Determine user access rights based on the user information and the data responsibility list, and verify the user access rights and the requested data;

[0091] In one embodiment, the user's access rights are determined based on the user information and the data responsibility list.

[0092] Compare user information to the data responsibility list to determine whether the user belongs to the department responsible for a data resource or has authorization from the responsible person. For example, if the user is a claims specialist, they may only be able to access data related to the claims case they are responsible for. If the data the user requests access belongs to a specific department, verify whether the user is a member of that department or has permission to access data across departments.

[0093] After determining the user's access rights, the platform verifies the user's rights against the requested data. This checks whether the user has access to the requested data, including read, write, and modify permissions. For example, if a user requests to read a table of basic customer information, the platform verifies whether their permissions include read access to that table.

[0094] In one embodiment, the requested data itself is verified, including but not limited to checking whether the data exists, is accessible, and complies with data security and privacy requirements. For example, certain sensitive data may be inaccessible outside of a specific time window or require desensitization before being provided to the user.

[0095] S304: When the authority verification is passed, the requested data is displayed to the user, and the data lineage relationship is recorded for data tracking.

[0096] In one embodiment, once user access rights are verified, the platform converts the requested data into an appropriate format based on the user's request and preferences and displays it to the user. For example, if the user requests to view data in a table format, the platform converts the query result set into a table format, including setting the table headers and arranging the rows and columns. For data analysts in the insurance industry, the distribution and trends of claims data may be displayed in intuitive statistical charts to facilitate business analysis.

[0097] In one embodiment, while providing data, the data lineage is recorded. This includes information such as the data's source, the processing steps it has gone through, user access information, and the data flow path. For example, when a user accesses a cleansed and converted claims data report, the source data table, the data transmission path, and the current user's access operation are recorded.

[0098] Data lineage records can be used to trace the source of data quality issues, analyze their impact, and conduct compliance audits. When data errors are discovered or require updates, lineage relationships can be used to quickly locate affected users and business processes. For example, if errors occur in insurance product data, lineage relationships can be used to track the department and personnel using that data for pricing analysis, allowing timely notification and action to be taken.

[0099] In the above embodiments, the data quality control platform enables users to flexibly access data, achieve cross-departmental collaboration, promote data sharing and process optimization, and improve the utilization of data resources.

[0100] Furthermore, when the permission verification passes, the request data is displayed to the user, including: when the permission verification passes, obtaining the data category of the request data; determining the data encryption rules of the request data based on the data category and the user access rights; encrypting the request data based on the data encryption rules, and displaying the encrypted request data to the user.

[0101] In one embodiment, after receiving a user request and passing permission verification, the system determines the data category of the requested data. Data is categorized based on business needs and sensitivity, such as public data, internal data, and confidential data. Based on data content, data can be categorized as name data, age data, date data, and so on. Different encryption rules are used for different data categories. For example, confidential data may require a higher-level encryption algorithm.

[0102] In one embodiment, encryption rules are determined based on the identified data category and the user's access rights. For example, for data containing sensitive personal information (such as ID numbers and bank account numbers), a strong encryption algorithm (such as the SM3 encryption algorithm) should be used regardless of the user's permissions. For general business data, an appropriate encryption method can be selected based on the user's permission level. Users with lower permissions can only see partially encrypted or desensitized data.

[0103] For example, when customer information is displayed externally, it is desensitized and encrypted:

[0104] 1. Take the first Chinese character of the individual's name (if the name is in a foreign language, take the first three bytes of its UTF-8 (8-bit Unicode Transformation Format, a variable-length character encoding) encoding), followed by the ID number (English characters in the ID number are converted to uppercase), to form a string (UTF-8 encoding, taking the resident ID number as an example, is 21 bytes);

[0105] 2. Take the SM3 hash value of the above string, which is a 64-byte string (in lowercase). SM3 is the cryptographic hash algorithm defined in GB / T32905-2016 Information Security Technology SM3 Cryptographic Hash Algorithm;

[0106] 3. Take the first 6 bytes of the UTF-8 encoding of the ID number, followed by the above SM3 hash value, to obtain a 70-byte string, which is the final desensitized value.

[0107] In one embodiment, the requested data is encrypted using a selected encryption rule, and the encrypted data is displayed to the user.

[0108] In another embodiment, during the data processing and storage process, corresponding encryption rules are also used to encrypt the data to ensure the security of the data, prevent the data from being stolen or tampered with during the transmission and storage process, and protect data privacy.

[0109] In the above embodiment, data encryption rules are flexibly determined according to data categories and user access rights to avoid data leakage and improve data security.

[0110] Furthermore, after step S304, it also includes: obtaining the business processing flow corresponding to the request data; based on the business processing flow and the data lineage relationship of the request data, verifying the data flow path of the request data to obtain a data path verification result; when the data path verification result is that the data flow path does not comply with the business processing flow, determining that the user has violated the rules, and generating a violation alarm based on the violation.

[0111] In one embodiment, the business processing flow includes data processing flows in different business scenarios. For example, the insurance claim settlement business process includes reporting, investigation, damage assessment, claim verification, and payment.

[0112] Based on the data lineage, all the flow links of the requested data from its generation to its current state are restored, including the data's source system, the processing steps it has gone through, which business links have used it, and other information to obtain the data flow path.

[0113] Compare the restored data flow path with the business processing flow. Check whether the data flows according to the order and rules specified by the process. For example, in insurance claims processing, whether the claim data enters the claim process only after completing the pre-process steps such as investigation, loss assessment, and claim verification.

[0114] If the data flow path does not conform to the business process, for example, the compensation data directly enters the compensation operation without going through the compensation verification stage, the platform will determine that the user may have violated the regulations.

[0115] The alert mechanism is triggered, generating a violation alert. The alert should describe in detail the nature of the violation, the data involved, and the steps involved. For example, "User A, when processing claim B, skipped the claim verification step and proceeded directly with the payment, posing a serious violation risk."

[0116] Violation alerts will be notified to relevant management personnel via email, SMS or in-site messages, and detailed information of violation alerts will be recorded in the platform, including alert time, violating user, involved data, processing status, etc., for auditing and processing.

[0117] In the above embodiment, the data flow path can be monitored through data lineage and compared with the business process to determine whether there is any violation, effectively monitor whether the data usage and user behavior are standardized, and improve the rationality of data access and data quality.

[0118] See also Figure 4 , Figure 4 1 is a schematic block diagram of a data quality control device provided in an embodiment of the present application, wherein the data quality control device is used to execute the aforementioned data quality control method.

[0119] like Figure 4 As shown, the data quality control device 400 includes:

[0120] The data acquisition module 401 is used to acquire the data to be controlled and the target quality assessment rules corresponding to the data to be controlled;

[0121] The quality assessment module 402 is configured to periodically execute a quality assessment task based on a preset time scheduling tool and the target quality assessment rule, and obtain a quality assessment result of the data to be managed, wherein the quality assessment result includes a quality verification result;

[0122] A strategy determination module 403 is configured to determine a data repair strategy according to the quality assessment result when the quality check result is failure, and generate a data repair task based on the data repair strategy;

[0123] The data repair module 404 is configured to execute the data repair task on a regular basis based on the timing scheduling tool and the data repair strategy when a preset data repair trigger condition is met, and store the repaired data in a preset data warehouse.

[0124] Furthermore, the data acquisition module 401 includes:

[0125] A data quality list obtaining unit, configured to monitor data changes of the data to be controlled and obtain a data quality list according to the data changes;

[0126] The quality assessment rule adjustment unit is used to adjust the preset quality assessment rules according to the data quality list to obtain target quality assessment rules.

[0127] Furthermore, the quality assessment module 402 includes:

[0128] A quality verification unit, configured to perform quality verification on the data to be controlled based on the timing scheduling tool and the target quality assessment rule, and obtain the quality verification result;

[0129] an anomaly monitoring unit, configured to, when the quality verification result is failure, mark the data to be controlled as data to be repaired, and call a preset anomaly detection model to perform anomaly detection on the data to be repaired, to obtain an anomaly problem and an anomaly level of the data to be repaired;

[0130] A result generating unit is configured to generate the quality assessment result based on the quality verification result, the abnormal problem and the abnormality level.

[0131] Furthermore, the data quality control device 400 further includes a data access control module, which includes:

[0132] A request receiving unit is used to obtain a data responsibility list based on a preset data quality control platform and receive user access requests;

[0133] A request parsing unit, configured to parse the user access request to obtain user information and request data;

[0134] an authority verification unit, configured to determine user access rights based on the user information and the data responsibility list, and verify the user access rights and the requested data;

[0135] The data display unit is used to display the requested data to the user when the permission verification is passed, and record the data lineage relationship for data tracking.

[0136] Furthermore, the data display unit includes:

[0137] A data category acquisition subunit, configured to acquire the data category of the requested data when the permission check passes;

[0138] an encryption rule determination subunit, configured to determine a data encryption rule for the requested data according to the data category and the user access rights;

[0139] The data encryption display subunit is used to encrypt the request data based on the data encryption rule and display the encrypted request data to the user.

[0140] Furthermore, the data quality control device 400 further includes a violation checking module, and the violation checking module includes:

[0141] A business processing flow acquisition unit, configured to acquire the business processing flow corresponding to the request data;

[0142] a path verification result obtaining unit, configured to verify the data flow path of the request data based on the business processing flow and the data lineage relationship of the request data, and obtain a data path verification result;

[0143] The violation alarm generating unit is used to determine that the user has violated the rules when the data path verification result shows that the data flow path does not comply with the business processing flow, and generate a violation alarm based on the violation.

[0144] Furthermore, the data acquisition module 401 includes:

[0145] A data acquisition unit, used to collect data in real time based on pre-set data capture and transmission tools;

[0146] The data preprocessing unit is used to preprocess the data based on a preset data preprocessing tool to obtain the data to be controlled.

[0147] It should be noted that those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0148] The above-mentioned device can be realized in the form of a computer program. The computer program can be used in Figure 5 Runs on the computer device shown.

[0149] See also Figure 5 , Figure 5 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device may be a server.

[0150] See Figure 5 The computer device includes a processor, a memory, and a network interface connected through a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.

[0151] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which, when executed, can cause a processor to perform any data quality control method.

[0152] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0153] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any data quality control method.

[0154] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0155] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0156] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0157] Obtaining the data to be controlled and the target quality assessment rules corresponding to the data to be controlled;

[0158] Based on a preset timing scheduling tool and the target quality assessment rule, a quality assessment task is executed regularly to obtain a quality assessment result of the data to be managed, wherein the quality assessment result includes a quality verification result;

[0159] When the quality check result is failure, determining a data repair strategy according to the quality assessment result, and generating a data repair task based on the data repair strategy;

[0160] When a preset data repair trigger condition is reached, the data repair task is executed regularly based on the timing scheduling tool and the data repair strategy, and the repaired data is stored in a preset data warehouse.

[0161] In one embodiment, when obtaining the target quality assessment rule corresponding to the data to be managed, the processor is configured to implement:

[0162] Monitor data changes of the data to be controlled, and obtain a data quality list based on the data changes;

[0163] According to the data quality list, the preset quality assessment rules are adjusted to obtain the target quality assessment rules.

[0164] In one embodiment, the processor, when implementing a preset timing scheduling tool and the target quality assessment rule, regularly executing a quality assessment task and obtaining a quality assessment result of the data to be managed, wherein the quality assessment result includes a quality verification result, is configured to implement:

[0165] Based on the timing scheduling tool and the target quality assessment rule, performing quality verification on the data to be controlled to obtain the quality verification result;

[0166] When the quality check result is that the check fails, the data to be controlled is marked as data to be repaired, and a preset anomaly detection model is called to perform anomaly detection on the data to be repaired to obtain the anomaly problem and anomaly level of the data to be repaired;

[0167] The quality assessment result is generated based on the quality verification result, the abnormal problem and the abnormality level.

[0168] In one embodiment, the processor, after executing the data repair task on a scheduled basis based on the scheduled scheduling tool and the data repair strategy when a preset data repair trigger condition is met and storing the repaired data in a preset data warehouse, is further configured to:

[0169] Based on the preset data quality control platform, obtain the data responsibility list and receive user access requests;

[0170] Parsing the user access request to obtain user information and request data;

[0171] Determine user access rights based on the user information and the data responsibility list, and verify the user access rights and the requested data;

[0172] When the permission check is passed, the requested data is displayed to the user and the data lineage relationship is recorded for data tracking.

[0173] In one embodiment, when the processor displays the request data to the user when the permission check passes, it is configured to implement:

[0174] When the permission check passes, obtaining the data category of the requested data;

[0175] Determining a data encryption rule for the requested data based on the data category and the user access rights;

[0176] The request data is encrypted based on the data encryption rule, and the encrypted request data is displayed to the user.

[0177] In one embodiment, after the processor displays the requested data to the user and records the data lineage relationship for data tracking when the permission check passes, it is further configured to:

[0178] Obtaining the business processing flow corresponding to the request data;

[0179] Based on the business processing flow and the data lineage relationship of the request data, verify the data flow path of the request data to obtain a data path verification result;

[0180] When the data path verification result is that the data flow path does not comply with the business processing flow, it is determined that the user has violated the rules, and a violation alarm is generated based on the violation.

[0181] In one embodiment, when obtaining the data to be managed and controlled, the processor is configured to implement:

[0182] Real-time data collection based on pre-set data capture and transmission tools;

[0183] Based on a preset data preprocessing tool, the data is preprocessed to obtain the data to be controlled.

[0184] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement any data quality control method provided in the embodiment of the present application.

[0185] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., equipped on the computer device.

[0186] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A data quality control method, characterized in that: include: Obtaining the data to be controlled and the target quality assessment rules corresponding to the data to be controlled; Based on a preset timing scheduling tool and the target quality assessment rule, a quality assessment task is executed regularly to obtain a quality assessment result of the data to be managed, wherein the quality assessment result includes a quality verification result; When the quality check result is failure, determining a data repair strategy according to the quality assessment result, and generating a data repair task based on the data repair strategy; When a preset data repair trigger condition is reached, the data repair task is executed regularly based on the timing scheduling tool and the data repair strategy, and the repaired data is stored in a preset data warehouse.

2. The data quality control method according to claim 1, characterized in that: The obtaining of target quality assessment rules corresponding to the data to be controlled includes: Monitor data changes of the data to be controlled, and obtain a data quality list based on the data changes; According to the data quality list, the preset quality assessment rules are adjusted to obtain the target quality assessment rules.

3. The data quality control method according to claim 1, characterized in that: The quality assessment task is executed regularly based on the preset timing scheduling tool and the target quality assessment rule to obtain the quality assessment result of the data to be managed, wherein the quality assessment result includes a quality verification result, including: Based on the timing scheduling tool and the target quality assessment rule, performing quality verification on the data to be controlled to obtain the quality verification result; When the quality check result is that the check fails, the data to be controlled is marked as data to be repaired, and a preset anomaly detection model is called to perform anomaly detection on the data to be repaired to obtain the anomaly problem and anomaly level of the data to be repaired; The quality assessment result is generated based on the quality verification result, the abnormal problem and the abnormality level.

4. The data quality control method according to claim 1, characterized in that: When a preset data repair trigger condition is met, the data repair task is executed regularly based on the timing scheduling tool and the data repair strategy, and the repaired data is stored in a preset data warehouse, further comprising: Based on the preset data quality control platform, obtain the data responsibility list and receive user access requests; Parsing the user access request to obtain user information and request data; Determine user access rights based on the user information and the data responsibility list, and verify the user access rights and the requested data; When the permission check is passed, the requested data is displayed to the user and the data lineage relationship is recorded for data tracking.

5. The data quality control method according to claim 4, characterized in that: When the permission check passes, the request data is displayed to the user, including: When the permission check passes, obtaining the data category of the requested data; Determining a data encryption rule for the requested data based on the data category and the user access rights; The request data is encrypted based on the data encryption rule, and the encrypted request data is displayed to the user.

6. The data quality control method according to claim 4, characterized in that: When the permission check is passed, the requested data is displayed to the user and the data lineage relationship is recorded for data tracking, and the following is also included: Obtaining the business processing flow corresponding to the request data; Based on the business processing flow and the data lineage relationship of the request data, verify the data flow path of the request data to obtain a data path verification result; When the data path verification result is that the data flow path does not comply with the business processing flow, it is determined that the user has violated the rules, and a violation alarm is generated based on the violation.

7. The data quality control method according to any one of claims 1 to 6, characterized in that: The obtaining of the data to be controlled includes: Real-time data collection based on pre-set data capture and transmission tools; Based on a preset data preprocessing tool, the data is preprocessed to obtain the data to be controlled.

8. A data quality control device, characterized in that: include: A data acquisition module is used to acquire the data to be controlled and the target quality assessment rules corresponding to the data to be controlled; A quality assessment module, configured to periodically execute quality assessment tasks based on a preset time scheduling tool and the target quality assessment rules, and obtain quality assessment results of the data to be managed, wherein the quality assessment results include quality verification results; a strategy determination module, configured to determine a data repair strategy according to the quality assessment result when the quality check result is a check failure, and generate a data repair task based on the data repair strategy; The data repair module is used to execute the data repair task regularly based on the timing scheduling tool and the data repair strategy when the preset data repair trigger condition is met, and store the repaired data in a preset data warehouse.

9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the data quality control method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the data quality control method according to any one of claims 1 to 7.