Data cleaning system and method

By designing a data cleaning system that includes source data module, screening rule module, screening result module and audit authorization module, the problem that the existing technology cannot achieve cross-system data consistency, and unified analysis and cleaning of cross-system data is realized, and data quality and consistency are improved.

CN120013656APending Publication Date: 2025-05-16LONGYING ZHIDA (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411379264.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing data cleaning technology is mainly aimed at a single system, and cannot achieve consistency of cross-system data, resulting in low data quality and inability to meet the requirements of information digitization and system standardization.

Method used

A data cleaning system was designed, including source data module, screening rule module, screening result module and audit authorization module. Through unified data extraction, screening rules and approval processes, unified analysis and cleaning of cross-system data is realized.

Benefits of technology

It realizes unified analysis and cleaning of cross-system data, ensures that the data complies with business and regulatory requirements, improves data quality and consistency, and supports the construction of information digitalization and system standardization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013656A_ABST
    Figure CN120013656A_ABST
Patent Text Reader

Abstract

The invention discloses a data cleaning system and method, and the system comprises a source data module which is used for defining a source data extraction time interval and carrying out data extraction; the screening rule module is used for defining a screening rule; the screening result module is used for supporting access to a screening result; and the auditing and authorizing module is used for supporting process approval and countersigning. According to the method and the system, the data of each system are ensured to meet the current business and supervision requirements, the consistency of the same information in different systems is also ensured, and the data quality is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data cleaning, and in particular relates to a data cleaning system and method. Background Art

[0002] With the development of society, banks are conducting more and more diverse businesses, and corresponding systems are being built to support more and more businesses. However, due to the different construction periods and business focuses of various systems, data differences in the same system at different times and differences in the same information between different systems have led to poor data quality over the years.

[0003] Existing data cleaning is single system data cleaning, which only meets the data quality requirements of the system and cannot achieve consistency of cross-system data. Therefore, there are great obstacles to the standardization and unification of system data. Summary of the invention

[0004] In view of the above problems, the present invention is proposed to provide a data cleaning system and method that overcomes the above problems or at least partially solves the above problems.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A data cleaning system, comprising:

[0007] Source data module, used to define the source data extraction time interval and perform data extraction;

[0008] Screening rules module, used to define screening rules;

[0009] Screening results module, used to support access to screening results;

[0010] Audit authorization module, used to support process approval and counter-signature.

[0011] Optionally, the data extraction includes:

[0012] Access the database to extract data;

[0013] Data synchronization through database logs;

[0014] Synchronize data with the file interface.

[0015] Optionally, the source data module performs real-time monitoring and notification of the extraction task through record number comparison, file size comparison and error log.

[0016] Optionally, the screening rules include single-field compliance verification, code value granularity verification, multi-field correlation verification, suspected customer determination rules and data consistency rules.

[0017] Optionally, the single-field compliance check includes mandatory check, length check, content check and format check.

[0018] Optionally, the screening result module supports access to the screening results through fixed reports, flexible queries, multi-dimensional analysis and data penetration.

[0019] Optionally, the review and authorization module encapsulates the SpringBoot workflow orchestration engine, and after encapsulation, the approval authority is set according to the organization, user and role.

[0020] The present invention also provides a data cleaning method, which is based on the data cleaning system described in any one of the above items, and comprises:

[0021] Extract data and generate files;

[0022] Clean the environment and load data;

[0023] After the data is loaded, execute the data exploration SQL to generate a list of problematic data;

[0024] Complete the processing of problem data according to the problem data processing rules;

[0025] Aggregate and merge manually corrected data and automatically corrected data.

[0026] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0027] 1. The present invention can make a unified analysis of the data and screen out non-compliant data, so that the data meets the business and regulatory requirements, thereby improving the data quality of the entire system to meet the requirements of information digitization.

[0028] 2. The present invention ensures that the data of each system meets the current business and regulatory requirements, and also ensures the consistency of the same information in different systems, thereby effectively improving the data quality. The cleaned high-quality data improves the accuracy of data analysis and data mining, provides strong data support for business decisions, and lays a solid foundation for the standardization of future systems.

[0029] 3. The present invention realizes unified analysis of cross-system data, comprehensive screening of non-compliant data, and unified comprehensive cleaning of cross-system data, avoiding repeated cleaning of different systems and inconsistent cleaning results of the same information items in different systems due to different system business focuses. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A schematic diagram of a data cleaning system flow provided in an embodiment of the present application;

[0031] Figure 2 A schematic diagram of the structure of a data cleaning system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0033] Embodiment 1:

[0034] See also Figure 1 and Figure 2 , this embodiment provides a data cleaning system, the system comprising:

[0035] The source data module is used to define the source data extraction time interval and perform data extraction.

[0036] The data extraction includes: accessing the database to extract data; synchronizing data through the database log; synchronizing data with the file interface. The extraction method can be freely selected according to requirements.

[0037] The source data module also has a complete data integrity check function and log recording function, which can monitor and notify the extraction task in real time through record number comparison, file size comparison and error log, so as to ensure data integrity and avoid data loss during the extraction process.

[0038] The screening rule module is used to define screening rules.

[0039] The screening rules include single-field compliance verification, code value granularity verification, multi-field correlation verification, suspected customer determination rules and data consistency rules. Single-field compliance verification, code value granularity verification, and multi-field correlation verification are mainly routine verifications within a single system. Single-field compliance verification includes mandatory verification, length verification, content verification, and format verification. Suspected customer determination rules and data consistency rules are screening rules unique to the data cleaning platform, which are different from the data verification rules of other systems. The data consistency rule is to perform consistency verification on common fields between multiple systems. Suspected customer determination rules can be applied to both a single system and multiple systems. This type of rule will use multiple conditions to determine whether different customers are the same natural person or legal person. The specific suspected customer determination rules are as follows:

[0040] Suspected customer determination rule 1: Customers with the same ID number but different ID types.

[0041] Suspected customer determination rule 2: The ID type is "51-second-generation resident ID card, 53-household register, 54-other personal ID". The 18-digit ID number removes the 7th, 8th, and 18th digits. The ID numbers are the same and they are considered a group of suspected customers.

[0042] Rule 3 for determining suspected customers: The numbers in the ID card and household registration booklet are different, but they are the same after being converted into 18-digit uppercase numbers.

[0043] The screening results module is used to support access to screening results.

[0044] The screening result module supports access to the screening results through fixed reports, flexible queries, multi-dimensional analysis and data penetration, which is convenient for business personnel to discover and carry out data cleaning work. The screening results can be accessed through PC and mobile terminals respectively. The PC terminal uses the intranet data cleaning platform to access, and the mobile terminal uses the function entrance of the enterprise WeChat APP to access.

[0045] Audit authorization module, used to support process approval and counter-signature.

[0046] The review and authorization module encapsulates the SpringBoot workflow orchestration engine, and after encapsulation, the approval authority is set according to the organization, user and role. It also supports process approval and multi-person countersignature functions, which not only ensures data security, but also accurately issues problem data, facilitating the smooth progress of data cleaning work.

[0047] This embodiment can perform unified analysis on the data and screen out non-compliant data to make the data meet business and regulatory requirements, thereby improving the data quality of the entire system to meet the requirements of information digitization.

[0048] This embodiment ensures that the data in each system meets the current business and regulatory requirements, and also ensures the consistency of the same information in different systems, thereby effectively improving the data quality. The high-quality data after cleaning improves the accuracy of data analysis and data mining, provides strong data support for business decisions, and lays a solid foundation for the standardization of future systems.

[0049] This embodiment realizes unified analysis of cross-system data, comprehensive screening of non-compliant data, and unified comprehensive cleaning of cross-system data, avoiding repeated cleaning of different systems and inconsistent cleaning results of the same information items in different systems due to different system business focuses.

[0050] Embodiment 2:

[0051] Embodiment 2 discloses a data cleaning method, which is based on the above data cleaning system and includes the following steps:

[0052] S1. Extract data and generate files. The data warehouse extracts the full amount of customer information from the core, credit and credit card system production environment, and generates files that are placed on the data exchange platform.

[0053] S2. Clean the environment and load data to execute the data loading script to complete the work of loading the source system production data into the production network segment data cleaning platform.

[0054] S3. After data loading is completed, execute the data exploration SQL to generate a list of problematic data, including system, table, customer information items, inspection rules, organization, and date. The problematic data generated in each round is registered in the problematic data list without deduplication.

[0055] S4. Complete the processing of problem data according to the problem data processing rules. The Operation Management Department takes the lead, and the Risk Management Department and the Credit Card Center cooperate to propose problem data processing plans. Complete the processing of problem data according to the automatic problem data processing rules proposed by the business department.

[0056] S5. Summarize and merge the manually corrected data and the automatically corrected data to form an integrated file for correcting the problems in ECIF data migration. The result file is used for ECIF data migration.

[0057] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A data cleaning system, characterized in that: The system comprises: Source data module, used to define the source data extraction time interval and perform data extraction; Screening rules module, used to define screening rules; Screening results module, used to support access to screening results; Audit authorization module, used to support process approval and counter-signature.

2. A data cleaning system as claimed in claim 1, characterized in that: The data extraction includes: Access the database to extract data; Data synchronization through database logs; Synchronize data with the file interface.

3. A data cleaning system as claimed in claim 1, characterized in that: The source data module performs real-time monitoring and notification of the extraction task through record number comparison, file size comparison and error log.

4. A data cleaning system as claimed in claim 1, characterized in that: The screening rules include single-field compliance verification, code value granularity verification, multi-field correlation verification, suspected customer determination rules and data consistency rules.

5. A data cleaning system as claimed in claim 4, characterized in that: The single-field compliance check includes mandatory check, length check, content check and format check.

6. A data cleaning system as claimed in claim 1, characterized in that: The screening result module supports access to the screening results through fixed reports, flexible queries, multi-dimensional analysis and data penetration.

7. A data cleaning system as claimed in claim 1, characterized in that: The review and authorization module encapsulates the SpringBoot workflow orchestration engine, and after encapsulation, the approval authority is set according to the organization, user and role.

8. A data cleaning method, the method being based on a data cleaning system according to any one of claims 1 to 7, characterized in that: The method comprises: Extract data and generate files; Clean the environment and load data; After the data is loaded, execute the data exploration SQL to generate a list of problematic data; Complete the processing of problem data according to the problem data processing rules; Aggregate and merge manually corrected data and automatically corrected data.