Distributed data management method and system, electronic equipment and medium
By employing distributed data governance methods, utilizing hash value deduplication and multi-dimensional quality assessment, the problem of efficient processing and precise control of massive heterogeneous data has been solved. This has enabled precise data quality classification and automated system governance, thereby improving data processing efficiency and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DIGITAL ZHEJIANG TECH OPERATION CO LTD
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies struggle to efficiently process and accurately manage massive amounts of heterogeneous data, exhibiting problems such as low deduplication efficiency, delayed change detection, weak structural adaptability, insufficient data quality assessment, and poor adaptability of data collection modes.
A distributed data governance approach is adopted, which uses hash values for data deduplication and change identification, performs multi-dimensional quality assessment based on a pre-built governance rule base, and performs data filtering and standardization processing. Combined with a distributed microservice architecture, an automated governance process is achieved.
It improves data deduplication efficiency and change detection accuracy, achieves precise data quality classification and reliability, meets the needs of efficient processing and precise management of massive heterogeneous data, reduces operation and maintenance costs, and improves the system's versatility and scalability.
Smart Images

Figure CN121901201A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data governance technology, and in particular to a distributed data governance method, system, electronic device, and medium. Background Technology
[0002] A public data platform refers to a platform used for storing, managing, and sharing public data. Government data from various regions and departments needs to be collected and processed on this public data platform. However, the technologies currently used in the field of government data governance are insufficient to meet the demands for efficient processing and precise control of massive amounts of heterogeneous data. Summary of the Invention
[0003] In view of this, the purpose of this invention is to provide a distributed data governance method, system, electronic device and medium that can meet the needs of efficient processing and precise control of massive heterogeneous data.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a distributed data governance method, comprising: synchronizing data tables from multiple data source front-end databases to a collection database based on a pre-built governance configuration table, and extracting data tables from the collection database to a cleaning database according to batch numbers; calculating the hash value of the data tables, and deduplicating and modifying the data tables in the cleaning database based on the hash value; performing multi-dimensional quality assessment on the deduplicated data tables based on a pre-built governance rule base to obtain quality assessment results; filtering the data tables based on the quality assessment results and standardizing the data tables; generating a result data table based on the filtered data tables, and exporting the result data table.
[0005] Optionally, data tables from multiple data source front-end databases are synchronized to the collection database, and data tables from the collection database are extracted to the cleaning database according to batch numbers. This includes: synchronizing data tables from the front-end databases to the collection database and generating batch numbers for the data tables according to time information; extracting the data tables to the cleaning database based on the batch numbers of the data tables, and storing the data tables according to batch partitions.
[0006] Optionally, the hash value of the data table is calculated, and the data tables in the cleaning database are deduplicated and modified based on the hash value, including: calculating the hash value of the business primary key of the data table, and deduplicating the data tables in the batch partition based on the hash value of the business primary key; calculating the hash value of the business field of the data table, and identifying the change information of the business field in the data table based on the hash value of the business field.
[0007] Optionally, based on a pre-built governance rule base, a multi-dimensional quality assessment is performed on the deduplicated data table to obtain the quality assessment results, including: determining the quality score of the deduplicated data table based on the rule operation functions in the pre-built governance rule base; and marking the data table for validity based on the quality score; wherein, the first value indicates high-quality data, and the second value indicates problematic data.
[0008] Optionally, the data tables are filtered and standardized based on the quality assessment results, including: generating a first partitioned table and a first full unpartitioned table based on validity marking and deleted records; wherein the first partitioned table stores validity-marked data and deleted records, and the first full unpartitioned table stores validity-marked data; and standardizing the data in the data tables to generate a second partitioned table and a second full unpartitioned table; wherein the second partitioned table stores the standardized data, and the second full unpartitioned table stores the original data.
[0009] Optionally, a result data table is generated based on the filtered data table, and the result data table is exported. This includes generating a problem data table, a statistical data table, and a deletion data table based on the filtered data table, and exporting the problem data table, the statistical data table, and the deletion data table. The problem data table stores the values of problem business fields, triggering rules, and data after rectification. The statistical data table stores the total amount of collected data, the total amount of deduplicated data, the amount of high-quality data, and the amount of problem data. The deletion data table stores data deletion records in batches.
[0010] Optionally, it also includes: periodically comparing the table structures of the data tables on the directory side and the collection side, identifying table structure change information, and updating the governance configuration table based on the table structure change information.
[0011] Secondly, this invention provides a distributed data governance system, comprising: a data synchronization module, used to synchronize data tables from multiple data source front-end databases to a collection database based on a pre-built governance configuration table, and to extract data tables from the collection database to a cleaning database according to batch numbers; a deduplication and change detection module, used to calculate the hash value of the data tables, and to perform deduplication and change detection on the data tables in the cleaning database based on the hash value; a rule calculation module, used to perform multi-dimensional quality assessment on the deduplicated data tables based on a pre-built governance rule base, and to obtain quality assessment results; a standardization module, used to filter the data tables based on the quality assessment results and to perform standardization processing on the data tables; and a result export module, used to generate a result data table based on the filtered data tables and to export the result data table.
[0012] Thirdly, the present invention provides an electronic device including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the steps of the method provided in any of the first aspects above.
[0013] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, performs the steps of the method provided in any of the first aspects above.
[0014] This invention brings the following beneficial effects: The distributed data governance method, system, electronic device, and medium provided by this invention first synchronize data tables from multiple data source front-end databases to a collection database based on a pre-built governance configuration table, and then extract data tables from the collection database to a cleaning database according to batch numbers. Second, the hash value of the data tables is calculated, and deduplication and modification are performed on the data tables in the cleaning database based on the hash value. Next, a multi-dimensional quality assessment is performed on the deduplicated data tables based on a pre-built governance rule base to obtain the quality assessment results. Then, the data tables are filtered and standardized based on the quality assessment results. Finally, a result data table is generated based on the filtered data tables and exported. In this method, data deduplication and modification identification are performed using hash values, improving deduplication efficiency and the accuracy of modification detection. Multi-dimensional quality assessment of data based on a pre-built governance rule base enables accurate data quality grading, improving data credibility. Simultaneously, data filtering and standardization based on the quality assessment results, while ensuring data authenticity and compliance, can meet the diverse needs of government business. In summary, the above methods can meet the needs for efficient processing and precise control of massive heterogeneous data.
[0015] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 A flowchart of a distributed data governance method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a data synchronization and deduplication detection process provided in an embodiment of the present invention; Figure 3 A flowchart illustrating a rule-based operation provided in an embodiment of the present invention; Figure 4 A schematic diagram of a high-quality data tagging process provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of a data standardization process provided in an embodiment of the present invention; Figure 6 A schematic diagram of a result export process provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of a distributed data governance system provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Currently, existing data governance technologies have the following core shortcomings: 1. Low efficiency in data deduplication: Traditional technology uses a full-field character-by-character comparison method for deduplication. In the case of petabyte-level government data, the comparison time increases linearly with the number of fields and the amount of data, resulting in high system resource consumption and long processing cycle.
[0021] 2. Change detection and synchronization lag: Traditional technology relies on full table scans to identify data changes, which cannot quickly capture changes in field values. During incremental synchronization, the amount of redundant data transmission is large, resulting in poor data timeliness.
[0022] 3. Weak adaptability to structural and rule changes: In traditional technologies, when adding fields, adjusting primary keys, or updating governance rules, intermediate tables need to be manually rebuilt, which can easily lead to inconsistent table structures and incompatibility between historical and new data, resulting in high operation and maintenance costs.
[0023] 4. Insufficient data quality assessment and traceability capabilities: Traditional technologies lack a multi-dimensional quality assessment system, resulting in low accuracy in labeling high-quality and problematic data; the data lineage tracing mechanism is imperfect, making it difficult to trace problematic data and failing to support full-process control.
[0024] 5. Poor adaptability of collection modes: In traditional technologies, the processing logic is not consistent between full collection and incremental collection scenarios, data storage and computing resources cannot be allocated flexibly, and the system lacks versatility and scalability.
[0025] Based on this, the distributed data governance method, system, electronic device and medium provided by the embodiments of the present invention can meet the needs of efficient processing and precise control of massive heterogeneous data.
[0026] To facilitate understanding of this embodiment, a distributed data governance method disclosed in this invention will first be described in detail. This method can be executed by electronic devices, such as smartphones, computers, and tablets. The method is applied to a distributed data governance system that adopts a distributed microservice architecture. It automates the entire data governance process through multiple collaborative functional modules. The core functional modules include: a configuration management module, a data synchronization module, a deduplication and change detection module, a rule calculation module, a standardization module, a result export module, and a collection and linkage module. All functional modules are containerized and elastically scaled based on Spark.
[0027] See Figure 1 The flowchart shown illustrates a distributed data governance method, which mainly includes the following steps S101 to S105: Step S101: Based on the pre-built governance configuration table, synchronize the data tables in the front-end databases of multiple data sources to the collection database, and extract the data tables in the collection database to the cleaning database according to the batch number.
[0028] In one implementation, a governance configuration table (pyodps_govern_config) is pre-built through the configuration management module to uniformly maintain catalog data items on the directory side, physical table structure on the aggregation side, field-level governance rules, and task configuration information, including core fields such as table unique identifier, project name, field definition, rule details, and lifecycle.
[0029] The system-built pluggable connector library provided in this embodiment of the invention supports access to various government data sources such as MySQL, Oracle, Hive, and API interfaces, and can automatically identify the data source type and load the corresponding adaptation strategy.
[0030] The configuration management module can also periodically compare the table structures of the data tables on the directory side and the aggregation side, identify table structure change information, and update the governance configuration table based on the table structure change information. Specifically, the configuration management module periodically compares the table structures of the data tables on the directory side and the aggregation side, identifies table structure change information such as field additions, deletions, and primary key changes, updates the governance configuration table, marks the change type, and provides a trigger basis for subsequent processing.
[0031] Based on this, when synchronizing data tables from multiple data source front-end databases to the collection database, and extracting data tables from the collection database to the cleaning database according to batch numbers, the following methods can be used, including but not limited to: First, synchronize the data tables from the front-end databases to the collection database, and generate batch numbers for the data tables according to time information; then, extract the data tables to the cleaning database based on the batch numbers of the data tables, and store the data tables according to batch partitions.
[0032] In practical implementation, the data synchronization module can synchronize data tables from the front-end database (databases maintained by various departments) to the collection database (databases maintained by the big data center). In the collection database, the data tables are updated with technical fields such as auto-incrementing primary key (dtz_id), data entry time (dtz_time), and operation flag (dtz_op). The corresponding batch number (batch_pt) is generated according to year, month, day, hour, minute, and second. This invention supports two modes: full collection and incremental collection.
[0033] Furthermore, the data synchronization module can use the `pipline_check.py` and `pipline_sync.py` jobs to extract data tables from the collection library to the cleaning library in batches, and store the data tables according to batch partitions. Specifically, the extracted raw data is converted into a unified format (Avro) and written to the raw data area of the data lake for persistent storage, ensuring consistency in subsequent processing. The extraction process can be processed in parallel, thereby improving data transmission efficiency.
[0034] In this embodiment of the invention, the configuration management module can realize multi-source data adaptation and configuration synchronization, improve the problems of inconsistent multi-source data formats and inconsistent configuration benchmarks, and provide a unified basis for subsequent governance; the data synchronization module can realize distributed data synchronization and batch extraction, improve the performance bottleneck of massive data synchronization, and realize efficient collection of full and incremental data.
[0035] Step S102: Calculate the hash value of the data table, and perform deduplication and modification on the data tables in the cleaning database based on the hash value.
[0036] In one implementation, when calculating the hash value of a data table and performing deduplication and modification on the data tables in the cleaning database based on the hash value, the following methods, including but not limited to, can be used: First, calculate the hash value of the business primary key of the data table, and then deduplicate the data tables within the batch partition based on the hash value of the business primary key.
[0037] In practical implementation, the deduplication and change detection module can calculate the MD5 hash value (hash_unique) of the business primary key of the data table, compress the multi-field combined primary key into a fixed 32-bit string, and deduplicate it according to the hash value within the batch partition, retaining the latest record with the largest hash value of the auto-incrementing primary key, thereby reducing the time complexity from O(N) to O(1).
[0038] Then, the hash value of the business field in the data table is calculated, and the change information of the business field in the data table is identified based on the hash value of the business field.
[0039] In practical implementation, the deduplication change detection module can calculate the MD5 hash value (hash_all) of all business fields in the data table, and quickly identify field value changes by comparing the hash values of the source table (i.e., the original data table in the aggregation database) with those of the target table (i.e., the calculated data table). In this embodiment of the invention, only change records are synchronized, which reduces redundant transmission.
[0040] In this embodiment of the invention, for the data table, additional fields such as source library name (src_project) and source table name (src_table) can be added to record the data source and provide support for subsequent lineage tracing.
[0041] In this embodiment of the invention, the deduplication and change detection module adopts dual hash-driven deduplication and change detection, which can improve the problems of low deduplication efficiency and lagging change detection, and achieve efficient deduplication and accurate change identification.
[0042] Step S103: Based on the pre-built governance rule base, perform a multi-dimensional quality assessment on the deduplicated data table to obtain the quality assessment results.
[0043] In one implementation, when performing a multi-dimensional quality assessment on the deduplicated data table based on a pre-built governance rule base to obtain the quality assessment result, the following methods can be used, including but not limited to: First, determine the quality score of the deduplicated data table based on the rule operation functions in the pre-built governance rule base; then, mark the data table for validity based on the quality score; wherein, a first value indicates high-quality data, and a second value indicates problematic data.
[0044] In practical implementation, the rule processing module calls UDF functions in the governance rule base through the pipeline_govern.py job to perform rule processing functions such as field non-empty validation, format validation, and numerical range detection, obtaining fields such as rule execution details (_exam_result), filtering score (_filter_score), standardization score (_std_score), and original value retention score (_keep_score) to construct a multi-dimensional quality assessment system. Furthermore, based on the quality score, data validity is marked (_is_valid), where 1 (first value) represents high-quality data and 0 (second value) represents problematic data, providing a basis for subsequent screening.
[0045] In this embodiment of the invention, the rule operation module can improve the problem of low accuracy in data quality assessment by performing rule operations and multi-dimensional quality assessment, and realize intelligent labeling of high-quality data and problem data.
[0046] Step S104: Filter the data table based on the quality assessment results and standardize the data table.
[0047] In one implementation, when filtering and standardizing data tables based on quality assessment results, the following methods can be used: First, a first partitioned table and a first full non-partitioned table are generated based on validity markers and data deletion records; wherein, the first partitioned table stores validity marker data and data deletion records, and the first full non-partitioned table stores validity marker data.
[0048] In practical implementation, the standardization module can filter high-quality data and generate a first partitioned table (collection table_valid_new) and a first full non-partitioned table (collection table_valid). The valid_new table can merge upstream data and deduplicate it according to the business primary key, retaining all valid data and data deletion records; the valid table removes data deletion records and only stores valid data.
[0049] Then, the data in the data table is standardized to generate a second partitioned table and a second full non-partitioned table; the second partitioned table stores the standardized data, and the second full non-partitioned table stores the original data.
[0050] In practical implementation, the standardization module can perform standardization transformation on the data, generating a second partition table (collection table _std_new) and a second full non-partition table (collection table _std). The std_new table performs standardization processing operations such as format unification (e.g., date format YYYY-MM-DD) and code standardization according to government data standards, and stores the transformed data; the valid series tables retain the original data for users to access as needed.
[0051] This invention also supports merging table data in both full and incremental scenarios. During incremental updates, only the changed partitions are processed, while during full updates, the table structure is rebuilt and the latest data is synchronized.
[0052] In this embodiment of the invention, the standardization module can perform high-quality data screening and standardization conversion, thereby improving the problem of inconsistent data formats and realizing dual-track storage of raw data and standardized data.
[0053] Step S105: Generate a result data table based on the filtered data table, and export the result data table.
[0054] In one implementation, when generating a result data table based on the filtered data table and exporting the result data table, the following methods can be used, including but not limited to: generating a problem data table, a statistical data table, and a deletion data table based on the filtered data table, and exporting the problem data table, the statistical data table, and the deletion data table; wherein, the problem data table stores the problem business field values, triggering rules, and rectified data; the statistical data table stores the total amount of collected data, the total amount of deduplicated data, the amount of high-quality data, and the amount of problem data; and the deletion data table stores data deletion records in batches.
[0055] In practical implementation, the results export module can export results by category. Specifically, it generates an issue data table (collection table_issue) based on the filtered data table, recording information such as issue field values, triggering rules, and data after rectification; it generates a statistical data table (collection table_stat) to count indicators such as the total amount of collected data, the total amount of deduplicated data, the amount of high-quality data, and the amount of issue data; and it generates a deletion data table (collection table_del) to store data deletion records in batches.
[0056] Furthermore, the results export module also provides an automatic change adaptation strategy. When a field is added, the primary key is changed, or the rule is updated, all intermediate tables except the statistical data table are automatically deleted and rebuilt, and the latest data is pulled to ensure compatibility. When a field without a configured rule is deleted, only the latest partition data is updated, and the table is not rebuilt.
[0057] The aggregation and linkage module adopts an aggregation and linkage mechanism. When the aggregation side rebuilds or mounts a table, it triggers a new cleanup through incremental notifications. When a task goes online or offline, it compares the differences between partitions and pulls out the unsynchronized data for cleanup.
[0058] In this embodiment of the invention, the result export module, namely the collection and linkage module, can improve the problem of missing governance loop by linking result export with automated governance, and realize problem tracing, statistical analysis and automated control.
[0059] The distributed data governance method provided in this invention improves deduplication efficiency and change detection accuracy by using hash values for data deduplication and change identification. Based on a pre-built governance rule base, it performs multi-dimensional quality assessment of data, enabling precise data quality grading and enhancing data credibility. Simultaneously, it filters data based on the quality assessment results and standardizes the data, ensuring both data authenticity and compliance while meeting the diverse needs of government operations. In summary, this method can meet the requirements for efficient processing and precise control of massive amounts of heterogeneous data.
[0060] For ease of understanding, this embodiment of the invention also provides a distributed data governance implementation process, which mainly includes the following steps one to six: Step 1: System deployment and configuration initialization.
[0061] Specifically, deploy the MaxCompute (ODPS) big data platform, create the project (e.g., ggsj_sjzc), divide the storage areas corresponding to the front-end library, aggregation library, and cleaning library, and deploy each functional module based on Kubernetes and configure elastic scaling strategies.
[0062] Configure the governance configuration table (pyodps_govern_config), and enter the government data source connection information, directory-side cataloging data items (such as id, name, age, birth, etc.), governance rules (such as non-empty validation of the id field), and task configuration (collection method: tool collection, lifecycle 0, already online). See Table 1 for details.
[0063] Table 1. Schematic diagram of governance configuration table
[0064] Load the plug-in connector, automatically identify the connected MySQL government data source, and complete the data source registration and adaptation.
[0065] In practical implementation, the configuration information in the governance configuration table includes: table information, collection table field information, cleaning table field information, and field governance configuration information, etc. The collection table name format can be: department abbreviation_collection business code_department number_source business table name, to ensure the global consistency of the collection table.
[0066] Step 2: Data synchronization and batch extraction.
[0067] Specifically, the front-end database pushes multiple batches of government data, including addition, update, and deletion operations. The collection database synchronizes the data and generates batch_pt batch numbers (such as 20250708162523), and adds technical fields such as auto-incrementing primary key (dtz_id), data entry time (dtz_time), and operation flag (dtz_op).
[0068] Run the pipeline_sync.py job to extract data from the collection library to the cleaning library in batches, convert it to the unified Avro format, and then write it to the raw data area of the data lake. At the same time, it triggers table structure comparison and updates the governance configuration table.
[0069] Step 3: Perform deduplication and change detection.
[0070] Specifically, the cleansing database generates a collection table (_unique table), and calculates the MD5 hash value of the business primary key of each record (hash_unique, based on the id field) and the MD5 hash value of all business fields (hash_all, based on the id, name, age, and birth fields).
[0071] Then, the hash_unique value is deduplicated by batch partition, keeping the latest record with the largest dtz_id and removing duplicate data; the hash_all values of the source table and the target table are compared to identify the change record where the age field is updated from 25 to 26, and only this changed data is synchronized.
[0072] For steps two to three mentioned above, see Figure 2 As shown, firstly, the system obtains the aggregation database connection information, table information, and field information, as well as the directory master data, directory data item information, and data warehouse table field information. Then, it compares the aggregation database table structure with the directory-side table structure. If the structure changes, the aggregation database table structure and configuration information are modified, and a primary key comparison is performed. If the primary key changes, data backup is performed, and the target table structure is created or updated. If the primary key remains unchanged, a rule comparison is performed. If a new table structure is added, the table information, table field information, and initial rule configuration information are recorded, and the target table structure is created or updated. If the table structure remains unchanged, a rule comparison is performed, and the aggregation database data is extracted to the target table in batches, and deduplication is performed to obtain the partitioned deduplication results.
[0073] Specifically, partition deduplication includes: (1) comparing whether the primary key, rules, and fields have changed; (2) calculating the hash value of the concatenated primary key field / business field; (3) adding the original space name and original table name fields; and (4) deduplicating within the partition according to the hash_unique hash value. The final result table is named the aggregation table name_unique, which includes: the MD5 hash value of the business primary key hash_unique, the MD5 hash value of the business field hash_all, the source database, the source table name, the processing timestamp, and the processing batch (partition), etc.
[0074] Step 4: Rule-based calculation and quality labeling.
[0075] Specifically, the pipline_govern.py job is run, which calls the governance rule base to perform non-empty checks on the id field and value range checks on the age field, generating the aggregation table name_exam.
[0076] Calculate the filtering score _filter_score, the standardized score _std_score, and the original value retention score _keep_score for each field. Mark records with a non-empty id field as _is_valid=1 (high-quality data) and records with an empty id field as _is_valid=0 (problematic data). The rule execution details _exam_result field records the violation details of an empty id field.
[0077] See Figure 3 As shown, firstly, the structure information of the `exam` table is obtained, and it is determined whether the table structure information exists. If it exists, the target table structure is generated, and upstream data is read. Then, the aggregation method is determined. If it is incremental aggregation, the incremental partition data is merged; if it is full aggregation, the latest partition data is obtained. Afterward, based on the governance rule base, a `pyodps udf` is run to generate governance rule results. Specifically, the running of the UDF includes general rules, data edge transformation rules, one-data-one-source verification rules, etc., and calculates the quality score, saving the result data. The result table is named `aggregation-table-name_exam`, and includes: the MD5 hash value of the business primary key (`hash_unique`), the MD5 hash value of the business field (`hash_all`), the rule result, the standardization result, the filtering score, the score for retaining the original value, the standardization score, the validity flag, the source database, the source table name, the processing timestamp, and the processing batch (partition), etc.
[0078] Step 5: High-quality data screening and standardization transformation.
[0079] Specifically, a valid_new partitioned table is generated based on the data in the exam table. Multiple batches of data are merged and deduplicated by primary key, retaining all data marked _is_valid and records with op=delete. A valid full table is generated, deleting delete records and storing only valid data.
[0080] Furthermore, generate the std_new partition table and the std full table, unify the format of the birth field to YYYY-MM-DD, such as converting 19990520 to 1999-05-20, and store the converted data; the valid series tables retain the original format data.
[0081] See Figure 4 As shown, the high-quality data tagging mainly includes: First, obtaining the structure information of the valid_new table and determining whether the table structure exists. If it exists, the target table structure is generated, and upstream data is read. Then, based on the governance rule base, a high-quality result table _valid_new is generated, and the full table _valid is initialized. Specifically, the high-quality result table is named collection table name _valid_new, and includes: business primary key MD5 hash value hash_unique, business field MD5 hash value hash_all, validity flag, source library, source table name, processing timestamp, and processing batch (partition) information. The full table is named ods_collection table name _valid, and includes: business primary key MD5 hash value hash_unique, business field MD5 hash value hash_all, validity flag (partition), source library, source table name, processing timestamp, and processing batch (partition) information.
[0082] See Figure 5 As shown, the data standardization process includes: First, obtaining the structure information of the `std_new` table and determining whether the table structure exists. If it exists, the target table structure is generated, and upstream data is read. Then, based on the governance rule base, a high-quality result table `_std_new` is generated, and the full table `_std` is initialized. Specifically, the high-quality result table is named `aggregate_table_name_std_new`, and includes fields such as: business primary key MD5 hash value `hash_unique`, business field MD5 hash value `hash_all`, standardized fields, validity flags, source database, source table name, processing timestamp, and processing batch (partition). The full table is named `ods_aggregate_table_name_std`, and includes fields such as: business primary key MD5 hash value `hash_unique`, business field MD5 hash value `hash_all`, standardized fields, validity flags (partition), source database, source table name, processing timestamp, and processing batch (partition).
[0083] Step Six: Results Export and Collaborative Governance.
[0084] Specifically, an issue data table is generated to record details of issue data where the id field is empty, the triggered non-empty verification rules, and rectification suggestions; a statistical data table is generated to count the total number of data collected in this batch (1000), the number of deduplicated data (800), the number of high-quality data (750), and the number of issue data (50), etc.
[0085] See Figure 6 As shown, the result export mainly includes: First, obtaining the structure information of the issue table and determining whether the table structure exists. If it exists, the target table structure is generated, and all data is read to obtain the issue data table. Then, the amount of aggregated data, the current amount of issue data, and the current amount of high-quality data are calculated, and the previous statistical information is read for calculation to generate a governance statistics table. The issue data table is named `aggregate_table_name_issue` and includes: English table name, Chinese table name, primary key, field value, post-rectification field value, governance mode, rule type, rule changed to, business primary key MD5 value, governance case ID, source space, processing timestamp, and whether rectification has been completed. The governance statistics table is named `aggregate_table_name_stat` and includes: source space, English table name, Chinese table name, total aggregated data, total cleaned data, total deduplicated data, total issue data, total high-quality data, cumulative total issue data, cumulative total rectified data, rectification time, processing timestamp, and processing batch (partition).
[0086] Furthermore, after the addition of the "ID number" field, the system automatically recognizes the structural change, deletes and rebuilds intermediate tables such as valid_new and std_new, and pulls the latest data for reprocessing; when the collection side rebuilds the tables, it triggers the governance system to clean the data again through incremental notifications to ensure data consistency.
[0087] The distributed data governance method provided in this embodiment of the invention has the following technical advantages compared with traditional methods: (1) Improve the efficiency of data deduplication under large data volume, reduce system resource consumption and shorten the processing cycle: The dual hash deduplication mechanism is adopted to replace the traditional full field comparison. The deduplication efficiency is improved by more than 50% in the scenario of millions of data, the system resource consumption is reduced by 40%, and the processing cycle is significantly shortened.
[0088] (2) Achieve rapid and accurate detection of data changes, reduce redundant data transmission, and ensure data timeliness: The hash_all field enables rapid identification of changes, reduces redundant data transmission by 80% during incremental synchronization, and improves data timeliness from the hour level to the minute level.
[0089] (3) Automatically adapts to various changes in table structure, primary key and governance rules without manual intervention, avoids table structure inconsistency, reduces operation and maintenance costs (60% reduction in operation and maintenance costs), and ensures data compatibility.
[0090] (4) Construct a multi-dimensional data quality assessment system, accurately mark the data quality level, improve data lineage tracing, and support problem tracing: The multi-dimensional quality assessment system realizes accurate data quality classification, and the improved lineage tracing mechanism supports the full-process tracing of problem data, significantly improving data credibility.
[0091] (5) Unify the processing logic of full and incremental collection scenarios, realize elastic resource allocation, and improve the system's versatility and scalability: support distributed deployment and elastic scaling, adapt to the construction of multi-level government data warehouses at the city and district levels, and meet the data governance needs of different scales.
[0092] (6) Significantly enhanced data availability: Dual-track storage of raw data and standardized data, combined with problem details and statistical indicators, takes into account both data authenticity and compliance, and meets the diverse needs of government business.
[0093] In addition to the distributed data governance method provided in the foregoing embodiments, this invention also provides a distributed data governance system, see [link to relevant documentation]. Figure 7 The diagram shown illustrates the structure of a distributed data governance system, indicating that the system mainly comprises the following components: The data synchronization module 701 is used to synchronize data tables from multiple data source front-end databases to the collection database based on a pre-built governance configuration table, and to extract data tables from the collection database to the cleaning database according to batch number.
[0094] The deduplication and change detection module 702 is used to calculate the hash value of the data table and perform deduplication and change on the data table in the cleaning database based on the hash value.
[0095] The rule calculation module 703 is used to perform multi-dimensional quality assessment on the deduplicated data table based on a pre-built governance rule base, and obtain the quality assessment results.
[0096] Standardization module 704 is used to filter data tables based on quality assessment results and standardize the data tables.
[0097] The result export module 705 is used to generate a result data table based on the filtered data table and export the result data table.
[0098] The distributed data governance system provided in this invention improves deduplication efficiency and change detection accuracy by using hash values for data deduplication and change identification. Based on a pre-built governance rule base, it performs multi-dimensional quality assessments of the data, enabling precise data quality grading and enhancing data credibility. Simultaneously, it filters data based on the quality assessment results and standardizes the data, ensuring both data authenticity and compliance while meeting the diverse needs of government operations. In summary, this system can meet the requirements for efficient processing and precise control of massive amounts of heterogeneous data.
[0099] It should be noted that the system provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the system embodiment can be referred to the corresponding content in the aforementioned method embodiment. The specific numerical values provided in the implementation of this invention are merely exemplary and are not intended to limit the scope of the invention.
[0100] This invention also provides an electronic device, specifically, the electronic device includes a processor and a storage device; the storage device stores a computer program, and the computer program, when run by the processor, executes the method described in any of the above embodiments.
[0101] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 80, a memory 81, a bus 82, and a communication interface 83. The processor 80, the communication interface 83, and the memory 81 are connected through the bus 82. The processor 80 is used to execute executable modules, such as computer programs, stored in the memory 81.
[0102] The memory 81 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 83 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.
[0103] Bus 82 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 8 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0104] The memory 81 is used to store programs. After receiving an execution instruction, the processor 80 executes the program. The method executed by the device for defining the flow process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 80 or implemented by the processor 80.
[0105] The processor 80 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 80 or by instructions in software form. The processor 80 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 81. The processor 80 reads the information in memory 81 and, in conjunction with its hardware, completes the steps of the above method.
[0106] The computer program product of the readable storage medium provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, please refer to the foregoing method embodiments, which will not be repeated here.
[0107] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0108] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A distributed data governance method, characterized in that, include: Based on the pre-built governance configuration table, the data tables in the front-end databases of multiple data sources are synchronized to the collection database, and the data tables in the collection database are extracted to the cleaning database according to the batch number. Calculate the hash value of the data table, and perform deduplication and modification on the data tables in the cleaning database based on the hash value; Based on a pre-built governance rule base, a multi-dimensional quality assessment is performed on the deduplicated data table to obtain the quality assessment results. The data table is filtered based on the quality assessment results, and the data table is then standardized. A result data table is generated based on the filtered data table, and the result data table is exported.
2. The method according to claim 1, characterized in that, Synchronize data tables from multiple data source front-end databases to the aggregation database, and extract data tables from the aggregation database to the cleaning database according to batch number, including: The data tables in the front-end database are synchronized to the collection database, and batch numbers of the data tables are generated according to time information. The data table is extracted to the cleaning library based on the batch number of the data table, and the data table is stored according to the batch partition.
3. The method according to claim 2, characterized in that, Calculate the hash value of the data table, and perform deduplication and modification on the data tables in the cleaning database based on the hash value, including: Calculate the hash value of the business primary key of the data table, and deduplicate the data tables in the batch partition based on the hash value of the business primary key; Calculate the hash value of the business field in the data table, and identify the change information of the business field in the data table based on the hash value of the business field.
4. The method according to claim 1, characterized in that, Based on a pre-built governance rule base, a multi-dimensional quality assessment is performed on the deduplicated data table to obtain the quality assessment results, including: Based on the rule operation functions in the pre-built governance rule base, the quality score of the deduplicated data table is determined; The data table is marked as valid based on the quality score; wherein, a first value indicates high-quality data, and a second value indicates problematic data.
5. The method according to claim 4, characterized in that, Based on the quality assessment results, the data table is filtered and standardized, including: Based on the validity markers and data deletion records, a first partitioned table and a first full non-partitioned table are generated; wherein, the first partitioned table stores validity marker data and data deletion records, and the first full non-partitioned table stores validity marker data; The data in the data table is standardized to generate a second partitioned table and a second full non-partitioned table; wherein, the second partitioned table stores the standardized data, and the second full non-partitioned table stores the original data.
6. The method according to claim 1, characterized in that, Generate a result data table based on the filtered data table, and export the result data table, including: Based on the filtered data table, a problem data table, a statistical data table, and a deletion data table are generated, and the problem data table, the statistical data table, and the deletion data table are exported. The problem data table stores the problem business field values, triggering rules, and data after rectification. The statistical data table stores the total amount of collected data, the total amount of deduplicated data, the amount of high-quality data, and the amount of problem data. The deletion data table stores data deletion records in batches.
7. The method according to claim 1, characterized in that, Also includes: Periodically compare the table structures of the data tables on the directory side and the collection side, identify table structure change information, and update the governance configuration table based on the table structure change information.
8. A distributed data governance system, characterized in that, include: The data synchronization module is used to synchronize data tables from multiple data source front-end databases to the collection database based on a pre-built governance configuration table, and to extract data tables from the collection database to the cleaning database according to batch number. The deduplication and change detection module is used to calculate the hash value of the data table and perform deduplication and change on the data table in the cleaning database based on the hash value; The rule calculation module is used to perform multi-dimensional quality assessment on the deduplicated data table based on a pre-built governance rule base, and obtain the quality assessment results. The standardization module is used to filter the data table based on the quality assessment results and to standardize the data table. The results export module is used to generate a results data table based on the filtered data table and export the results data table.
9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the steps of the method described in any one of claims 1 to 7.