Data verification method and device, equipment, storage medium and program product

By acquiring field identification information and generating rule mapping relationships, and combining a distributed computing framework to perform parallel verification of Hive table data, the problem of low efficiency in Hive table data quality monitoring in existing technologies is solved, achieving efficient full verification and accurate problem location.

CN121833683APending Publication Date: 2026-04-10INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2025-12-30
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing Hive table data quality monitoring solutions are difficult to adapt to massive data volumes, resulting in low data quality verification efficiency and weak problem tracing capabilities.

Method used

By obtaining field identification information, generating rule mapping relationships, and combining preset rules and custom rules, a distributed computing framework is used to perform parallel verification of Hive table data, thereby achieving field-level data quality monitoring.

Benefits of technology

It significantly improves the efficiency and accuracy of Hive table data quality verification, enhances the precision of data problem location, realizes efficient full-scale verification in large-scale data scenarios, and strengthens the comprehensiveness and real-time nature of data quality verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833683A_ABST
    Figure CN121833683A_ABST
Patent Text Reader

Abstract

The invention provides a data verification method and device, equipment, a storage medium and a program product, and relates to the technical field of big data. According to the method, field identification information is obtained, and a rule mapping relation is generated based on the field identification information; and executing a data verification task based on the rule mapping relationship. According to the method, through rule configuration driven by the field identifier, the efficiency and accuracy of Hive table data quality verification are remarkably improved, the accuracy of data problem positioning is improved, efficient full-amount verification under a large-scale data scene is achieved, the comprehensiveness and real-time performance of data quality verification are improved, and the verification efficiency of full-amount data is also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, and in particular to a data verification method, apparatus, device, storage medium, and program product. Background Technology

[0002] With the rapid development of big data technology, data has become a core strategic asset driving business innovation and optimizing operational decisions. Hive table data quality monitoring is a crucial link in ensuring data credibility and the reliability of business decisions in the big data field, and is widely used in data-intensive industries such as finance, e-commerce, telecommunications, and the Internet of Things.

[0003] Currently, in large-scale data scenarios, most existing Hive table data quality monitoring solutions adopt a combination of full scan and offline verification. However, existing Hive table data quality monitoring solutions are difficult to adapt to massive data volumes, resulting in low efficiency in data quality verification and weak problem tracing capabilities. Summary of the Invention

[0004] This application provides a data verification method, apparatus, device, storage medium, and program product to solve the problem that existing Hive table data quality monitoring solutions are difficult to adapt to massive data scales, resulting in low efficiency of data quality verification and weak problem traceability.

[0005] Firstly, this application provides a data verification method, including:

[0006] Obtain field identification information, which is used to describe the attributes of the field;

[0007] A rule mapping relationship is generated based on the field identification information. The rule mapping relationship includes preset rules and custom rules. The preset rules include uniformly configured rules, and the custom rules include user-defined rules.

[0008] A data verification task is performed based on the rule mapping relationship, and the data verification task includes rule verification of Hive table data.

[0009] Secondly, this application provides a data verification device, comprising:

[0010] The acquisition module is used to acquire field identification information, which is used to describe the attributes of the field.

[0011] The processing module is used to generate a rule mapping relationship based on the field identification information. The rule mapping relationship includes preset rules and custom rules. The preset rules include uniformly configured rules, and the custom rules include user-defined rules.

[0012] The processing module is also used to perform a data verification task based on the rule mapping relationship, the data verification task including rule verification of Hive table data.

[0013] Thirdly, this application provides an electronic device, comprising:

[0014] A processor, and a memory communicatively connected to the processor;

[0015] The memory stores computer-executed instructions;

[0016] The processor executes computer execution instructions stored in the memory to implement the data verification method as described in the first aspect and various possible implementations of the first aspect above.

[0017] Fourthly, this application provides a computer storage medium storing computer execution instructions thereon, which are executed by a processor to implement the data verification method as described in the first aspect and various possible implementations of the first aspect above.

[0018] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the data verification method as described in the first aspect and various possible implementations of the first aspect.

[0019] The data verification method provided in this application obtains field identification information and matches it with preset rules, while also receiving user-inputted custom rules. It then obtains a rule mapping relationship based on the preset and custom rules. Using a distributed computing framework, it splits Hive table data into multiple partitions and verifies the Hive table data in multiple partitions in parallel based on the rule mapping relationship. This method, through field-identifier-driven rule configuration, solves the problems of complex rule configuration and difficulty in adapting to changing business needs in existing technologies. It significantly improves the efficiency and accuracy of Hive table data quality verification, enhances the precision of data problem location, and forms a full-link traceability capability from data anomalies to the root cause of the problem. It achieves efficient full-volume verification in large-scale data scenarios, improving not only the comprehensiveness and real-time performance of data quality verification but also the efficiency of full-volume data verification. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0021] Figure 1 This is a flowchart illustrating the data verification method provided in this application. Figure 1 .

[0022] Figure 2This is a flowchart illustrating the data verification method provided in this application. Figure 2 .

[0023] Figure 3 This is a schematic diagram of the data verification device provided in this application.

[0024] Figure 4 This is a schematic diagram of the structure of the electronic device provided in this application.

[0025] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0026] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0027] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.

[0028] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.

[0029] It should be noted that the data verification method, apparatus, equipment, storage medium and program products provided in this application can be used in the field of big data technology, or in any field other than big data. The application fields of the data verification method, apparatus, equipment, storage medium and program products in this application are not limited.

[0030] First, the terms used in this application will be explained.

[0031] Hive is a distributed data warehouse tool that provides users with large-scale data query and management services by mapping structured data files to database tables and providing a SQL-like query language.

[0032] SQL: Structured Query Language, abbreviated as SQL, is the standard language for operating relational databases / data warehouses.

[0033] Spark is a fast, general-purpose, and scalable big data analytics engine that is based on in-memory computing and supports a variety of tasks, including batch processing, interactive queries, real-time stream processing, machine learning, and graph computing.

[0034] Flink is an open-source stream and batch processing framework. It is based on a pure stream computing model, supports high-throughput, low-latency stateful computing, and provides exact-once semantics and event-time processing capabilities.

[0035] MapReduce is a distributed computing framework that enables parallel processing of large datasets through two phases: Map and Reduce. It is suitable for offline batch processing tasks.

[0036] With the rapid development of big data technology, data has become a core strategic asset driving business innovation and optimizing operational decisions. Hive table data quality monitoring is a crucial link in ensuring data credibility and the reliability of business decisions in the big data field, and is widely used in data-intensive industries such as finance, e-commerce, telecommunications, and the Internet of Things.

[0037] Hive, as a distributed data warehouse tool, has become a core component for building data storage and analysis systems across various industries. The data quality at the field level in Hive tables is a crucial guarantee for the entire data value chain, specifically including quality indicators such as the accuracy, completeness, and consistency of field data. This directly impacts the effectiveness of subsequent data modeling, data analysis, and business applications. For example, risk control models in the financial industry rely on accurate data from customer transaction amounts and credit information fields in Hive tables, while user profiling analysis on e-commerce platforms requires complete data from order fields and user behavior fields. Therefore, monitoring Hive table data quality has become a core aspect of ensuring data credibility in the big data field, and the performance and coverage of its technical solutions have attracted significant industry attention.

[0038] Currently, the core requirements for Hive table data quality monitoring focus on multi-dimensional verification of field data, including but not limited to: field value validity verification (such as range constraints for numeric fields and format specifications for string fields), data integrity verification (such as the proportion of null values ​​in non-null fields and the missing rate of key fields), data consistency verification (such as the consistency of values ​​of corresponding fields of the same entity in different tables and the matching degree between field values ​​and business rules), and data timeliness verification (such as the update frequency of incremental fields and the deviation between data writing and business occurrence time).

[0039] In large-scale data scenarios, most existing Hive table data quality monitoring solutions adopt a mode of full scan combined with offline verification.

[0040] However, existing Hive table data quality monitoring solutions are difficult to adapt to massive data volumes, resulting in low efficiency in data quality verification.

[0041] In addition, most solutions only focus on basic rules (such as not null, field type, and value range), and lack support for complex business rules (such as multi-field association validation: the "payment amount" in the order table must be greater than the "discount amount", and the "user ID" in the log table must match the user table), resulting in full data validation and completion, which is difficult to adapt to complex business needs.

[0042] To address the aforementioned issues, this application provides a data verification method.

[0043] The data verification method provided in this application constructs a configurable and automated data quality monitoring framework. It uses a UI interface to annotate field attributes (such as primary keys and enumeration fields) and automatically adapts to preset rules (such as requiring uniqueness verification for primary key fields). It utilizes the Spark distributed computing framework to perform a full data scan of Hive tables, dynamically executing verification logic in conjunction with a rule engine. By recording the data flow path (such as extraction, transformation, and loading stages in the ETL process) and combining log analysis technology, it quickly locates the root cause of problems, achieving full verification and dynamic rule adaptation of Hive table fields. Combined with a full-link problem tracing and closed-loop management mechanism, it addresses the core pain points of existing technologies, such as low efficiency, narrow coverage, rigid rules, and difficulty in problem localization. This method significantly improves the efficiency and accuracy of Hive table data quality verification through field-identified rule configuration, distributed verification, and full-link problem tracing. It achieves efficient full verification in large-scale data scenarios, reduces the possibility of sampling omissions, and not only improves the comprehensiveness and real-time performance of data quality verification but also enhances the efficiency of full data verification.

[0044] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0045] Figure 1 This is a flowchart illustrating the data verification method provided in this application. Figure 1 The executing entity in this embodiment can be, for example, a data quality verification system. Figure 1 As shown, the data verification method provided in this embodiment includes:

[0046] S101: Get field identification information.

[0047] Among them, the field identification information is used to describe the attributes of the field.

[0048] In this embodiment of the application, the data quality verification system verifies the quality of Hive table data and determines whether there is abnormal data in the Hive table data through logical verification, thereby achieving accurate location of Hive table data problems.

[0049] Hive table data refers to the set of all data values ​​corresponding to a single field (column) in a structured data table built on the Hive data warehouse. It is the smallest business data unit that constitutes the structured data of Hive tables. The field definitions in Hive table data (including field type, storage format, parsing rules, etc.) are parsed and extracted from physical files (such as Parquet, ORC, TextFile, etc.) stored in the underlying distributed file system (such as HDFS). They carry specific business attribute values ​​and are the core objects of Hive table data storage, query analysis, and business applications.

[0050] Understandably, the metadata information of Hive table fields, including field names, data types, and business meanings, is obtained through the Hive Metastore; field identification information defines, describes, and sets rules and constraints for Hive table data. For example, the field identification information obtained in this instance is used to identify whether the corresponding field is a "primary key field," "enumeration field," or "numeric field."

[0051] Specifically, Hive Metastore is one of the core components of Hive. Essentially, it is a centralized storage and service module that manages all Hive metadata. It is responsible for storing structured metadata information such as Hive tables, databases, partitions, columns, storage formats, and data paths, and provides metadata query, modification, and synchronization services for Hive clients and query engines. It is a key support for Hive to realize "structured mapping of distributed file data".

[0052] For example, in a financial scenario, attributes (such as primary key and enumeration fields) are labeled for Hive table fields through the user interface and stored as field identification information. The field identification information obtained at this time can be: the field "User ID" is labeled as "primary key field", the field "gender" is labeled as "enumeration field", and the field "transaction amount" is labeled as "numeric field".

[0053] S102: Generate rule mapping relationships based on field identification information.

[0054] The rule mapping relationship includes preset rules and custom rules. Preset rules include uniformly configured rules, while custom rules include user-defined rules.

[0055] As can be understood, rule mapping relationship refers to the association between field identification information and rules. The rule mapping relationship generated based on field identification information is a structured association between preset rules (such as "primary key field must be uniquely validated") and custom rules (such as "transaction amount must be ≥0") with field identification information (covering basic attributes, business attributes, quality attributes, management attributes, etc.) as the core association dimension. In other words, rule mapping relationship includes system preset rules and user-defined rules.

[0056] The preset rules include uniformly configured rules; these preset rules are designed for general scenarios of Hive table data quality control. They are standardized rules uniformly configured by the data quality verification system, and have the characteristics of universality and generality. They can cover the field quality verification needs in most industries or scenarios (such as field non-empty verification, basic format verification, general value range verification, etc.).

[0057] Custom rules include rules set by the user; these custom rules are non-standard rules defined by the user (data developers or business personnel) to meet the personalized needs of specific business scenarios and specific fields. They provide personalized supplementary rules for Hive table data quality control when the preset rules cannot cover personalized quality control needs (such as industry-specific field logic verification, cross-table field consistency verification, business scenario-specific value constraints, etc.).

[0058] By matching the features of field identifier information, the system can accurately associate and coordinate preset rules with custom rules, avoid rule redundancy, and ensure the comprehensiveness and flexibility of field data quality verification.

[0059] For example, in a financial scenario, the field "User ID" is marked as a "primary key field", and the system automatically associates it with a "uniqueness check" rule; the field "Transaction Amount" is marked as a "numeric field", and the user defines the rule "transaction amount ≥ 0", and it is stored in JSON format.

[0060] S103: Perform data verification tasks based on rule mapping relationships.

[0061] The data validation task includes rule validation of Hive table data.

[0062] Understandably, the purpose of data validation tasks, namely rule validation of Hive table data, is to ensure the validity, accuracy, consistency and integrity of Hive table data, and to provide a reliable data foundation for downstream applications (analysis, modeling, business decision-making, etc.). In other words, the data validation task is a process of performing validation on Hive table data based on rule mapping relationships, specifically including rule parsing, data scanning and anomaly marking.

[0063] By using rule-based validation (such as field value range validation, format validation, and association consistency validation), we can identify and filter erroneous data that does not conform to business logic or technical specifications (such as negative order amounts, incorrect mobile phone number formats, and user IDs that do not exist in the user table). This prevents erroneous data from entering the data analysis, report generation, or model training process and avoids drawing misleading business conclusions based on erroneous data (such as revenue statistics deviations or distorted user behavior analysis).

[0064] In this embodiment of the application, the data quality verification system is equipped with a rule engine, which can parse rule mapping relationships. The rule engine reads, identifies, and decomposes the structured content of the rule mapping relationship according to preset logic and syntax, extracts the core elements, and transforms them into instructions or logical models that the engine can execute. In other words, the rule engine can transform the rule mapping relationship from a static storage form into dynamic logic that the rule engine can understand, schedule, and execute, providing data support for the implementation of data verification tasks.

[0065] For example, after the rule engine parses the rule mapping relationship, it performs validation on the "User ID" and "Transaction Amount" fields of the Hive table and marks data records that do not conform to the rules during the validation process (such as duplicate user IDs or negative transaction amounts).

[0066] The data verification method provided in this embodiment obtains field identification information and generates rule mapping relationships based on the field identification information; then, it executes data verification tasks based on the rule mapping relationships. This method, through field identification-driven rule configuration, significantly improves the efficiency and accuracy of Hive table data quality verification, enhances the precision of data problem location, and achieves efficient full-volume verification in large-scale data scenarios. It not only improves the comprehensiveness and real-time performance of data quality verification but also increases the efficiency of full-volume data verification.

[0067] Figure 2 This is a flowchart illustrating the data verification method provided in this application. Figure 2 .like Figure 2 As shown, in this embodiment... Figure 1 Based on the embodiments, the data verification method is described in detail. The data verification method shown in this embodiment includes:

[0068] S201: Get field identification information.

[0069] Steps S201-S202 are similar to steps S101-S102 above, and will not be repeated here.

[0070] S202: Match preset rules based on field identifier information.

[0071] S203: Receive user-inputted custom rules.

[0072] S204: Obtain the rule mapping relationship based on preset rules and custom rules.

[0073] The preset rules include the validation logic corresponding to the field type; the custom rules include the field validation conditions and validation thresholds.

[0074] Understandably, preset rules are rules that are automatically associated with field identification information, such as the "uniqueness check" rule associated with the primary key field; custom rules are rules defined by the user, including field validation conditions (such as "transaction amount ≥ 0") and validation thresholds (such as "inventory ≤ 100,000").

[0075] By associating the field identifier information determined in this instance with the corresponding preset rules and the custom rules input by the user, a mapping table between field identifiers and rules is generated, that is, a complete rule mapping relationship is obtained, which is used for the automated configuration of subsequent data validation tasks.

[0076] For example, the data quality verification system automatically matches the "uniqueness verification" rule based on the "primary key field" identifier of the "user ID" field. At the same time, users can input a custom rule "transaction amount ≥ 0", which the system associates with the "transaction amount" field, and also associates the user-input custom rule "inventory ≥ 0 and ≤ 100000" with the same field "product inventory". By combining preset rules and custom rules, the system improves rule adaptation efficiency, is suitable for multi-dimensional business scenarios, and supports flexible expansion of complex rules, ensuring that the rule mapping relationship can cover field characteristics and business requirements.

[0077] In some embodiments, the custom rules include at least one of the following:

[0078] Field value format validation rules are used to verify whether field values ​​conform to a preset format;

[0079] Field value range validation rules are used to verify whether a field value is within a specified range.

[0080] The reference data validation rules for field values ​​are used to verify whether a field value exists in a preset reference dataset.

[0081] As you can understand, a field value is the specific content of a row of data corresponding to a column (field) in a Hive table, and it is the smallest unit of data; a field is the attribute definition of data (such as a mobile phone number), and a field value is the specific value of that attribute (such as a user's mobile phone number).

[0082] The format of field values ​​refers to the syntax, structure, or encoding rules that field values ​​must follow. This verifies whether the data conforms to the preset format to avoid problems such as garbled characters, structural errors, and inconsistent formats. These preset formats include, for example, string formats, numeric formats, and date formats. For string formats, mobile phone numbers must be 11 digits, email addresses must contain "@" and end with a domain name, and order numbers must begin with "OD" followed by the date and serial number (e.g., "OD20251210001"). Numeric formats include amounts needing to be rounded to two decimal places and ID numbers needing to be 18 digits (the last digit can be "X"). Date formats must be uniformly formatted as yyyy-MM-dd HH:mm:ss (avoiding the mixing of "2025 / 12 / 10" and "2025-12-10").

[0083] The range of field values ​​refers to the preset numerical intervals, enumeration sets, or logical boundaries that field values ​​must meet. This is used to verify the logical rationality of the data, ensuring that the data conforms to business rules or common sense, thereby filtering out abnormal data that does not conform to business logic (such as negative amounts or non-existent status values) and preventing erroneous data from interfering with analysis results (such as including abnormally large orders when calculating revenue, leading to data distortion). This specified numerical range may include numerical ranges, enumeration ranges, and time ranges. For example, a numerical range may require payment amounts > 0 and ≤ 100,000 (to avoid negative numbers or abnormally large orders), and age ≥ 0 and ≤ 120. An enumeration range may allow payment statuses only ["Unpaid", "Paid", "Refunding in Progress", "Refunded"], and product categories only ["Digital", "Clothing", "Food"]. A time range may require order creation time ≥ system launch time and ≤ current time (to avoid orders from future times).

[0084] Reference data for field values ​​refers to verifying the existence or consistency of field values ​​based on other data sources (such as related tables, benchmark databases, and external system data). This eliminates inconsistencies caused by data silos, ensures data matches external benchmarks or related data, and prevents "isolated data" (such as orders pointing to non-existent users) from affecting business processes (e.g., failed shipments, incorrect user profiles). Reference data verification for field values ​​mainly includes: cross-table join verification, benchmark data verification, and cross-system consistency verification. Cross-table join verification, for example, can be verifying whether the `user_id` in the order table exists in the `user_id` in the user table (ensuring the order belongs to a valid user), and whether the `category_id` in the product table matches the `id` in the category table. Benchmark data verification, for example, can be verifying that the user's mobile phone number matches the number segment database provided by the operator (avoiding invalid number segments), and that the product price references the brand's official pricing range. Cross-system consistency verification, for example, can be verifying that the order amount in the Hive table matches the order amount in the business system.

[0085] In this embodiment, the custom rules cover the multidimensional requirements of field data quality by combining three verification types: format, range, and reference data, thus achieving multidimensional adaptability of rule configuration.

[0086] For example, a format validation rule "Conforms to email format" is defined for the "User Email" field, a range validation rule "≥0 and ≤1000000" is defined for the "Transaction Amount" field, and a reference data validation rule "Value must be 'Male' or 'Female'" is defined for the "Gender" field. By diversifying the rule types, the data quality requirements of different fields can be adapted, significantly improving the flexibility of rule configuration.

[0087] In other embodiments, custom rules can also be generated based on machine learning models.

[0088] The machine learning model was trained using historical monitoring data and manually labeled rule samples.

[0089] For example, the machine learning model generates a new rule "order amount ≥ 0 and ≤ 100000" based on the historical rule sample "transaction amount ≥ 0".

[0090] S205: Based on a distributed computing framework, Hive table data is split into multiple partitions.

[0091] S206: Parallel verification of Hive table data in multiple partitions based on rule mapping relationships.

[0092] Among them, distributed computing frameworks are systems that support parallel computing, such as Spark, Flink, or MapReduce. That is, Spark can be used to perform distributed tasks to validate the "transaction amount" field of a Hive table.

[0093] Understandably, using a distributed computing framework to break down the full Hive table data into multiple partitions that can be processed in parallel, and then performing batch verification on multiple partitions according to rules, can reduce the resource consumption of a single task and improve the overall verification efficiency of Hive table data.

[0094] After splitting the Hive table data, a corresponding verification rule is matched for each split partition. Utilizing the multi-node parallel capability of the distributed framework, each partition can be assigned to different computing nodes of the distributed framework to achieve parallel data processing. This allows for the simultaneous verification of data from multiple partitions, significantly reducing the overall verification time (e.g., reducing full verification from hours to minutes) and avoiding resource exhaustion or timeouts caused by full scans.

[0095] For example, the distributed computing framework Spark splits 100 million transaction records into 10 partitions. Each partition independently performs the "transaction amount ≥ 0" rule verification, and finally aggregates the verification results. This improves verification efficiency through parallel computing and is suitable for large-scale data scenarios.

[0096] In some embodiments, when splitting Hive table data into multiple partitions based on a distributed computing framework, the splitting can be performed in at least one of the following ways:

[0097] Split according to native partitions of Hive table data; split according to data volume; split according to field dimensions.

[0098] For example, splitting according to the native partitions of a Hive table is partitioning by business dimensions, such as dt=date and region=region. Specifically, splitting by dt (day): splitting date partitions such as dt=20251201, dt=20251202, etc., into independent task units; splitting by region: splitting region partitions such as region=North China, region=East China, etc.; splitting by data volume is performed when the Hive table has no native partitions or the data volume of a single partition is too large (such as more than 1 billion rows of data per day). Sharding, specifically, is based on the sharding mechanism of a distributed framework (such as Spark's RDD partitioning or Flink's parallelism), splitting by the number of rows or file size, dividing the data of a single partition into 10 million rows per shard, generating 100 shards for parallel processing; splitting by field dimension involves splitting according to different fields, and performing independent rule validation for different fields (such as phone number format and payment amount range can be validated separately). Specifically, "format fields" and "range fields" are split into different partitions, each matched with corresponding rules, and data validation is performed based on these rules.

[0099] In some embodiments, during the parallel verification of Hive table data in multiple partitions based on rule mapping relationships, computing resources for each partition can be dynamically allocated according to the cluster resource status.

[0100] The cluster resource status refers to the real-time status of available resources in the distributed computing framework, including CPU, memory, and / or storage resources.

[0101] Understandably, distributed computing frameworks dynamically allocate computing resources by real-time monitoring of cluster resource status (such as CPU utilization and memory availability), thereby optimizing the utilization efficiency of computing resources and improving the response speed of core business table verification. For example, high-priority tasks (such as core business table verification) are allocated more resources to ensure fast execution.

[0102] For example, in financial scenarios, transaction table verification tasks can be completed within minutes due to their high resource allocation priority, avoiding delays caused by insufficient resources; or, in the distributed computing framework Spark, if the CPU utilization of the cluster is 30%, more Executors can be dynamically allocated to execute specific data verification tasks.

[0103] In other embodiments, log information can be generated based on the data verification results obtained from performing the data verification task; the log information can be used to locate the problematic parts when performing the data verification task.

[0104] The log information includes abnormal data samples and the execution path of verification rules.

[0105] Understandably, data verification results record abnormal data samples (such as "transaction amount is negative") and the execution path of verification rules (such as "transaction amount ≥ 0" rule being triggered) through log information; by tracing the log chain, data support is provided for problem localization, quickly locating the problem link, thereby significantly improving the accuracy of problem localization and accelerating the closed-loop processing of data quality issues.

[0106] For example, the data quality verification system records the abnormal data sample "Transaction ID=12345" and its corresponding rule execution path for the current verification, for subsequent problem tracing.

[0107] In some embodiments, a data quality report is generated based on log information.

[0108] The data quality report includes statistics on abnormal data, rule execution efficiency, and root cause analysis of problems.

[0109] Understandably, data verification is one of the core aspects of data governance. Continuous rule verification can establish data quality reports, recording data quality status, anomaly types, and processing results. The data quality report summarizes abnormal data statistics (such as "number of records with negative transaction amounts"), rule execution efficiency (such as "verification task took 5 minutes"), and root cause analysis (such as "ETL conversion step error") through log information, forming a complete view of data quality verification results.

[0110] For example, the data quality verification system generates a corresponding data quality report based on the log information obtained in the current instance. Specifically, it may be: "The transaction table's data quality score this month is 92 points, an increase of 3 points compared to last month." This report is used to assist users in optimizing their data governance strategies and forming a virtuous cycle of continuous improvement in data quality.

[0111] In some embodiments, the data verification task can be scheduled in three modes: timed scheduling, event-triggered scheduling (such as triggered after the ETL task is completed), and manual-triggered scheduling; users can choose flexibly according to the real-time requirements of the data.

[0112] For example, for real-time data warehouse Hive tables with high real-time requirements, you can configure "trigger monitoring (verification) immediately after ETL task is completed" to ensure that problems are detected as soon as the data is loaded; for offline data warehouse tables, you can configure "full verification at 2 am every day" to avoid occupying resources during peak business periods and achieve efficient collaboration between monitoring tasks and business processes.

[0113] In other embodiments, precise early warnings that "match the severity of the problem with the speed of response" are achieved by classifying warning levels (urgent, important, general, alert) and using multiple channels for notification (email, SMS, etc.).

[0114] For example, if more than 10% of the key fields in a core business table have data errors, an "emergency" alert is triggered, and the responsible person is notified via both SMS and email, requiring a response within 1 hour. Minor data anomalies in non-core tables trigger a "hint" alert, which is only notified via system message to avoid excessive disruption. At the same time, an alert frequency limit function is also set up to prevent repeated alerts for the same problem, thereby improving the effectiveness of alerts.

[0115] The data verification method provided in this embodiment obtains field identification information and matches it with preset rules, while also receiving user-inputted custom rules. It then obtains a rule mapping relationship based on the preset and custom rules. Using a distributed computing framework, it splits Hive table data into multiple partitions and verifies the Hive table data in multiple partitions in parallel based on the rule mapping relationship. This method, through field-identifier-driven rule configuration, solves the problems of complex rule configuration and difficulty in adapting to changing business needs in existing technologies. It significantly improves the efficiency and accuracy of Hive table data quality verification, enhances the precision of data problem location, and forms a full-link traceability capability from data anomalies to the root cause of the problem. It achieves efficient full-volume verification in large-scale data scenarios, improving not only the comprehensiveness and real-time performance of data quality verification but also the efficiency of full-volume data verification.

[0116] Figure 3 This is a schematic diagram of the data verification device provided in this application. Figure 3 As shown, this application provides a data verification device, the data verification device 300 including:

[0117] The acquisition module 301 is used to acquire field identification information, which is used to describe the attributes of the field.

[0118] The processing module 302 is used to generate rule mapping relationships based on field identification information. The rule mapping relationships include preset rules and custom rules. The preset rules include uniformly configured rules, and the custom rules include user-defined rules.

[0119] The processing module 302 is also used to perform data verification tasks based on rule mapping relationships. The data verification tasks include rule verification of Hive table data.

[0120] Optionally, the processing module 302 is also used to match preset rules based on field identification information, the preset rules including the validation logic corresponding to the field type.

[0121] The acquisition module 301 is also used to receive user input of custom rules, which include field validation conditions and validation thresholds.

[0122] The processing module 302 is also used to obtain the rule mapping relationship based on preset rules and custom rules.

[0123] Optionally, the processing module 302 is also used to split the Hive table data into multiple partitions based on a distributed computing framework; the distributed computing framework is a framework that supports parallel computing.

[0124] The processing module 302 is also used to perform parallel verification of Hive table data in multiple partitions based on rule mapping relationships.

[0125] Optionally, the processing module 302 is also used to dynamically allocate computing resources for each partition according to the cluster resource status. The cluster resource status is the real-time status of available resources in the distributed computing framework, and the computing resources include CPU, memory and / or storage resources.

[0126] Optionally, the processing module 302 is also used to generate log information based on the data verification results obtained from the execution of the data verification task. The log information includes abnormal data samples and the execution path of the verification rules.

[0127] The processing module 302 is also used to locate the problematic parts when performing data verification tasks based on log information.

[0128] Optionally, the processing module 302 is also used to generate a data quality report based on log information. The data quality report includes statistics on abnormal data, rule execution efficiency, and root cause analysis of problems.

[0129] Figure 4 A schematic diagram of the structure of the electronic device provided in this application. Figure 4 As shown, this application provides an electronic device 400, which includes a receiver 401, a transmitter 402, a processor 403, and a memory 404.

[0130] Receiver 401 is used to receive instructions and data;

[0131] Transmitter 402 is used to send commands and data;

[0132] Memory 404 is used to store instructions executed by the computer;

[0133] Processor 403 is used to execute computer execution instructions stored in memory 404 to implement the various steps of the data verification method in the above embodiments. For details, please refer to the relevant descriptions in the foregoing data verification method embodiments.

[0134] Optionally, the memory 404 can be either standalone or integrated with the processor 403.

[0135] When the memory 404 is set up independently, the electronic device also includes a bus for connecting the memory 404 and the processor 403.

[0136] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the data verification method performed by the aforementioned electronic device.

[0137] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the data verification method of any of the foregoing embodiments.

[0138] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0139] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0140] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.

[0141] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.

[0142] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.

[0143] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0144] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

[0145] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0146] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A data verification method, characterized in that, The method includes: Obtaining field identification information, where the field identification information is used to describe the attributes of a field; Generating a rule mapping relationship based on the field identification information, where the rule mapping relationship includes preset rules and custom rules, the preset rules include rules configured uniformly, and the custom rules include rules set by users; Performing a data verification task based on the rule mapping relationship, where the data verification task includes rule verification of Hive table data.

2. The method according to claim 1, characterized in that, The generating a rule mapping relationship based on the field identification information includes: Matching the preset rules according to the field identification information, where the preset rules include verification logics corresponding to field types; Receiving the custom rules input by the user, where the custom rules include field verification conditions and verification thresholds; Obtaining the rule mapping relationship according to the preset rules and the custom rules.

3. The method according to claim 2, characterized in that, The custom rules include at least one of the following: A format verification rule for field values, used to verify whether a field value conforms to a preset format; A range verification rule for field values, used to verify whether a field value is within a specified numerical range; A reference data verification rule for field values, used to verify whether a field value exists in a preset reference dataset.

4. The method according to any one of claims 1-3, characterized in that, The performing a data verification task based on the rule mapping relationship includes: Splitting the Hive table data into multiple partitions based on a distributed computing framework; the distributed computing framework is a framework that supports parallel computing; Parallelly verifying the Hive table data in the multiple partitions based on the rule mapping relationship.

5. The method according to claim 4, characterized in that, During the process of parallelly verifying the Hive table data in the multiple partitions based on the rule mapping relationship, the following steps are also performed: Dynamically allocating computing resources for each of the partitions according to the cluster resource status, where the cluster resource status is the real-time status of available resources in the distributed computing framework, and the computing resources include CPU, memory, and / or storage resources.

6. The method according to any one of claims 1-3, characterized in that, The method further includes: Generating log information based on the data verification result obtained by performing the data verification task, where the log information includes abnormal data samples and verification rule execution paths; Locating the link with problems when performing the data verification task based on the log information.

7. The method according to claim 6, characterized in that, The method further includes: Generating a data quality report based on the log information, where the data quality report includes abnormal data statistics, rule execution efficiency, and root cause analysis of problems.

8. A data verification device, including: An obtaining module, used to obtain field identification information, where the field identification information is used to describe the attributes of a field; A processing module, used to generate a rule mapping relationship based on the field identification information, where the rule mapping relationship includes preset rules and custom rules, the preset rules include rules configured uniformly, and the custom rules include rules set by users; The processing module is further used to perform a data verification task based on the rule mapping relationship, where the data verification task includes rule verification of Hive table data.

9. An electronic device, characterized in that, Including: A processor, and a memory communicatively connected to the processor; The memory stores computer execution instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.