Data verification method, device, system and storage medium

By initializing big data and tag comparison in the environment of Spark 2.0 or above, the problem of low efficiency of big data verification and inability to be applicable to complex data structures in the existing technology is solved, and efficient data verification and scalability are achieved.

CN112783855BActive Publication Date: 2025-05-16BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201911081372.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-11-07
Publication Date
2025-05-16
Estimated Expiration
2039-11-07

AI Technical Summary

Technical Problem

The prior art is inefficient in verifying the results of big data processing and cannot be used for the verification of complex data structures.

Method used

In the application environment of Spark 2.0 or above, Spark Session is constructed to initialize the source data and the data to be verified, and the data frames are labeled and compared according to the loaded template rules to obtain the data verification results.

Benefits of technology

It realizes effective verification of complex data structures, improves data verification efficiency, and has the characteristics of strong scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112783855B_ABST
    Figure CN112783855B_ABST
Patent Text Reader

Abstract

The present application provides a data verification method, device, system and storage medium, the method comprising: respectively initializing source data and data to be verified to obtain source data frame and data frame to be verified; respectively extracting labels from the source data frame and data frame to be verified according to the loaded template rule to obtain label data corresponding to the source data and label data of the data to be verified; comparing the label data corresponding to the source data with the label data of the data to be verified to obtain a data verification result. The present application is applicable to the verification processing of complex data structures, can effectively improve the verification efficiency of data, and has strong scalability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data technology, and in particular to a data verification method, device, system and storage medium. Background Art

[0002] With the continuous development of big data processing technology, more and more scenarios involve the processing of massive data. In order to ensure the accuracy of the data, the results of big data processing need to be verified.

[0003] At present, the common practice is to use big data calculation result comparison tools to generate Hive SQL use cases, query the calculation results through the MapReduce calculation engine, and compare the calculation results pairwise through left joins.

[0004] However, the above verification method is inefficient and cannot be applied to the verification of complex data structures. Summary of the invention

[0005] The present application provides a data verification method, device, system and storage medium, which can be applicable to the verification processing of complex data structures, can effectively improve the verification efficiency of data, and has strong scalability.

[0006] In a first aspect, an embodiment of the present application provides a data verification method, including:

[0007] Initialize the source data and the data to be verified respectively to obtain a source data frame and a data frame to be verified;

[0008] According to the loaded template rule, respectively extract labels from the source data frame and the data frame to be verified to obtain label data corresponding to the source data and label data of the data to be verified;

[0009] The label data corresponding to the source data is compared with the label data of the data to be verified to obtain a data verification result.

[0010] In a possible design, respectively initializing the source data and the data to be verified to obtain a source data frame and a data frame to be verified includes:

[0011] In an application environment with Spark 2.0 or higher installed, the source data and the data to be verified are initialized by constructing a Spark Session to obtain the source data frame and the data frame to be verified.

[0012] In a possible design, the template rule includes: a label vocabulary of a source table, a label vocabulary of a table to be verified, and a mapping relationship between the label vocabulary of the source table and the label vocabulary of the table to be verified;

[0013] The step of extracting labels from the source data frame and the data frame to be verified respectively according to the loaded template rule includes:

[0014] Extracting label data corresponding to the source data from the source data frame according to the label vocabulary of the source table;

[0015] Extracting label data of the data to be verified from the data to be verified according to the label vocabulary of the table to be verified;

[0016] The tag data corresponding to the source data is compared with the tag data of the data to be verified to obtain a data verification result, including:

[0017] According to the mapping relationship between the label word list of the source table and the label word list of the table to be verified, the label data corresponding to the source data is compared with the label data of the data to be verified to obtain a data verification result.

[0018] In a possible design, extracting label data corresponding to the source data from the source data frame according to the label vocabulary of the source table includes:

[0019] Extracting label data of all users from the source data frame according to the label vocabulary of the source table;

[0020] According to the user identification and label category, the labels of all users are classified and deduplicated to obtain the label data corresponding to the source data frame.

[0021] In a possible design, the label data corresponding to the source data includes: a first label corresponding to the same user, and a statistical value corresponding to each first label;

[0022] The label data of the data to be verified includes: second labels corresponding to the same user, and statistical values ​​corresponding to each second label.

[0023] In a possible design, comparing the label data corresponding to the source data with the label data of the data to be verified to obtain a data verification result includes:

[0024] According to a mapping relationship between a label vocabulary of a source table and a label vocabulary of a table to be verified, searching for the first label corresponding to the second label;

[0025] If the first tag corresponding to the second tag is not found, it is determined that the data verification fails, and the failure details are marked as the first category;

[0026] If the first tag corresponding to the second tag is found, comparing whether the statistical value corresponding to the second tag is the same as the statistical value corresponding to the first tag;

[0027] If the statistical value corresponding to the second tag is different from the statistical value corresponding to the first tag, it is determined that the data verification fails, and the failure details are marked as the second category;

[0028] If the first tag corresponding to the second tag is found, and the statistical value corresponding to the second tag is the same as the statistical value corresponding to the first tag, it is determined that the data verification is successful.

[0029] In one possible design, it also includes:

[0030] Generate a log file according to the data verification result; the log file contains the record data of all data frames marked as failure details;

[0031] The log file is pushed to the tester's terminal.

[0032] In a second aspect, an embodiment of the present application provides a data verification device, including:

[0033] An initialization module is used to initialize the source data and the data to be verified respectively to obtain a source data frame and a data frame to be verified;

[0034] An extraction module, used to extract labels from the source data frame and the data frame to be verified respectively according to the loaded template rule, to obtain label data corresponding to the source data and label data of the data to be verified;

[0035] The comparison module is used to compare the label data corresponding to the source data with the label data of the data to be verified to obtain a data verification result.

[0036] In a possible design, the processing module is specifically used to:

[0037] In an application environment with Spark 2.0 or higher installed, the source data and the data to be verified are initialized by constructing a Spark Session to obtain the source data frame and the data frame to be verified.

[0038] In a possible design, the template rule includes: a label vocabulary of a source table, a label vocabulary of a table to be verified, and a mapping relationship between the label vocabulary of the source table and the label vocabulary of the table to be verified;

[0039] Wherein, the extraction module is specifically used for:

[0040] Extracting label data corresponding to the source data from the source data frame according to the label vocabulary of the source table;

[0041] Extracting label data of the data to be verified from the data to be verified according to the label vocabulary of the table to be verified;

[0042] Wherein, the comparison module is specifically used for:

[0043] According to the mapping relationship between the label word list of the source table and the label word list of the table to be verified, the label data corresponding to the source data is compared with the label data of the data to be verified to obtain a data verification result.

[0044] In a possible design, the extraction module is further used to:

[0045] Extracting label data of all users from the source data frame according to the label vocabulary of the source table;

[0046] According to the user identification and label category, the labels of all users are classified and deduplicated to obtain the label data corresponding to the source data frame.

[0047] In a possible design, the label data corresponding to the source data includes: a first label corresponding to the same user, and a statistical value corresponding to each first label;

[0048] The label data of the data to be verified includes: second labels corresponding to the same user, and statistical values ​​corresponding to each second label.

[0049] In a possible design, the comparison module is specifically used for:

[0050] According to a mapping relationship between a label vocabulary of a source table and a label vocabulary of a table to be verified, searching for the first label corresponding to the second label;

[0051] If the first tag corresponding to the second tag is not found, it is determined that the data verification fails, and the failure details are marked as the first category;

[0052] If the first tag corresponding to the second tag is found, comparing whether the statistical value corresponding to the second tag is the same as the statistical value corresponding to the first tag;

[0053] If the statistical value corresponding to the second tag is different from the statistical value corresponding to the first tag, it is determined that the data verification fails, and the failure details are marked as the second category;

[0054] If the first tag corresponding to the second tag is found, and the statistical value corresponding to the second tag is the same as the statistical value corresponding to the first tag, it is determined that the data verification is successful.

[0055] In a possible design, the system further includes: a log generation module, configured to:

[0056] Generate a log file according to the data verification result; the log file contains the record data of all data frames marked as failure details;

[0057] The log file is pushed to the tester's terminal.

[0058] In a third aspect, an embodiment of the present application provides a data verification system, comprising: a memory and a processor, wherein the memory stores executable instructions of the processor; wherein the processor is configured to execute the data verification method described in any one of the first aspects by executing the executable instructions.

[0059] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data verification method described in any one of the first aspects.

[0060] In a fifth aspect, an embodiment of the present application provides a program product, comprising: a computer program, wherein the computer program is stored in a readable storage medium, at least one processor of a server can read the computer program from the readable storage medium, and the at least one processor executes the computer program so that the server executes any data verification method described in the first aspect.

[0061] The present application provides a data verification method, device, system and storage medium, which respectively initialize the source data and the data to be verified to obtain the source data frame and the data frame to be verified; according to the loaded template rule, the source data frame and the data frame to be verified are respectively subjected to label extraction to obtain the label data corresponding to the source data and the label data of the data to be verified; the label data corresponding to the source data is compared with the label data of the data to be verified to obtain the data verification result. The present application is applicable to the verification processing of complex data structures, can effectively improve the data verification efficiency, and has strong scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0063] Figure 1 This is a schematic diagram of the principle of an application scenario of this application;

[0064] Figure 2 A flowchart of a data verification method provided in Example 1 of the present application;

[0065] Figure 3 A flowchart of a data processing method provided in Embodiment 2 of the present application;

[0066] Figure 4 A schematic diagram of the structure of a data verification device provided in Embodiment 3 of the present application;

[0067] Figure 5 A schematic diagram of the structure of a data verification device provided in Embodiment 4 of the present application;

[0068] Figure 6 This is a structural diagram of the data verification system provided in Example 5 of the present application.

[0069] The above drawings show clear embodiments of the present disclosure, which will be described in more detail below. These drawings and text descriptions are not intended to limit the scope of the present disclosure in any way, but to illustrate the concepts of the present disclosure to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0070] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0071] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0072] The technical solution of the present application is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0073] With the continuous development of big data processing technology, more and more scenarios involve the processing of massive data. Based on the underlying global user basic data of a certain platform, its obvious characteristics are: Hive data warehouse storage, a large number of entries (1.1 billion / day), a large number of tag types and continuous expansion (152 types), complex data storage (using Map to compress and store tag data, the Key of Map is the tag, and the value is the tag statistical value [the number of times the user generates the behavior described by the tag]), a large number of source tables and a large amount of data (currently 21 underlying tables, the source bottom table of advertising exposure has the largest amount of data, at the daily level of 2 billion). In order to ensure the accuracy of the data, it is necessary to verify the results of big data processing.

[0074] At present, the common practice is to use big data calculation result comparison tools to generate Hive SQL use cases, query the calculation results through the MapReduce calculation engine, and compare the calculation results pairwise through left joins.

[0075] However, the above verification method is only applicable to the comparison between two tables, and the data retrieval SQL logic of the two tables should not be complicated, otherwise the SQL length and template maintenance difficulty will be very large; the verification support for results containing complex data structures is not friendly enough, such as Map data structures. Hive lacks a built-in method for looping out all key and value values ​​in the Map structure. If the subscript values ​​are enumerated one by one, and the label statistics are obtained according to the subscript and compared with the original table, there will be a certain implementation cost. Therefore, this method is inefficient and cannot be applied to the verification of complex data structures.

[0076] In response to the above technical problems, the present application provides a data verification method, device, system and storage medium, which can be applicable to the verification processing of complex data structures, can effectively improve the verification efficiency of data, and has strong scalability. Figure 1 This is a schematic diagram of the principle of an application scenario of this application, such as Figure 1As shown, in Spark 2.0 and above, Spark uses Spark Session to load, convert and process data. It supports loading data from Hive data source and converting it into data frame, and then SQL statements can be used to operate data. Therefore, the resource initialization step is to construct Spark Session, add HiveSupport, set the application name, set general operation parameters, and further implement the loading, conversion and processing of source data and data to be verified to achieve source data frame and data frame to be verified. Template rules are stored in the form of configuration files with pre-set source table label word list, to-be-verified table label word list and mapping relationship between the two. The mapping relationship can be a one-to-many mapping relationship. After the template rules are loaded, they can be parsed by the parser to generate source table label extraction rules, to-be-verified table label extraction rules, and source table and to-be-verified table label mapping rules. Then data collection is performed through the data collector, which includes source table data collector and to-be-verified table data collector. The source table data collector extracts the label data corresponding to the source data from the source data frame according to the label vocabulary of the source table, including: extracting the label data of all users from the source data frame according to the label vocabulary of the source table; classifying and deduplicating the labels of all users according to the user identification and label category to obtain the label data corresponding to the source data frame. The data collector of the table to be verified extracts the label data of the data to be verified from the data to be verified according to the label vocabulary of the table to be verified. Then, the comparison engine can compare the label data corresponding to the source data with the label data of the data to be verified according to the mapping relationship between the label vocabulary of the source table and the label vocabulary of the table to be verified to obtain the data verification result. In the specific implementation process, the first label corresponding to the second label can be searched according to the mapping relationship between the label vocabulary of the source table and the label vocabulary of the table to be verified; if the first label corresponding to the second label is not found, it is determined that the data verification fails, and the failure details are marked as the first category. The first category refers to the data content corresponding to the data to be verified that does not exist in the source data. If the first tag corresponding to the second tag is found, the statistical value corresponding to the second tag is compared to see if it is the same as the statistical value corresponding to the first tag; if the statistical value corresponding to the second tag is different from the statistical value corresponding to the first tag, it is determined that the data verification has failed, and the failure details are marked as the second category. The second category refers to the inconsistency between the data content corresponding to the source data and the data to be verified. If the first tag corresponding to the second tag is found, and the statistical value corresponding to the second tag is the same as the statistical value corresponding to the first tag, it is determined that the data verification is successful. Finally, a log file can be generated based on the data verification results, and the log file contains the record data of all data frames marked as failure details. Then, the log file is pushed to the tester's terminal.

[0077] The above method is applicable to the verification processing of complex data structures, which can effectively improve the verification efficiency of data and has strong scalability.

[0078] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0079] Figure 2 This is a flow chart of the data verification method provided in Example 1 of the present application, such as Figure 2 As shown, the method in this embodiment may include:

[0080] S101 , respectively initializing source data and data to be verified to obtain a source data frame and a data frame to be verified.

[0081] In this embodiment, in an application environment with Spark 2.0 or higher installed, the constructed SparkSession can be used to perform initialization processing on the source data and the data to be verified to obtain the source data frame and the data frame to be verified.

[0082] Specifically, in Spark version 2.0 and above, Spark uses Spark Session to load, convert, and process data. It supports loading data from Hive data sources and converting them into data frames, and then SQL statements can be used to operate the data. Therefore, the initialization steps of this process are to construct Spark Session, add Hive Support, set the application name, set general operating parameters, and further implement the loading, conversion, and processing of source data and data to be verified to reach the source data frame and data frame to be verified.

[0083] S102 . Extract labels from source data frames and data frames to be verified respectively according to the loaded template rules to obtain label data corresponding to the source data and label data corresponding to the data to be verified.

[0084] In this embodiment, the template rules include: a label word list of the source table, a label word list of the table to be verified, and a mapping relationship between the label word list of the source table and the label word list of the table to be verified. According to the loaded template rules, according to the label word list of the source table, the label data corresponding to the source data is extracted from the source data frame; according to the label word list of the table to be verified, the label data of the data to be verified is extracted from the data to be verified.

[0085] Specifically, the template rule stores the preset source table label word list, the to-be-verified table label word list, and the mapping relationship between the two in the form of a configuration file. The mapping relationship can be a one-to-many mapping relationship. After the template rule is loaded, it can be parsed by a parser to generate source table label extraction rules, to-be-verified table label extraction rules, and source table and to-be-verified table label mapping rules.

[0086] Optionally, the label data corresponding to the source data includes: a first label corresponding to the same user, and a statistical value corresponding to each first label. The first label is used to indicate the label details extracted from the source data, and the number can be multiple. The label data of the data to be verified includes: a second label corresponding to the same user, and a statistical value corresponding to each second label. The second label is used to indicate the label details extracted from the data to be verified, and the number can be multiple.

[0087] Optionally, according to the label vocabulary of the source table, label data corresponding to the source data is extracted from the source data frame, including: according to the label vocabulary of the source table, label data of all users is extracted from the source data frame; according to the user identification and label category, the labels of all users are classified and deduplicated to obtain the label data corresponding to the source data frame.

[0088] Specifically, you can configure the source base table, the fields to be extracted, the partitions, whether repartitioning is needed (to improve performance and increase parallelism), and the fields based on which repartitioning is done. Then, parse the configuration file, call the general parameterized data collector, and obtain the source table label details. Then, combine (union) the label details of each source table. Finally, aggregate the data frame after the union again: deduplicate according to the user dimension, accumulate each label dimension, and the final label source data frame structure is user / label 1 / label 2 / … / label N (label X represents the statistical value of each label), so that one user corresponds to one piece of data, achieving the purpose of data deduplication.

[0089] According to the label word list of the table to be verified, the label data of the data to be verified is extracted from the data to be verified, which specifically includes: configuring the table to be verified, the fields to be extracted, and the partitions. Then, the configuration file is parsed, the general parameterized data collector is called, the label statistical data information is obtained, and the label data of the data to be verified is generated. The data structure is user / label Map ({label 1: label 1 statistical value, label 2: label 2 statistical value, ..., label N: label N statistical value}).

[0090] S103: Compare the label data corresponding to the source data with the label data of the data to be verified to obtain a data verification result.

[0091] In this embodiment, the label data corresponding to the source data can be compared with the label data of the data to be verified according to the mapping relationship between the label word list of the source table and the label word list of the table to be verified to obtain the data verification result. In the specific implementation process, the first label corresponding to the second label can be found according to the mapping relationship between the label word list of the source table and the label word list of the table to be verified; if the first label corresponding to the second label is not found, it is determined that the data verification fails, and the details of the failure are marked as the first category. The first category refers to the fact that there is no data content corresponding to the data to be verified in the source data. If the first label corresponding to the second label is found, the statistical value corresponding to the second label is compared to see whether it is the same as the statistical value corresponding to the first label; if the statistical value corresponding to the second label is different from the statistical value corresponding to the first label, it is determined that the data verification fails, and the details of the failure are marked as the second category. The second category refers to the inconsistency of the data content corresponding to the source data and the data to be verified. If the first label corresponding to the second label is found, and the statistical value corresponding to the second label is the same as the statistical value corresponding to the first label, it is determined that the data verification is successful.

[0092] Specifically, according to the loaded template rules, the detailed data of each label of the source data frame is mapped to the label of the table to be verified through the mapping relationship between the label vocabulary of the source table and the label vocabulary of the table to be verified, and stored in the form of key-value pairs ({label 1: label 1 statistical value, label 2: label 2 statistical value, ..., label N: label N statistical value}), the primary key is the label to be verified, and the value is the cumulative value of the label statistical value. Then, the source table and the table to be verified are left-connected, associated through the user pin, and the label statistical Map data of the source and the table to be verified are taken out for comparison. The comparison scheme is as follows:

[0093] a) Traverse the data of each user's pin, add the "Failed" field, the default value is False, add the "Failure Details" field, the default value is "";

[0094] b) If the tag statistics map of the tag source data frame user pin granularity is not empty and the tag statistics map of the data frame to be verified is empty, then the data is marked with a failure tag (set "failure" to True), and the failure details are marked as the first category, for example, the "failure details" field is recorded as "the tag statistics map to be compared is empty";

[0095] c) If the tag statistics Map of the tag source data frame and the data frame to be verified are not empty, traverse the source tag statistics Map, and obtain the statistical values ​​of the tag of the source data frame and the data frame to be verified according to the primary key value (tag name). If the two are not equal, mark the failure tag (set "failure" to True), and mark the failure details as the second category, for example, the "failure details" field adds "inconsistent tag XX, source data frame statistical value XX, and data frame statistical value XX" in the form of append.

[0096] d) If the first tag corresponding to the second tag is found, and the statistical value corresponding to the second tag is the same as the statistical value corresponding to the first tag, then it is determined that the data verification is successful. Continue to traverse the next user until the traversal is completed and exit the program.

[0097] You can also save the results to a new data frame in the format of user pin / tag source tag statistics map / to-be-verified table tag statistics map / failed / failure details.

[0098] Experimental results show that the technical solution of this application can complete the data comparison of 1.1 billion people + 21 underlying source tables and the table to be verified within hours, reducing the execution cost by 30% compared to Hive SQL on Map Reduce. After being configured as a sustainable integration task, snapshot data at different times can be verified without supervision. In subsequent version upgrades, the code itself does not need to be changed, only the rule template and the base table configuration file need to be modified to complete the subsequent result verification task. If the data structure is converted from Map to Struct, the comparison method in the comparison engine can be modified for compatibility.

[0099] In this embodiment, the source data and the data to be verified are respectively initialized to obtain the source data frame and the data to be verified frame; according to the loaded template rules, the source data frame and the data to be verified frame are respectively extracted to obtain the label data corresponding to the source data and the label data of the data to be verified; the label data corresponding to the source data is compared with the label data of the data to be verified to obtain the data verification result. This application is suitable for the verification processing of complex data structures, which can effectively improve the data verification efficiency and has strong scalability.

[0100] Figure 3 This is a flow chart of a data processing method provided in Example 2 of the present application, such as Figure 3 As shown, the method in this embodiment may include:

[0101] S201 , respectively initialize the source data and the data to be verified to obtain a source data frame and a data frame to be verified.

[0102] S202 . Extract labels from the source data frame and the data frame to be verified respectively according to the loaded template rule to obtain label data corresponding to the source data and label data corresponding to the data to be verified.

[0103] S203: Compare the label data corresponding to the source data with the label data of the data to be verified to obtain a data verification result.

[0104] In this embodiment, the specific implementation process and technical principles of steps S201 to S203 are shown in Figure 2The relevant descriptions of steps S101 to S103 in the method shown are not repeated here.

[0105] S204: Generate a log file based on the data verification result, and push the log file to the tester's terminal.

[0106] In this embodiment, a log file may be generated according to the data verification result, wherein the log file contains the recorded data of all data frames marked as failure details. Then, the log file is pushed to the terminal of the tester.

[0107] Specifically, you can filter out the records marked with failure labels (records with the "failed" field set to True, it is recommended to limit the number of output records) from the data frame of the data verification result and output them to the log file. Then, call the email sending module to send the failure report and notify relevant personnel of the test results.

[0108] In this embodiment, the source data and the data to be verified are respectively initialized to obtain the source data frame and the data to be verified frame; according to the loaded template rules, the source data frame and the data to be verified frame are respectively extracted to obtain the label data corresponding to the source data and the label data of the data to be verified; the label data corresponding to the source data is compared with the label data of the data to be verified to obtain the data verification result. This application is suitable for the verification processing of complex data structures, which can effectively improve the data verification efficiency and has strong scalability.

[0109] In addition, this embodiment can also generate a log file according to the data verification result, and push the log file to the terminal of the tester, thereby effectively improving the data verification efficiency and having strong scalability.

[0110] Figure 4 This is a schematic diagram of the structure of the data verification device provided in Example 3 of the present application, such as Figure 4 As shown, the data verification device of this embodiment may include:

[0111] An initialization module 31 is used to initialize the source data and the data to be verified respectively to obtain a source data frame and a data frame to be verified;

[0112] The extraction module 32 is used to extract labels from the source data frame and the data frame to be verified according to the loaded template rules, and obtain label data corresponding to the source data and label data of the data to be verified;

[0113] The comparison module 33 is used to compare the label data corresponding to the source data with the label data of the data to be verified to obtain a data verification result.

[0114] In a possible design, the initialization module 31 is specifically used for:

[0115] In an application environment with Spark 2.0 or higher installed, the constructed Spark Session is used to perform initialization processing on the source data and the data to be verified, thereby obtaining the source data frame and the data frame to be verified.

[0116] In a possible design, the template rule includes: a label vocabulary of a source table, a label vocabulary of a table to be verified, and a mapping relationship between the label vocabulary of the source table and the label vocabulary of the table to be verified;

[0117] The extraction module 32 is specifically used for:

[0118] According to the label vocabulary of the source table, the label data corresponding to the source data is extracted from the source data frame;

[0119] Extracting label data of the data to be verified from the data to be verified according to the label vocabulary of the table to be verified;

[0120] Wherein, the comparison module 33 is specifically used for:

[0121] According to the mapping relationship between the label word list of the source table and the label word list of the table to be verified, the label data corresponding to the source data is compared with the label data of the data to be verified to obtain the data verification result.

[0122] In a possible design, the extraction module 32 is further configured to:

[0123] According to the label vocabulary of the source table, extract the label data of all users from the source data frame;

[0124] According to user identification and tag category, all users' tags are classified and deduplicated to obtain the tag data corresponding to the source data frame.

[0125] In a possible design, the label data corresponding to the source data includes: first labels corresponding to the same user, and statistical values ​​corresponding to each first label;

[0126] The label data of the data to be verified includes: second labels corresponding to the same user, and statistical values ​​corresponding to each second label.

[0127] In a possible design, the comparison module 33 is specifically used for:

[0128] According to the mapping relationship between the label vocabulary of the source table and the label vocabulary of the table to be verified, searching for the first label corresponding to the second label;

[0129] If the first tag corresponding to the second tag is not found, it is determined that the data verification has failed, and the failure details are marked as the first category;

[0130] If a first tag corresponding to the second tag is found, then comparing whether the statistical value corresponding to the second tag is the same as the statistical value corresponding to the first tag;

[0131] If the statistical value corresponding to the second label is different from the statistical value corresponding to the first label, it is determined that the data verification fails, and the failure details are marked as the second category;

[0132] If the first tag corresponding to the second tag is found, and the statistical value corresponding to the second tag is the same as the statistical value corresponding to the first tag, it is determined that the data verification is successful.

[0133] The data verification device of this embodiment can execute Figure 2 For the technical solution in the method shown, its specific implementation process and technical principles can be found in Figure 2 The relevant descriptions in the method shown will not be repeated here.

[0134] In this embodiment, the source data and the data to be verified are respectively initialized to obtain the source data frame and the data to be verified frame; according to the loaded template rules, the source data frame and the data to be verified frame are respectively extracted to obtain the label data corresponding to the source data and the label data of the data to be verified; the label data corresponding to the source data is compared with the label data of the data to be verified to obtain the data verification result. This application is suitable for the verification processing of complex data structures, which can effectively improve the data verification efficiency and has strong scalability.

[0135] Figure 5 A schematic diagram of the structure of the data verification device provided in the fourth embodiment of the present application is shown in FIG. Figure 5 As shown, the data verification device of this embodiment is Figure 4 Based on the device shown, it can also include:

[0136] The log generation module 34 is used to:

[0137] Generate a log file based on the data verification result; the log file contains the record data of all data frames marked as failure details;

[0138] Push log files to the tester's terminal.

[0139] The data verification device of this embodiment can execute Figure 2 , Figure 3 For the technical solution in the method shown, its specific implementation process and technical principles can be found in Figure 2 , Figure 3 The relevant descriptions in the method shown will not be repeated here.

[0140] In this embodiment, the source data and the data to be verified are respectively initialized to obtain the source data frame and the data to be verified frame; according to the loaded template rules, the source data frame and the data to be verified frame are respectively extracted to obtain the label data corresponding to the source data and the label data of the data to be verified; the label data corresponding to the source data is compared with the label data of the data to be verified to obtain the data verification result. This application is suitable for the verification processing of complex data structures, which can effectively improve the data verification efficiency and has strong scalability.

[0141] In addition, this embodiment can also generate a log file according to the data verification result, and push the log file to the terminal of the tester, thereby effectively improving the data verification efficiency and having strong scalability.

[0142] Figure 6 This is a schematic diagram of the structure of the data verification system provided in Example 5 of the present application, such as Figure 6 As shown, the data verification system 40 of this embodiment may include: a processor 41 and a memory 42 .

[0143] The memory 42 is used to store programs; the memory 42 may include volatile memory (English: volatile memory), such as random-access memory (English: random-access memory, abbreviated: RAM), such as static random-access memory (English: static random-access memory, abbreviated: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviated: DDR SDRAM), etc.; the memory may also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory 42 is used to store computer programs (such as applications, functional modules, etc. that implement the above method), computer instructions, etc., and the above computer programs, computer instructions, etc. can be partitioned and stored in one or more memories 42. And the above computer programs, computer instructions, data, etc. can be called by the processor 41.

[0144] The above-mentioned computer programs, computer instructions, etc. may be stored in one or more memories 42 in partitions. And the above-mentioned computer programs, computer instructions, data, etc. may be called by the processor 41.

[0145] The processor 41 is used to execute the computer program stored in the memory 42 to implement the various steps in the method involved in the above embodiment.

[0146] For details, please refer to the relevant description in the previous method embodiment.

[0147] The processor 41 and the memory 42 may be independent structures or integrated structures. When the processor 41 and the memory 42 are independent structures, the memory 42 and the processor 41 may be coupled and connected via a bus 43 .

[0148] The data verification system of this embodiment can execute Figure 2 , Figure 3 For the technical solution in the method shown, its specific implementation process and technical principles can be found in Figure 2 , Figure 3 The relevant descriptions in the method shown will not be repeated here.

[0149] In addition, an embodiment of the present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When at least one processor of the user device executes the computer-executable instructions, the user device executes the above-mentioned various possible methods.

[0150] Among them, computer-readable media include computer storage media and communication media, wherein the communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general or special-purpose computer. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in a user device. Of course, the processor and the storage medium can also be present in a communication device as discrete components.

[0151] The present application also provides a program product, which includes a computer program, which is stored in a readable storage medium. At least one processor of the server can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the server implements any method of the above-mentioned embodiments of the present application.

[0152] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A data verification method, characterized in that: include: Initialize the source data and the data to be verified respectively to obtain a source data frame and a data frame to be verified; According to the loaded template rule, respectively extract labels from the source data frame and the data frame to be verified to obtain label data corresponding to the source data and label data of the data to be verified; Compare the label data corresponding to the source data with the label data of the data to be verified to obtain a data verification result; The template rule includes: a label word list of a source table, a label word list of a table to be verified, and a mapping relationship between the label word list of the source table and the label word list of the table to be verified; The step of extracting labels from the source data frame and the data frame to be verified respectively according to the loaded template rule includes: Extracting label data corresponding to the source data from the source data frame according to the label vocabulary of the source table; Extracting label data of the data to be verified from the data to be verified according to the label vocabulary of the table to be verified; The tag data corresponding to the source data is compared with the tag data of the data to be verified to obtain a data verification result, including: According to the mapping relationship between the label word list of the source table and the label word list of the table to be verified, the label data corresponding to the source data is compared with the label data of the data to be verified to obtain a data verification result.

2. The method according to claim 1, characterized in that: The initialization processing is performed on the source data and the data to be verified respectively to obtain the source data frame and the data frame to be verified, including: In an application environment where Spark uses Spark Session to implement data loading, conversion, and processing, the source data and the data to be verified are initialized through the constructed Spark Session to obtain the source data frame and the data frame to be verified.

3. The method according to claim 1, characterized in that Extracting label data corresponding to the source data from the source data frame according to the label vocabulary of the source table includes: Extracting label data of all users from the source data frame according to the label vocabulary of the source table; According to the user identification and label category, the labels of all users are classified and deduplicated to obtain the label data corresponding to the source data frame.

4. The method according to claim 1, characterized in that: The label data corresponding to the source data includes: a first label corresponding to the same user, and a statistical value corresponding to each first label; The label data of the data to be verified includes: second labels corresponding to the same user, and statistical values ​​corresponding to each second label.

5. The method according to claim 4, characterized in that Comparing the label data corresponding to the source data with the label data of the data to be verified to obtain a data verification result, including: According to a mapping relationship between a label vocabulary of a source table and a label vocabulary of a table to be verified, searching for the first label corresponding to the second label; If the first tag corresponding to the second tag is not found, it is determined that the data verification fails, and the failure details are marked as the first category; If the first tag corresponding to the second tag is found, comparing whether the statistical value corresponding to the second tag is the same as the statistical value corresponding to the first tag; If the statistical value corresponding to the second tag is different from the statistical value corresponding to the first tag, it is determined that the data verification fails, and the failure details are marked as the second category; If the first tag corresponding to the second tag is found, and the statistical value corresponding to the second tag is the same as the statistical value corresponding to the first tag, it is determined that the data verification is successful.

6. The method according to any one of claims 1 to 5, characterized in that Also includes: Generate a log file according to the data verification result; the log file contains the record data of all data frames marked as failure details; The log file is pushed to the tester's terminal.

7. A data processing device, characterized in that: include: An initialization module is used to initialize the source data and the data to be verified respectively to obtain a source data frame and a data frame to be verified; An extraction module, used to extract labels from the source data frame and the data frame to be verified respectively according to the loaded template rule, to obtain label data corresponding to the source data and label data of the data to be verified; The template rule includes: a label word list of a source table, a label word list of a table to be verified, and a mapping relationship between the label word list of the source table and the label word list of the table to be verified; A comparison module, used to compare the label data corresponding to the source data with the label data of the data to be verified to obtain a data verification result; The extraction module is specifically used to extract the label data corresponding to the source data from the source data frame according to the label word list of the source table; and to extract the label data of the data to be verified from the data to be verified according to the label word list of the table to be verified; The comparison module is specifically used to compare the label data corresponding to the source data with the label data of the data to be verified according to the mapping relationship between the label vocabulary of the source table and the label vocabulary of the table to be verified, so as to obtain a data verification result.

8. A data verification system, characterized in that: include: A memory and a processor, wherein executable instructions of the processor are stored in the memory; wherein the processor is configured to execute the data verification method according to any one of claims 1 to 6 by executing the executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the data verification method described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for evaluating accuracy of data of different data sources

    CN106777235A

  • Data verification method and device, and electronic equipment

    CN107122368A