Data verification method and device, electronic equipment, storage medium and program product
By preprocessing and CRC verification of the data files before and after the calculation engine switch, the problem of inaccurate data consistency verification results is solved, and efficient and reliable data migration verification is achieved.
Patent Information
- Application Number
- CN202510409507.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the reliability of the data consistency verification results before and after the calculation engine switches are poor, resulting in low accuracy of the data migration verification results.
By obtaining the running results of the first and second data files, preprocessing is performed to unify the format of the preset type field, and the processed text is checked using the cyclic redundancy verification (CRC) algorithm to determine the consistency of the data files.
Improve the efficiency and reliability of data consistency verification, ensure the accuracy of data migration verification results, and reduce errors caused by operating system differences.
Smart Images

Figure CN120336077A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data verification method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] With the change of computing engines and the introduction of optimized parameters in the process of big data production, existing big data tasks need to be migrated. In related technologies, only when the data passes the consistency check can it be confirmed that the task can be migrated smoothly. Usually, the consistency of the data is checked by comparing the results of double-running the task (that is, the same task is simulated and executed once in the environment before the change and once in the environment after the change). However, there may be differences in the logic of data processing by different engines. For example, there are differences in the processing order of data, hash logic, or calculation accuracy of data. Therefore, in related technologies, the method of simulating double-running and comparing the operation results before and after switching the computing engine is likely to result in poor reliability of the obtained data consistency check results. Summary of the Invention
[0003] The present disclosure provides a data verification method, apparatus, electronic device, storage medium, and program product, which are used to solve the technical problem of low accuracy of data migration verification results obtained based on related technologies.
[0004] In a first aspect, the present application provides a data verification method, and the method includes:
[0005] Obtain a first data file and a second data file, where the first data file includes N rows of operation results generated by a first operating system running a target task, and the second data file includes M rows of operation results generated by a second operating system running the target task;
[0006] When N is equal to M, preprocess fields of a preset type in the first data file to obtain a first text, and preprocess the fields of the preset type in the second data file to obtain a second text; wherein, the preprocessing is used to unify the formats of the fields of the preset type;
[0007] Perform cyclic redundancy check (CRC) processing on the first text and the second text respectively to obtain a first CRC processing result corresponding to the first text and a second CRC processing result corresponding to the second text;
[0008] Determine a consistency check result of the first data file and the second data file according to the first CRC processing result and the second CRC processing result.
[0009] In a second aspect, the present application further provides a data verification apparatus, and the apparatus includes:
[0010] A data acquisition module for acquiring a first data file and a second data file, where the first data file includes N rows of operation results generated by a first operating system running a target task, and the second data file includes M rows of operation results generated by a second operating system running the target task;
[0011] A first processing module for preprocessing fields of a preset type in the first data file to obtain a first text and preprocessing the fields of the preset type in the second data file to obtain a second text when N is equal to M; wherein, the preprocessing is used to unify the formats of the fields of the preset type;
[0012] A second processing module for performing cyclic redundancy check (CRC) processing on the first text and the second text respectively to obtain a first CRC processing result corresponding to the first text and a second CRC processing result corresponding to the second text;
[0013] A result determination module for determining a consistency check result of the first data file and the second data file according to the first CRC processing result and the second CRC processing result.
[0014] In a third aspect, the present application provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, the steps of the method described in the first aspect are implemented.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0016] In a fifth aspect, the present application provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0017] In this application, the number of lines in the first data file and the second data file is compared. When the number of lines in the first data file is the same as that in the second data file, preprocessing is performed on the first data file and the second data file. Through preprocessing, the unified format of the preset type fields in the first data file and the second data file is ensured, the error between texts is reduced, and the accuracy of the verification process is improved. Subsequently, by performing CRC processing on the first text and the second text obtained through preprocessing respectively, and using the CRC processing results of both as the basis for consistency verification, a reliable basis is provided for the consistency verification results of the first data file and the second data file. In this way, through efficient preprocessing, accurate CRC algorithms, and reliable verification result feedback in the embodiments of this application, the efficiency and reliability of data consistency verification are effectively improved, the rapid identification and verification of data consistency are realized, and the accuracy of data migration verification results is improved. Description of the Drawings
[0018] Figure 1 is a schematic diagram of the double-run verification data process in the related art;
[0019] Figure 2 is one of the schematic diagrams of the data verification method provided by the embodiments of this application;
[0020] Figure 3 is another schematic diagram of the data verification method provided by the embodiments of this application;
[0021] Figure 4 is a schematic diagram of the structure of the data verification device provided by the embodiments of this application;
[0022] Figure 5 is a schematic diagram of the structure of an electronic device provided by the embodiments of this application. Detailed Embodiments
[0023] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0024] Before further elaborating on the embodiments of this application, the related technologies and terms involved in the embodiments of this application are described.
[0025] The double-run phase, that is, the phase where the same tasks are run in two production environments. The same tasks are run separately in the two environments and two running results are generated. In the data processing or verification process, the same data set is run independently twice to ensure the consistency and reliability of the results. This process can be widely applied to data transmission or task execution scenarios that require high accuracy. For example, Figure 1 As shown, taking the upgrade of the Spark computing engine from version 3.1 to version 3.5 as an example, the computing engines of version 3.1 and version 3.5 each simulate and execute the same Structured Query Language (SQL) task once. First, to eliminate the incompatibility in syntax or functionality between different versions of the Spark computing engine and ensure that the task can be correctly executed in both environments, the SQL task needs to be rewritten. Subsequently, the rewritten SQL tasks are executed separately to obtain two separate running results, namely temporary table A and temporary table B. To check whether the data is consistent, temporary table A and temporary table B are compared item by item to check the data consistency and determine whether there are differences. Only by passing the consistency check can it be confirmed that the task can be migrated smoothly.
[0026] Cyclic Redundancy Check (CRC) is a data transmission error detection technology, specifically used to verify whether errors occur during data transmission or storage. It can effectively detect burst errors and information loss.
[0027] Next, in combination with the accompanying drawings, a data verification method, device, electronic device, storage medium, and program product provided by an embodiment of the present application will be described in detail through specific embodiments and their application scenarios.
[0028] For example, Figure 2 As shown, a data verification method in the present application specifically includes the following steps:
[0029] Step 101: Obtain a first data file and a second data file. The first data file includes N rows of running results generated by a first operating system running a target task, and the second data file includes M rows of running results generated by a second operating system running the target task.
[0030] It should be noted that the first data file and the second data file can be two running results obtained by two different operating systems executing the same task in the double-run phase.
[0031] For example, in the scenario of updating and upgrading the computing engine version, the first operating system can be the computing engine before the update (the historical version, which is relatively stable and can participate in the data verification process as the operating system of the control group), and the second operating system is the computing engine after the update (the updated version, whose stability is to be verified, and participates in the data verification process as the operating system of the experimental group). Therefore, the first data file can be used as the operation result generated by the control group during the double-run stage, and the second data file can be used as the operation result generated by the experimental group during the double-run stage.
[0032] Among them, the first data file includes N lines of operation results, and the second data file includes M lines of operation results. Both N and M are integers greater than or equal to 1, and N and M may be equal or unequal.
[0033] Exemplarily, the first data file and the second data file can also be two temporary tables that respectively store operation results, and the two temporary tables include row data and column data. The temporary table corresponding to the first data file contains N rows of row data, and the temporary table corresponding to the second data file contains M rows of row data. Thus, it is convenient to perform a lightweight line number verification on the first data file and the second data file to check the consistency of the data files.
[0034] Step 102: When N is equal to M, preprocess the fields of the preset type in the first data file to obtain a first text, and preprocess the fields of the preset type in the second data file to obtain a second text; wherein, the preprocessing is used to unify the formats of the fields of the preset type.
[0035] In this application, when the number of lines in the first data file and the second data file is the same, it indicates that the number of data entries in the first data file and the second data file is equal, and there is no significant difference in the quantity level between the two data files. Therefore, more in-depth verification can be performed (for example, performing CRC verification processing after standardizing the data files). In addition, if the number of lines is inconsistent, it means that data has been lost or incorrect during the extraction or processing process. Therefore, it can be directly determined that further investigation and processing are required.
[0036] Among them, since different operating systems may have different logics for processing target tasks. For example, there are differences in the data processing order, the adopted hash logic, and the calculation accuracy of data among different operating systems. Therefore, it is necessary to preprocess the fields of the preset type in the first data file or the second data file to unify the format types of the fields, reduce data errors caused by different operating system task processing logics, etc., and improve the accuracy of data verification.
[0037] It should be noted that the above-mentioned fields of the preset type can be fields in the data file that are susceptible to factors such as processing order, processing logic, and calculation accuracy. For example, arrays, maps, integers, floating-point numbers (Float / Double), etc. are not listed one by one here, and the fields of the preset type can be set according to actual application requirements.
[0038] Exemplarily, the fields of the preset type can include fields of the type storing a set of key-value pairs or the array type. If different operating systems have different processing orders or hash logics for data, although the data within the set of key-value pairs or the array is not affected, the arrangement order of the elements within the set of key-value pairs or the array is different. Thus, the fields of the type storing a set of key-value pairs or the array type can be preprocessed, that is, the arrangement order of the fields of the type storing a set of key-value pairs or the array type can be unified.
[0039] Step 103: Perform cyclic redundancy check (CRC) processing on the first text and the second text respectively to obtain a first CRC processing result corresponding to the first text and a second CRC processing result corresponding to the second text.
[0040] In this application, the first text and the second text can be used as the texts for CRC processing. The first text and the second text have achieved unified field formats after preprocessing, reducing the possibility of data errors easily caused by inconsistent field formats. For the first text or the second text, specific mathematical operations are performed on the data blocks therein using CRC to generate a check code of a fixed length. Subsequently, the CRC algorithm reads each character in the text and calculates the corresponding CRC check value according to a specific polynomial function.
[0041] The above-mentioned first CRC processing result corresponding to the first text represents the CRC check value of the first text, and the second CRC processing result corresponding to the second text represents the CRC check value of the second text. In this way, since the obtained CRC check values of the first text and the second text are small and convenient for storage, fast comparison can be realized, improving the data verification efficiency.
[0042] In addition, even if there are only minor changes between the first text and the second text, their corresponding CRC check values will be very different. Therefore, the consistency of the text can be effectively checked, improving the accuracy of data verification.
[0043] Step 104: Determine the consistency verification result of the first data file and the second data file according to the first CRC processing result and the second CRC processing result.
[0044] Among them, the first CRC processing result and the second CRC processing result are compared. When the first CRC processing result is the same as the second CRC processing result, it can be determined that the consistency check result passes the consistency check. Thus, it is possible to provide a basis for reliability for the subsequent data migration process, reduce the impact on user services, and ensure the continuity of business operations.
[0045] When the first CRC processing result is different from the second CRC processing result, it can be determined that the consistency check result fails the consistency check. The data source can be checked to confirm the specific reasons for the inconsistency between the first data file and the second data file, such as data loss, format error, field missing, etc. Data review can also be performed to review the first data file and the second data file line by line, compare the corresponding fields, and confirm the specific situation of the differences. Inconsistent data can also be repaired according to the first CRC processing result and the second CRC processing result, re-extract the different data, update the error records, or delete redundant data.
[0046] In the embodiment of the present application, first, the number of lines of the first data file and the number of lines of the second data file are compared. When the number of lines of the first data file is the same as the number of lines of the second data file, it indicates that there is no significant difference between the two data files in terms of quantity at this time, and the first data file and the second data file can be further preprocessed. Then, by preprocessing the first data file and the second data file respectively, the unified format of the preset type fields in the first data file and the second data file can be realized, and the error between texts can be reduced.
[0047] Subsequently, by performing CRC processing on the first text and the second text obtained by preprocessing respectively, the CRC processing results of the two can be used as the basis for consistency check, providing a reliable basis for the consistency check result of the first data file and the second data file. In this way, the embodiment of the present application effectively improves the efficiency and reliability of data consistency check, realizes the rapid identification and check of data consistency, and improves the accuracy of data migration check results.
[0048] In one embodiment, the fields of the preset type include a first type field, a second type field, and a third type field. The first type field is a field for storing a key-value pair set type or an array type, the second type field is a field for storing a structure, and the third type field is a field for storing a floating point number;
[0049] The preprocessing includes:
[0050] Perform a first operation on the operation result generated by the operating system running the target task to obtain a target text; wherein, the first operation includes at least one of the following: sorting the first type of fields in lexicographical order; converting the second type of fields into strings; according to the value range where the integer part of the third type of fields is located, retaining the K-digit decimal part corresponding to the value range in the third type of fields;
[0051] Convert the target text into a string in a preset format to obtain the first text or the second text.
[0052] In this embodiment, by performing a first operation on the fields of a preset type in the first data file and the second data file, and converting them into a unified field format through the first operation, the error caused by inconsistent fields is reduced, the accuracy of data consistency verification is effectively improved, and by converting complex data structures into strings, the complexity of data management can be reduced and the efficiency of data analysis and processing can be improved.
[0053] In some embodiments, the operation results generated when running the target task may include fields of types such as key-value pair sets / arrays, structures, and floating-point numbers. For example, a user information record may simultaneously include a Map (such as grades), a structure (Struct) (such as user attributes), and floating-point values, such as calculation precision, etc. Fields with different storage formats may result in inconsistent results obtained by different computing engines. The preset type of fields may include fields for storing key-value pair sets or array types, such as field types like Map and Array. It also includes fields for storing structures, such as the Struct field type, which can store structures. It may also include fields for storing floating-point numbers, such as floating-point numbers or high-precision numbers like Float, Double, and Decimal.
[0054] Among them, the preprocessing is the first operation such as sorting, converting to a string, and precision control performed on the above-mentioned preset type of fields, and the purpose is to eliminate the non-essential differences caused by storage differences. For example: after arranging the Map keys in lexicographical order, even if the initial orders of the two systems are different, the content of the generated target text will be the same; converting the structure into a fixed-format string (such as JSON) can unify the representation methods of different systems. In this way, after the preprocessed data is integrated into the target text, it is further converted into a string in a preset format, and finally the first text or the second text is generated.
[0055] Among them, a first operation is performed on fields of a preset type, and different operation processes can be carried out on fields of different types. Exemplarily, for fields of a key-value pair collection type such as Map or for fields of an array type such as Array, the first operation performed on them can be: for the Map type, the key-value pairs in the collection or the elements in the array can be sorted in lexicographical order according to the keywords (keys); for the Array, after converting the elements therein into text strings, lexicographical sorting can be performed.
[0056] The above lexicographical sorting may refer to a method of comparing elements one by one according to the American Standard Code for Information Interchange (ASCII) encoding or Unicode encoding of characters and sorting them in the way words are arranged in a dictionary.
[0057] Exemplarily, for the Struct field type, it can be directly converted into a string, for example, into a Json-formatted string. In this way, by converting the Struct field type into a standard JSON format, the consistency of data in various environments and scenarios can be ensured, and errors caused by inconsistent formats can be reduced.
[0058] For example, for fields of floating-point numbers or high-precision numbers such as Float, Double, and Decimal, since different operating systems may be written in different programming languages, and there are differences in the precision definitions of floating-point numbers in different programming language environments, it is inevitable that there are differences in precision between different operating systems. Therefore, in this application, corresponding standards can be set for the floating-point numbers in the first data file and the second data file to ensure that the integer parts of the floating-point numbers are consistent, and there may be slight differences in their decimal parts, such as a difference of one ten-thousandth. Specifically, the following operations can be performed on floating-point numbers such as Float, Double, and Decimal: according to the value range where the integer part of the floating-point number is located, retain the decimal part with the corresponding number of digits in the value range of the floating-point number.
[0059] For example, if the integer part of the floating-point number is greater than 10000, only the integer part of the floating-point number can be retained, and the decimal part can be removed. If the integer part of the floating-point number is greater than 1000, one significant decimal digit of the floating-point number can be retained. If the integer part of the floating-point number is greater than 100, two significant decimal digits of the floating-point number can be retained. Correspondingly, if the integer part of the floating-point number is greater than 10, three significant decimal digits of the floating-point number can be retained. If the integer part of the floating-point number is less than 1, four significant decimal digits of the floating-point number can be retained.
[0060] In an application, if both the first data file and the second data file contain the following fields:
[0061] The first type of field (a field of key-value pair set type or a field of numerical type): For example, it represents a list of movies a user likes, such as ["Inception", "Interstellar", "Dunkirk"];
[0062] The second type of field (a field of struct type): For example, user details, such as: {"name": "Alice", "age": 30, "city": "New York"};
[0063] The third type of field (a field of floating-point type): For example, user ratings, such as 4.5, 3.75, 5.0.
[0064] Thus, the following first operation steps can be performed on the above fields respectively:
[0065] Sort the field of key-value pair set type or the field of numerical type. The movie names in the first type of field can be sorted in lexicographical order. For example, sort ["Inception", "Dunkirk", "Interstellar"], and the sorted field is ["Dunkirk", "Inception", "Interstellar"];
[0066] Convert the struct to a string. The second type of field (user details) can be converted to a string format. For example, directly convert "name:Alice,age:30,city:New York" to {"name": "Alice", "age": 30, "city": "New York"};
[0067] Retain decimal places for floating-point numbers. If the range of the integer part of the pre-set floating-point number is from 0 to 10, the K decimal places of the original values 4.5, 3.75, 5.0 can all be retained. If it is specified to retain two decimal places, the result is still 4.50, 3.75, 5.00.
[0068] In this way, after completing the above first operation processing on different types of fields respectively, the obtained results can form an integrated text, that is, the target text. For example, the target text can be expressed as:
[0069] {
[0070] "favorite_movies": ["Dunkirk", "Inception", "Interstellar"],
[0071] "user_details": {"name": "Alice", "age": 30, "city": "New York"},
[0072] "ratings": [4.50, 3.75, 5.00]
[0073] }
[0074] It should be understood that the above target text includes the fields obtained after performing the first operation on the fields of the preset type, and also includes other types of fields that have not been processed by the first operation in the operation result. Other types of fields do not need to perform the first operation, but will also be text-converted together with the fields obtained after performing the first operation to achieve the splicing process of the entire line of fields in the target text.
[0075] Furthermore, the target text can be converted into a string of a preset format, for example, uniformly converted into a text string, and the strings are spliced according to the arrangement order of each string in each line to obtain the spliced long text, that is, the first text or the second text. In this way, by splicing multiple strings into a long text, it can be ensured that the data conforms to the unified format specification during storage and transmission, reducing the number of parsing and reading operations on short texts, and helping to quickly perform consistency check and other processes, such as reliability judgment by calculating the overall hash value or CRC.
[0076] In one embodiment, step 103 includes:
[0077] Step 1031: Perform CRC processing on the first text to obtain the CRC check value of each line of data in the first text, and based on the CRC check value of each line of data in the first text, obtain the first CRC processing result as the sum of the CRC check values of each line of data in the first text;
[0078] Step 1032: Perform CRC processing on the second text to obtain the CRC check value of each line of data in the second text, and based on the CRC check value of each line of data in the second text, obtain the second CRC processing result as the sum of the CRC check values of each line of data in the second text.
[0079] In this embodiment, the CRC algorithm is executed on each line of data in the first text to obtain the CRC check value for each line in the first text. The CRC algorithm in the same manner is executed on each line of data in the second text to obtain the CRC check value for each line in the second text. Moreover, the CRC check values for each line in the first text or the second text are aggregated to obtain the first CRC processing result or the second CRC processing result, thereby determining the detection ability of the overall data consistency. Any change in a single line of data can be reflected in the first CRC processing result or the second CRC processing result. In this way, the embodiments of the present application can effectively improve the accuracy and efficiency of data processing, transmission, and verification, laying a good foundation for subsequent data operations.
[0080] Exemplarily, the first text may be a temporary table storing data in a unified field format. The data in the first text can be arranged in rows and columns, and the data for each row is a long text concatenated string. When performing CRC processing, the row data in the first text can be directly verified to obtain the check value for each row. Thus, the first CRC processing result for the entire first text can be obtained as the sum of the CRC check values for each row of data in the first text. Similarly, the second text can adopt a processing method similar to that for determining the first CRC processing result corresponding to the first text, which is not elaborated in the embodiments of the present application.
[0081] In one embodiment, after step 101, the method further includes:
[0082] Step 105: When N is not equal to M, or when the CRC processing result of the first text and the CRC processing result of the second text are different, obtain a third data file generated by the first operating system running the target task. The third data file includes L lines of operation results;
[0083] Step 106: When L is not equal to N, confirm that the consistency check result of the first data file and the second data file passes the data consistency check.
[0084] In this embodiment, when it is determined that the number of lines in the first data text is inconsistent with the number of lines in the second text, that is, there is a difference in the quantity level between the two data files. Since it cannot be excluded whether the data inconsistency is caused by the randomness of the target task itself (for example, the target task contains a Random function), the environment of the control group can be re - simulated, the control group can be re - run, and the target task can be run again using the first operating system to obtain the third data file. Or, when the CRC processing result of the first text and the CRC processing result of the second text are different, the control group is re - run to obtain the third data file.
[0085] Among them, the CRC processing results of the first text and the second text are different, that is, at least one line of data in the two texts has changed. During the data generation or processing process, some data in the data file is modified, deleted, or added, resulting in different generated CRC check values. Or, although the number of lines in the two texts is the same, there are differences in the data format or content within the lines. Or, the data file may be damaged or not fully written during the movement or storage process, resulting in the loss or tampering of some line data, which leads to the mismatch of the CRC check values.
[0086] In this application, rerunning the first operating system to obtain the third data file is because the control group usually represents the historical version of stability and consistency. By rerunning the control group, it can be verified whether the obtained operation results are consistent without external changes (such as code updates, configuration adjustments, etc.). By rerunning the control group, not only can a new round of operation results be collected, but also it can ensure that the data file obtained in a consistent environment configuration is more credible, providing effective support for subsequent comparisons. Rerunning the control group can also help identify potential problems, such as whether the system performance is abnormal due to specific input or parameter combinations, and this abnormal performance may not be detected in the original results.
[0087] Exemplarily, N is the number of lines of the operation result generated by the first operating system when executing the target task for the first time, and the third data file is the operation result generated by repeating the execution in exactly the same environment (i.e., the first operating system executes the target task). The L lines in the third data file are the number of lines of the operation result generated by the first operating system when executing the target task again. Since the conditions (including input data, system status, etc.) for the first operating system to execute the task twice are exactly the same, it is expected that the same results should be generated twice. If the number of lines in the first data file is not equal to the number of lines in the third data file, combined with the nature of the task (for example, there is a random function in the target task), it can be inferred that the target task itself has randomness. For example, even if the input conditions are the same, the output may still be different due to random factors inside some algorithms or functions. Therefore, the inconsistency between L and N can point to the instability of the first operating system during the execution of the target task.
[0088] In this case, the inconsistent number of lines of the results obtained from the two runs means that during the construction of the result set, it is indeed affected by random elements, confirming the random nature of the task itself and the instability of the results. That is, the data set may be affected by the inherent characteristics of the task rather than external factors. At this time, no judgment is made on the data consistency check result, reducing the complexity of data verification.
[0089] In one embodiment, after the step 106, the method further includes:
[0090] Step 107, when L is equal to N, preprocess a preset field of a preset type in the third data file to obtain a third text;
[0091] Step 108, perform CRC processing on the third text to obtain a third CRC processing result corresponding to the third text;
[0092] Step 109, when the third CRC processing result is the same as the first CRC processing result, confirm that the consistency check result of the first data file and the second data file fails the data consistency check.
[0093] In this embodiment, the third data file is obtained by repeatedly executing the target task in the same environment as the first data file, and when L is equal to M, it means that the number of rows of the third data file and the first data file is the same, thus providing a suitable basis for comparing the CRC processing results. If the third CRC processing result of the third text is the same as the first CRC processing result of the first text, it indicates that the data contents of the third text and the first text are actually the same, reflecting that the result sets are not affected by random factors.
[0094] Based on the above, when the third CRC processing result is the same as the first CRC processing result, it indicates that the target task is not random, but the first CRC processing result and the second CRC processing result show inconsistency, indicating that the consistency check result fails the data consistency check at this time. Furthermore, when the third CRC processing result is different from the first CRC processing result, it indicates that the target task is random. Although the first CRC processing result and the second CRC processing result show inconsistency, due to the influence of the randomness of the target task, it can be considered that the consistency check result passes the data consistency check at this time, improving the accuracy of data verification.
[0095] In one embodiment, after the step 103, the method further includes:
[0096] Step 110, when the first CRC processing result and the second CRC processing result are different, perform a second operation on a preset type of field in the first data file to obtain a third text, and perform the second operation on the preset type of field in the second data file to obtain a fourth text;
[0097] Step 111, perform CRC processing on the third text and the fourth text respectively to obtain a third CRC processing result and a fourth CRC processing result; wherein, the third CRC processing result includes the CRC check values of each column field in the third text, and the fourth CRC processing result includes the CRC check values of each column field in the fourth text;
[0098] Step 112: Obtain the target column, where the target column is the column in the fourth text whose CRC check value is different from the corresponding column in the third text.
[0099] Wherein, the second operation includes at least one of the following: sorting the fields for storing key-value set types or array types in lexicographical order; converting the fields for storing structures into strings; according to the value range where the integer part of the floating-point field is located, retaining the K-digit decimal part corresponding to the value range in the floating-point field.
[0100] It is worth mentioning that a column is the natural unit of field attributes. In a data file, each column represents a specific field or attribute. If the data in a certain column is inconsistent in two files, it indicates that there are systematic problems in this field (such as incorrect data source, calculation logic differences, etc.), rather than accidental errors in a single record. If only a certain column needs to be repaired, the data of this column can be processed specifically. If checked row by row, the entire record may be misjudged as invalid or all content needs to be retransmitted, wasting resources. Therefore, the embodiments of the present application take column data as the granularity.
[0101] In this embodiment, since the first CRC processing result is different from the second CRC processing result, the inconsistency of the data in the first data file and the second data file can be determined. Therefore, by performing CRC processing on the column data in the first data file and the column data in the second data file respectively, the columns in the second data file that are different from the first data file can be quickly identified. In this way, the specific fields or columns in the second data file that have changes during operation can be clearly understood, the data quality management ability can be improved, the repeated operation of all data is avoided, and the processing efficiency can be significantly improved by processing at the granularity of column data, and the data management and processing ability can be improved.
[0102] In some embodiments, the second operation is performed on the fields of the preset type in each column of the first data file to obtain the third text, and the second operation is performed on the fields of the preset type in each column of the second data file to obtain the fourth text. The specific processing process of the fields of the preset type and the second operation can refer to the specific description in the foregoing embodiments of the present application. The second operation can be implemented in a similar processing manner as the first operation, and will not be repeated here to avoid redundancy.
[0103] In this application, at the granularity of column data, CRC processing is performed on each column field in the first data file after the second operation to obtain the CRC check values of each column field in the first data file. Therefore, the obtained third CRC processing result includes the CRC check values of each column field in the first data file. Similarly, the fourth CRC processing result can be obtained in a manner similar to the third CRC processing result, and details will not be repeated here to avoid redundancy.
[0104] Among them, the target column can be determined by comparing the CRC check values of each column in the third text and the CRC check values of each column in the fourth text. In the case where no random changes occur in the third text and the fourth text, each column field in the third text will correspond to a column field in the fourth text, that is, the CRC check value of the a-th column in the third text should be the same as the CRC check value of the a-th column in the fourth text. For the target column, the CRC check value of the b-th column in the third text is different from the CRC check value of the b-th column in the fourth text. At this time, the b-th column in the third text is the target column, and both a and b are integers greater than or equal to 1.
[0105] Exemplarily, to ensure that the first data text and the second data text are as consistent as possible in format, type, and content, the second operation is performed on the relevant column fields in the first data text and the second data text to obtain the third text and the fourth text. For each column in the third text and each column in the fourth text, calculate their CRC check values respectively. Calculate the CRC check value of each column in the third text and record it in a list or array, such as crc_values_C. Calculate the CRC check value of each column in the fourth text and record it in another list or array, such as crc_values_D. Subsequently, traverse each column i and compare crc_values_C[i] with crc_values_D[i]. If crc_values_C[i] is different from crc_values_D[i], mark this column as the target column. A list can also be used to store the indices or names of all target columns, such as target_columns. Finally, output the detailed information of the target column.
[0106] It is worth mentioning that in this application, the target column is determined by comparing the column data in the first data file and the second data file. This is because checking by column can significantly reduce the complexity of comparison. Since the number of rows in the data sets of the first data file and the second data file is large, comparing row by row will involve a large number of repeated operations. By processing columns, the integrity and consistency of each column can be checked centrally. If processed by row, the entire record may be misjudged as invalid or the entire content may need to be retransmitted, resulting in waste of resources. Therefore, processing by column can greatly improve efficiency. In addition, data tables are usually constructed based on columns, and each column represents a specific data attribute (such as user ID, behavior type, score, etc.). Therefore, comparing columns can directly identify which specific attributes are inconsistent in different data sources, making the processing of data fields more logical.
[0107] Thus, by obtaining the target column in the embodiment of this application, the source of differences between the first data file and the second data file is located. Compared with row-level comparison, column-level processing can find problems in specific fields faster and discover the error sources in the data file in a timely manner, thereby optimizing the data processing flow and improving the overall data quality. It significantly improves the efficiency and accuracy of data management and processing.
[0108] For ease of understanding, as Figure 3 shown in a data verification method in the embodiment of this application, it is illustrated as follows:
[0109] Step 201: Obtain a first data file, where the first data file includes N rows of operation results generated by a first operating system running a target task;
[0110] Step 202: Obtain a second data file, where the second data file includes M rows of operation results generated by a second operating system running the target task;
[0111] Step 203: Determine whether N is equal to M; in the case where N is not equal to M, execute Step 208; in the case where N is equal to M, execute Step 204;
[0112] Step 204: In the case where N is equal to M, preprocess the fields of a preset type in the first data file to obtain a first text, and preprocess the fields of the preset type in the second data file to obtain a second text; wherein, the preprocessing is used to unify the formats of the fields of the preset type;
[0113] Step 205: Perform cyclic redundancy check (CRC) processing on the first text and the second text respectively to obtain a first CRC processing result corresponding to the first text and a second CRC processing result corresponding to the second text;
[0114] Step 206: Determine whether the first CRC processing result is the same as the second CRC processing result;
[0115] Step 207: When the first CRC processing result is the same as the second CRC processing result, confirm that the consistency check result of the first data file and the second data file passes the data consistency check, and end the execution;
[0116] Step 208: When the first CRC processing result is different from the second CRC processing result, obtain a third data file, where the third data file includes L rows of operation results generated by the first operating system running the target task again;
[0117] Step 209: Determine whether N is equal to L; when N is not equal to L, execute Step 207;
[0118] Step 210: When N is equal to L, preprocess a preset field of a preset type in the third data file to obtain a third text; perform CRC processing on the third text to obtain a third CRC processing result corresponding to the third text;
[0119] Step 211: Determine whether the first CRC processing result is the same as the third CRC processing result; when the second CRC processing result is different from the third CRC processing result, execute Step 207;
[0120] Step 212: When the first CRC processing result is the same as the third CRC processing result, confirm that the consistency check result of the first data file and the second data file fails the data consistency check.
[0121] It should be noted that by comparing whether the first CRC processing result is the same as the third CRC processing result, if they are different, it means that the first operating system obtains different numbers of rows of results when executing the same target task twice, and the task itself has randomness and the result set is unstable. If they are the same, it means that the task does not have randomness and the result set is stable. Therefore, it can be determined that the data of the double-run task is inconsistent, and it can be confirmed that the consistency check result fails the data consistency check.
[0122] Based on the above, a data verification method in an embodiment of the present application is used to solve the technical problem of low accuracy of the data migration verification result obtained based on the related technology. Among them, the embodiments of the present application can refer to the specific implementations of the foregoing Figure 2 shown in the method of each embodiment, and the repeated parts will not be described again.
[0123] See Figure 4 , Figure 4 is a data verification device provided by an embodiment of the present disclosure, as Figure 5As shown, the data verification device 300 includes:
[0124] A data acquisition module 301, configured to acquire a first data file and a second data file, where the first data file includes N lines of operation results generated by a first operating system when running a target task, and the second data file includes M lines of operation results generated by a second operating system when running the target task;
[0125] A first processing module 302, configured to, when N is equal to M, preprocess fields of a preset type in the first data file to obtain a first text, and preprocess the fields of the preset type in the second data file to obtain a second text; wherein, the preprocessing is used to unify the formats of the fields of the preset type;
[0126] A second processing module 303, configured to perform cyclic redundancy check (CRC) processing on the first text and the second text respectively to obtain a first CRC processing result corresponding to the first text and a second CRC processing result corresponding to the second text;
[0127] A result determination module 304, configured to determine a consistency verification result of the first data file and the second data file according to the first CRC processing result and the second CRC processing result.
[0128] In one embodiment, the fields of the preset type include a first type field, a second type field, and a third type field. The first type field is a field used to store a key-value pair set type or an array type, the second type field is a field used to store a structure, and the third type field is a field used to store a floating point number;
[0129] The preprocessing is used for:
[0130] Performing a first operation on the operation results generated by the operating system when running the target task to obtain a target text; wherein, the first operation includes at least one of the following: sorting the first type fields in lexicographical order; converting the second type fields into strings; according to the value range where the integer part of the third type field is located, retaining K decimal places corresponding to the value range in the third type field;
[0131] Converting the target text into a string of a preset format to obtain the first text or the second text.
[0132] In one embodiment, the first processing module 302 includes:
[0133] A first processing unit, configured to perform CRC processing on the first text to obtain a CRC check value for each line of data in the first text;
[0134] A second processing unit, configured to perform CRC processing on the second text to obtain a CRC check value for each line of data in the second text;
[0135] A first obtaining unit, configured to obtain, based on the CRC check value of each line of data in the first text, the first CRC processing result as the sum of the CRC check values of each line of data in the first text;
[0136] A second obtaining unit, configured to obtain, based on the CRC check value of each line of data in the second text, the second CRC processing result as the sum of the CRC check values of each line of data in the second text.
[0137] In one embodiment, the data verification device 300 further includes:
[0138] A first obtaining module, configured to obtain a third data file generated by the first operating system when running the target task, where the third data file includes L lines of operation results, when N is not equal to M, or when the CRC processing results of the first text and the second text are different;
[0139] A first confirmation module, configured to confirm that the consistency verification result of the first data file and the second data file passes the data consistency verification when L is not equal to N.
[0140] In one embodiment, the data verification device 300 further includes:
[0141] A third processing module, configured to preprocess a preset field of a preset type in the third data file to obtain a third text when L is equal to N;
[0142] A fourth processing module, configured to perform CRC processing on the third text to obtain a third CRC processing result corresponding to the third text;
[0143] A second confirmation module, configured to confirm that the consistency verification result of the first data file and the second data file fails to pass the data consistency verification when the third CRC processing result is the same as the first CRC processing result.
[0144] In one embodiment, the data verification device 300 further includes:
[0145] A first execution module, configured to perform a second operation on a preset type of field in the first data file to obtain a third text, and perform the second operation on the preset type of field in the second data file to obtain a fourth text, when the CRC processing results of the first text and the second text are different;
[0146] A fifth processing module is configured to perform CRC processing on the third text and the fourth text respectively to obtain a third CRC processing result and a fourth CRC processing result. Wherein, the third CRC processing result includes CRC check values of each column field in the third text, and the fourth CRC processing result includes CRC check values of each column field in the fourth text.
[0147] A second obtaining module is configured to obtain a target column, where the target column is a column in the fourth text whose CRC check value is different from the CRC check value of the corresponding column in the third text.
[0148] Wherein, the second operation includes at least one of the following: sorting the fields for storing a set in lexicographical order; converting the fields for storing a structure into a string; and retaining the K-bit decimal part corresponding to the value range in the floating-point field according to the value range where the integer part of the floating-point field is located.
[0149] The data verification device 300 provided by the embodiments of the present disclosure can implement each process in the data verification method embodiments as shown above Figure 2 or Figure 3 as shown. To avoid repetition, details are not described herein again.
[0150] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0151] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 400 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0152] As Figure 5As shown, device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0153] Multiple components in device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0154] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include but are not limited to a central processing unit (CPU), a graphic process unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the data verification method. For example, in some embodiments, the data verification method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the data verification method described above can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the data verification method in any other appropriate way (e.g., by means of firmware).
[0155] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0156] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on the remote machine or server.
[0157] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0158] As used herein, the term "machine-readable medium" refers to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) that provides machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that provides machine instructions and / or data to a programmable processor.
[0159] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0160] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0161] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server that incorporates a blockchain.
[0162] An embodiment of this application also provides a computer program product, including computer instructions, which when executed by a processor, implement each process of the method embodiment shown above Figure 1 and can achieve the same technical effects. To avoid repetition, details are not described herein again.
[0163] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution of this disclosure can be achieved, and no limitations are imposed herein.
[0164] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A data verification method, characterized in that, Including: Obtain a first data file and a second data file, where the first data file includes N lines of operation results generated by a first operating system running a target task, and the second data file includes M lines of operation results generated by a second operating system running the target task; When N is equal to M, preprocess the fields of a preset type in the first data file to obtain a first text, and preprocess the fields of the preset type in the second data file to obtain a second text; wherein, the preprocessing is used to unify the formats of the fields of the preset type; Perform cyclic redundancy check (CRC) processing on the first text and the second text respectively to obtain a first CRC processing result corresponding to the first text and a second CRC processing result corresponding to the second text; Determine the consistency check result of the first data file and the second data file according to the first CRC processing result and the second CRC processing result.
2. The method according to claim 1, wherein The fields of the preset type include a first type field, a second type field, and a third type field. The first type field is a field used to store a key-value pair set type or an array type, the second type field is a field used to store a structure, and the third type field is a field used to store a floating point number; The preprocessing includes: Perform a first operation on the operation results generated by the operating system running the target task to obtain a target text; wherein, the first operation includes at least one of the following: Sort the first type fields in lexicographical order; Convert the second type fields into strings; According to the value range where the integer part of the third type field is located, retain K decimal places corresponding to the value range in the third type field; Convert the target text into a string in a preset format to obtain the first text or the second text.
3. The method according to claim 1, wherein The performing cyclic redundancy check (CRC) processing on the first text and the second text respectively to obtain a first CRC processing result corresponding to the first text and a second CRC processing result corresponding to the second text includes: Perform CRC processing on the first text to obtain the CRC check value of each line of data in the first text, and based on the CRC check values of each line of data in the first text, obtain the first CRC processing result as the sum of the CRC check values of each line of data in the first text; Perform CRC processing on the second text to obtain the CRC check value of each line of data in the second text, and based on the CRC check values of each line of data in the second text, obtain the second CRC processing result as the sum of the CRC check values of each line of data in the second text.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: When N is not equal to M, or when the CRC processing results of the first text and the second text are different, obtain a third data file generated by the first operating system running the target task, where the third data file includes L lines of operation results; When L is not equal to N, confirm that the consistency check result of the first data file and the second data file passes the data consistency check.
5. The method according to claim 4, wherein The method further includes: When L is equal to N, preprocess a preset field of a preset type in the third data file to obtain a third text; Perform CRC processing on the third text to obtain a third CRC processing result corresponding to the third text; When the third CRC processing result is the same as the first CRC processing result, confirm that the consistency check result of the first data file and the second data file fails the data consistency check.
6. The method according to claim 1, wherein The method further includes: When the first CRC processing result is different from the second CRC processing result, perform a second operation on a preset type of field in the first data file to obtain a third text, and perform the second operation on the preset type of field in the second data file to obtain a fourth text; Perform CRC processing on the third text and the fourth text respectively to obtain a third CRC processing result and a fourth CRC processing result; wherein, the third CRC processing result includes CRC check values of each column field in the third text, and the fourth CRC processing result includes CRC check values of each column field in the fourth text; Obtain a target column, where the target column is a column in which the CRC check value in the fourth text is different from the CRC check value of the corresponding column in the third text; Wherein, the second operation includes at least one of the following: Sort the fields for storing key-value pair sets or array types in lexicographical order; Convert the fields for storing structures into strings; According to the value range where the integer part of the floating-point field is located, retain the K-bit decimal part corresponding to the value range in the floating-point field.
7. A data verification device, characterized in that, Includes: A data acquisition module, configured to acquire a first data file and a second data file, where the first data file includes N rows of operation results generated by a first operating system running a target task, and the second data file includes M rows of operation results generated by a second operating system running the target task; A first processing module, configured to, when N is equal to M, preprocess a preset type of field in the first data file to obtain a first text, and preprocess the preset type of field in the second data file to obtain a second text; wherein, the preprocessing is used to unify the formats of the preset type of fields; A second processing module, configured to perform cyclic redundancy check (CRC) processing on the first text and the second text respectively to obtain a first CRC processing result corresponding to the first text and a second CRC processing result corresponding to the second text; A result determination module, configured to determine the consistency check result of the first data file and the second data file according to the first CRC processing result and the second CRC processing result.
8. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the data verification method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the steps of the data verification method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, It includes computer instructions. When the computer instructions are executed by the processor, it implements the steps of the data verification method according to any one of claims 1 to 6.