Data verification method and device, equipment and storage medium

By performing the execution results of the override processing tasks in the source processing engine and the target processing engine, the data processing tasks are automatically checked, and the problem of inaccuracy and low efficiency is solved before migration is solved, ensuring the accuracy and stability of data migration.

CN120492434APending Publication Date: 2025-08-15HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510356065.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the verification accuracy and efficiency of data processing tasks before migration are low, and are often performed manually, resulting in time-consuming and error-prone.

Method used

By executing the execution results of the rewrite processing tasks in the source processing engine and the target processing engine, the original processing tasks are automatically checked and the tasks to be migrated are determined.

Benefits of technology

It realizes automated verification before data processing tasks migration, improves accuracy and efficiency, avoids the impact on the production environment, and ensures data stability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492434A_ABST
    Figure CN120492434A_ABST
Patent Text Reader

Abstract

The invention provides a data verification method and device, equipment and a storage medium. The data verification method comprises the steps that multiple original processing tasks executed in a source processing engine are determined; rewriting the plurality of original processing tasks to obtain a plurality of rewritten processing tasks; the multiple rewriting processing tasks are sent to a source processing engine and a target processing engine, so that the multiple rewriting processing tasks are executed through the source processing engine and the target processing engine; receiving a plurality of source execution results returned by the source processing engine and a plurality of target execution results returned by the target processing engine; and checking the plurality of original processing tasks according to the plurality of source execution results and the plurality of target execution results, and determining a to-be-migrated processing task which is directly migrated to the target processing engine for execution in the plurality of original processing tasks, so as to improve the accuracy and efficiency of checking the data processing task before migration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a data verification method, apparatus, device, and storage medium. Background Art

[0002] Data processing engines such as Hive, Spark, and Impala play a key role in the data processing field. Migrating data processing tasks (or processing tasks) executed on less efficient and more expensive engines to more efficient and cost-effective engines can reduce costs and increase efficiency. Data processing tasks can be Structured Query Language (SQL) tasks, which refer to data processing tasks that contain SQL statements, such as report generation and data analysis. For example, SQL tasks executed in Hive can be migrated to Spark.

[0003] Before migrating a data processing task, the data processing task can be verified to ensure that the migrated data processing task can be executed normally. Currently, verification is often done manually, which has the problem of low accuracy and efficiency. Summary of the Invention

[0004] The present application provides a data verification method, apparatus, device and storage medium, which can improve the accuracy and efficiency of verifying data processing tasks before migration.

[0005] In a first aspect, a data verification method is provided, comprising: determining a plurality of original processing tasks executed in a source processing engine; rewriting the plurality of original processing tasks respectively to obtain a plurality of rewritten processing tasks; sending the plurality of rewritten processing tasks to the source processing engine and the target processing engine, so that the plurality of rewritten processing tasks are executed respectively by the source processing engine and the target processing engine; receiving a plurality of source execution results returned by the source processing engine and a plurality of target execution results returned by the target processing engine; verifying the plurality of original processing tasks according to the plurality of source execution results and the plurality of target execution results, and determining the processing tasks to be migrated from the plurality of original processing tasks for direct migration to the target processing engine for execution.

[0006] In the second aspect, a data verification device is provided, including: a first determination module, used to determine multiple original processing tasks executed in a source processing engine; a task rewriting module, used to rewrite the multiple original processing tasks respectively to obtain multiple rewritten processing tasks; a task sending module, used to send the multiple rewritten processing tasks to the source processing engine and the target processing engine, so that the multiple rewritten processing tasks are executed by the source processing engine and the target processing engine respectively; a result receiving module, used to receive multiple source execution results returned by the source processing engine and multiple target execution results returned by the target processing engine; a second determination module, used to verify the multiple original processing tasks based on the multiple source execution results and the multiple target execution results, and determine the processing tasks to be migrated from the multiple original processing tasks for direct migration to the target processing engine for execution.

[0007] In a third aspect, an electronic device is provided, comprising: a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the first aspect or its various implementations.

[0008] In a fourth aspect, a computer-readable storage medium is provided for storing a computer program, wherein the computer program enables a computer to execute the method according to the first aspect or its various implementations.

[0009] In a fifth aspect, a computer program product is provided, comprising computer program instructions, which enable a computer to execute the method in the first aspect or its various implementations.

[0010] In a sixth aspect, a computer program is provided, which enables a computer to execute the method in the first aspect or its various implementations.

[0011] In summary, the present application can automatically obtain the original processing task and rewrite the original processing task to obtain the rewritten processing task; after that, the rewritten processing task can be executed in the source processing engine and the target processing engine respectively to obtain two execution results; finally, the two execution results can be automatically verified to determine the processing task that can be directly migrated. Therefore, not only can the automated verification of data processing tasks before migration be achieved, the accuracy and efficiency of data verification can be improved; moreover, there is no need to build a test environment, but the production environment is directly used, that is, the source engine and the target engine are directly used to execute the processing task, therefore, a lot of time and resources and potential problems such as insufficient compatibility can be avoided, and comprehensive testing of the processing task can be achieved; in addition, since the processing task executed in the source engine and the target engine is a rewritten processing task, it can also avoid directly using the original task executed from the production environment, that is, the original processing task, to affect the online production data, thereby ensuring the stability and reliability of the online data. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The following is an introduction to the drawings required for describing the embodiments.

[0013] Figure 1 A flowchart of a data verification method provided in an embodiment of the present application;

[0014] Figure 2 A schematic diagram of a data verification method provided in an embodiment of the present application;

[0015] Figure 3 A schematic diagram of another data verification method provided in an embodiment of the present application;

[0016] Figure 4 A schematic diagram of another data verification method provided in an embodiment of the present application;

[0017] Figure 5 A schematic diagram of another data verification method provided in an embodiment of the present application;

[0018] Figure 6 A schematic diagram of a data verification device 600 provided in an embodiment of the present application;

[0019] Figure 7 Schematic diagram of an electronic device 700 provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] The following will introduce various embodiments of the technical solution of this application in conjunction with the drawings in this application.

[0021] It should be noted that the information, data (including, but not limited to: data used for analysis, stored data, displayed data, etc., such as original processing tasks, rewritten processing tasks, source execution results, target execution results, original tables, target tables, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the data processing tasks involved in this application and the operations performed on the data processing tasks are all obtained with full authorization.

[0022] In one embodiment, the technical solution of the present application can be used in data migration scenarios. For example, it can be applied to scenarios where data processing tasks executed on a data processing engine are migrated, and specifically, it can be applied to data verification scenarios such as syntax compatibility and data consistency before actual migration, but is not limited thereto.

[0023] In one embodiment, the solution provided in this application can be executed by any electronic device with data processing capabilities. For example, the electronic device can be a server, specifically an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In another example, the electronic device can be a terminal device, specifically a tablet computer, a laptop computer, or a desktop computer. In another example, the electronic device can be a combination of a server and a terminal device, wherein the server and terminal device in the combination can communicate via wireless or wired means. This application does not impose any specific restrictions on the electronic device.

[0024] In one embodiment, the processing engine involved in the present application may be an online analytical processing (OLAP) engine, for example, a data processing engine such as Hive, Spark, Impala, or Clickhouse.

[0025] The source processing engine can be a low-efficiency, high-cost engine, while the target processing engine can be a high-efficiency, low-cost engine. For example, the source processing engine can be Hive, and the target processing engine can be Spark. The original processing task executed in the source processing engine can be a data processing task containing SQL statements, specifically tasks such as generating reports and performing data analysis.

[0026] It should be noted that the migration of data processing tasks consists of two phases: a pre-migration preparation phase and the actual migration phase. The technical solution of this application primarily relates to the pre-migration preparation phase, which allows for identifying the processing tasks to be migrated, i.e., the original processing tasks, that can be directly migrated to the target processing engine for execution. The following embodiments provide a detailed description of this phase.

[0027] Figure 1 This is a flow chart of a data verification method provided in an embodiment of the present application, which can be executed by the electronic device described above. Figure 1 As shown, the method includes:

[0028] S110: Determine a plurality of original processing tasks executed in a source processing engine;

[0029] S120: rewriting the multiple original processing tasks respectively to obtain multiple rewritten processing tasks;

[0030] S130: Sending the multiple rewriting processing tasks to the source processing engine and the target processing engine, so that the source processing engine and the target processing engine respectively execute the multiple rewriting processing tasks;

[0031] S140: receiving multiple source execution results returned by the source processing engine and multiple target execution results returned by the target processing engine;

[0032] S150: Verify the multiple original processing tasks according to the multiple source execution results and the multiple target execution results, and determine the processing tasks to be migrated for direct migration to the target processing engine for execution among the multiple original processing tasks.

[0033] The following describes various embodiments using an electronic device as a server.

[0034] In one embodiment, S110 may include at least one of the following: receiving a first original processing task submitted by a user based on a client, and receiving a second original processing task collected and executed in a source processing engine. In other words, the original processing task may include at least one of the following: a first original processing task submitted by a user based on a client to a server, and a second original processing task collected and executed in a source processing engine by the server.

[0035] The above-mentioned second original processing task collected and executed in the source processing engine includes: collecting multiple third original processing tasks executed in each data interaction session in the source processing engine; sending the multiple third original processing tasks to a message queue; for the multiple third original processing tasks in the message queue, obtaining the third original processing task corresponding to any execution identifier under the same task identifier to obtain the second original processing task; wherein the task identifier is used to uniquely identify the third original processing task, and the execution identifier is used to uniquely identify the number of executions of the third original processing task. The task identifier or execution identifier can be a universally unique identifier (UUID).

[0036] Specifically, the above-mentioned obtaining the third original processing task corresponding to any execution identifier under each same task identifier to obtain the second original processing task may include: obtaining the third original processing task corresponding to the maximum execution identifier under each same task identifier to obtain the second original processing task.

[0037] For example, assuming that the source processing engine is Hive, the server can automatically collect Hive SQL tasks executed in Hive based on Hive Hook (a hook used to perform corresponding custom operations at different stages of query execution. For example, the "QueryLifeTimeHook" in the following content is a type of Hook). Specifically, Figure 2As shown, the server can first use Hive QueryLifeTimeHook (query lifecycle hook or query lifecycle hook mechanism) to collect all processing tasks executed in each session (data interaction session) in HiveServer2 (the Hive server that provides an interface for remote access to Hive) and send them to Kafka. A processing task includes at least one processing statement, which is divided into two types: SQL-type processing statements and non-SQL-type processing statements, such as commands (commands or instructions, sometimes called "stored procedures" in stored procedures and "conditional statements" in control flow statements). For example, Figure 2 "Session1`s SQL&command" represents the SQL type processing statement and command type processing statement in data interaction dialogue 1, and "Session2`s SQL&command" represents the SQL type processing statement and command type processing statement in data interaction dialogue 2; then, the data in Kafka can be consumed regularly through HiveSQLParser (referring to the component used to parse SQL statements executed in Hive), that is, the above-mentioned collected processing tasks, for example, the collected processing tasks are as follows Figure 2 As shown in "Sessionncommand...SQL...Session m..." in the figure. A processing task on HiveServer2 may be executed multiple times a day (for example, hourly tasks). Each execution of a processing task generates a new execution identifier (Id). To reduce the number of tasks processed in subsequent stages, HiveSQLParser can retain the processing task corresponding to the largest execution identifier in a day. Finally, HiveSQLParser can output the parse result (consumption result or parsing result) and save it to MySQL (MyStructured Query Language, a relational database). The parse result includes: Meta information and all statements included in the processing task. All statements included in the processing task can be SQL statements and command statements. Meta information can include the Session ID (the identifier of the data interaction session), execution ID, task name, project name, and user name. SQL statements can be associated with Session IDs, and one Session ID can correspond to multiple SQL statements and command statements.

[0038] Through the above content, not only can we receive processing tasks submitted by users and verify them before migration, but we can also realize the automated collection of processing tasks executed in the source processing engine to avoid omissions or erroneous collection, and ensure the comprehensiveness and accuracy of processing tasks.

[0039] Moreover, when collecting the second original processing tasks, multiple third original processing tasks executed in each data interaction session in the source processing engine can be screened through the task identifier and execution identifier to avoid collecting duplicate processing tasks, thereby reducing subsequent workload.

[0040] In addition, in the related technology, data processing tasks can be exported from the middle-office scheduling system, but the correct execution order of the processing tasks may not be guaranteed, and some processing tasks may even be missing or operation records unrelated to the migration may be exported. The present application collects and executes multiple third original processing tasks in each data interaction session in the source processing engine. Therefore, it can not only ensure that the collected processing tasks are comprehensive, but also ensure that the collected processing tasks have the correct execution order.

[0041] In one embodiment, for S120, the following steps may be included: for any target original processing task among multiple original processing tasks, splitting the target original processing task into at least one original processing statement; for any target original processing statement among at least one original processing statement, parsing the target original processing statement to obtain a first grammar structure tree of the target original processing statement; according to a preset statement reconstruction rule, adjusting the specific nodes indicated by the statement reconstruction rule in the first grammar structure tree to obtain a second grammar structure tree; deserializing the second grammar structure tree to obtain a rewritten processing statement; assembling at least one rewritten processing statement according to the order of at least one original processing statement in the target original processing task to obtain a rewritten processing task corresponding to the target original processing task.

[0042] Exemplarily, before splitting the target original processing task into at least one original processing statement, the process further includes removing comments in the target original processing task that explain the original processing statement. This prevents content unrelated to the rewritten processing statement from interfering with the statement rewriting, thereby improving the accuracy of statement rewriting while reducing the amount of data processed.

[0043] Exemplarily, the statement reconstruction rule includes at least one of the following: a mapping relationship between an original table and a target table used to replace the original table; and an original field and a replacement rule for the original field.

[0044] Correspondingly, the above-mentioned adjustment of the specific nodes indicated by the sentence reconstruction rules in the first grammar structure tree according to the preset sentence reconstruction rules to obtain the second grammar structure tree may include: traversing the first grammar structure tree to determine the first specific node containing the original table and / or the second specific node containing the original field in the first grammar structure tree; replacing the original table in the first specific node with the corresponding target table according to the mapping relationship, and / or adjusting the original field in the second specific node according to the replacement rule to obtain the second grammar structure tree.

[0045] The statement reconstruction rules may include rules submitted by the user based on the client. In other words, the server can rewrite the original processing statement according to the user-defined rules, which can improve the diversity of task rewriting and user experience.

[0046] In addition, the original table and / or original field can be predetermined so that when traversing the first syntax structure tree, the first specific node containing the original table and / or the second specific node containing the original field can be quickly identified, thereby improving the recognition speed.

[0047] For example, you can predetermine the original fields, including the fields corresponding to the following types of SQL: CREATE TABLE, CREATE TABLE LIKE, CREATE TABLE AS SELECT, ALTER TABLE RENAME TO; and predetermine the library table name corresponding to the original table, including ALTER TABLE.

[0048] For example, some fields in the original processing statement (e.g., fields that do not affect online tasks when executed in the processing engine) may not be replaced, such as those related to Roles, Show, Describe, Explainplan, Queries / Select, Operators and UDFs, Locks, and Authorization. These fields that do not need to be replaced can be recorded in the replacement rules so that these fields are skipped during traversal.

[0049] It is understood that the mapping relationship between the original table and the target table can achieve accurate and comprehensive replacement of the original table, avoiding missing parts of the original table that need to be replaced. In addition, the replacement rules can achieve accurate and fast replacement of the corresponding original fields.

[0050] For example, suppose the mapping relationship between the original table and the target table is:<t,tmp.uuid0_t> , where t is the original table and tmp.uuid0_t is the target table; if the original processing statement is:

[0051] CREATE TABLE t(c int)USING parquet; (uuid0);

[0052] INSERT INTO TABLE t SELECT*FROM t2; (uuid1);

[0053] The rewritten processing statement after rewriting according to the mapping relationship is:

[0054] CREATE TABLE t(c int)USING parquet; (uuid0);

[0055] INSERT INTO TABLE tmp.uuid0_t SELECT*FROM t2;(uuid1).

[0056] For another example, suppose the original processing statement is: INSERT INTO tb; if there is a mapping relationship<tb,uuid_tb> , you can get the uuid_tb corresponding to tb according to the mapping relationship, and directly replace the table INSERT INTO in the statement, and get the rewritten processing statement INSERT INTO uuid_tb; if the mapping relationship does not exist, create a mapping relationship<tb,uuid_tb> , and add the table creation SQL before the INSERT statement to obtain the following rewritten statement:

[0057] CREATE TABLE uuid_tb LIKE tb

[0058] INSERT INTO uuid_tb.

[0059] For example, the replacement rules are exemplified below, wherein the “operation” in each of the following replacement rules refers to the operation / field involved in the original processing statement.

[0060] (1) The replacement rules for the library are shown in Table 1 below.

[0061] Table 1

[0062]

[0063]

[0064] (2) The replacement rules for the table are shown in Table 2 below.

[0065] Table 2

[0066]

[0067] (3) The replacement rules for partitions in the replacement rules are shown in Table 3 below.

[0068] Table 3

[0069]

[0070] (4) The replacement rules for fields in the replacement rules are shown in Table 4 below.

[0071] Table 4

[0072]

[0073]

[0074] (5) The replacement rules for views in the replacement rules are shown in Table 5 below.

[0075] Table 5

[0076]

[0077] (6) The replacement rules for materialized views are shown in Table 6 below.

[0078] Table 6

[0079]

[0080] (7) The replacement rules for the index Index in the replacement rules are shown in Table 7 below.

[0081] Table 7

[0082]

[0083] (8) The replacement rules for Function in the replacement rules are shown in Table 8 below.

[0084] Table 8

[0085]

[0086]

[0087] (9) The replacement rules for Roles are shown in Table 9 below.

[0088] Table 9

[0089]

[0090] (10) The replacement rules related to Load are shown in Table 10 below.

[0091] Table 10

[0092]

[0093] (11) The replacement rules related to Insert are shown in Table 11 below.

[0094] Table 11

[0095]

[0096] (12) The replacement rules for parameters in the replacement rules are shown in Table 12 below.

[0097] Table 12

[0098]

[0099]

[0100] (13) The replacement rules related to Update are shown in Table 13 below.

[0101] Table 13

[0102]

[0103] (14) The replacement rules related to Delete are shown in Table 14 below.

[0104] Table 14

[0105]

[0106] (15) The replacement rules for Merge are shown in Table 15 below.

[0107] Table 15

[0108]

[0109] (16) The replacement rules related to Import / Export in the replacement rules are shown in Table 16 below.

[0110] Table 16

[0111]

[0112] In one embodiment, Figure 3 As shown, the server can first determine the original processing task, that is, the task before rewriting, such as Figure 3As shown in the top box of ; then, the SQL preprocessor can be used to remove all useless comments in the original processing task and split the original processing task into single independent original processing statements arranged in sequence (the sequence here refers to the sequence between the statements in the original processing task, such as the execution order), such as SOL statements, as shown in Figure 3 The content of the second box from the top is shown in the figure, such as "set k=v"; then, the SQL statement can be parsed by the SQL rewriter based on Antlr (ANother Tool for Language Recognition) or Calcite (Calcite framework or dynamic data management framework) to generate the following Figure 4 The Abstract Syntax Tree (AST) shown in (a) in the figure can be used. Then, some nodes in the abstract syntax tree can be replaced based on the statement reconstruction replacement rules to generate a reconstructed Antlr TokenRewriteStream character stream (the result obtained after using TokenRewriteStream, TokenRewriteS tream is a class of the Antlr tool, which is used to modify the token stream (TokenStrea m, which is a structured data sequence generated after lexical analysis of the processing statement) during the parsing process, for example, to replace the content or adjust the structure); then, the reconstructed character stream can be reversely parsed to generate a rewritten SQL statement, that is, a rewritten processing statement; finally, the rewritten SQL statement can be assembled into a complete rewritten task according to the processing order of the SQL preprocessor to obtain a rewritten processing task, such as Figure 3 The contents of the last box are shown in .

[0113] Among them, such as Figure 4As shown in the figure, assuming that the original processing statement is "CREATE EXTERNAL TABLE IF NOTEXISTS t(iint,d desc)STORED AS orc", it can be parsed into an AST tree by using the SQL rewriter based on the Antlr / Calcite tool using the Hive syntax parsing file (for example, HiveParser.g4). The main structure of the AST tree is createTableStatement, where uppercase characters such as KW_CREATE, KW_EXTERNAL, KW_TABLE, and LPAREN represent keywords in the syntax tree, and subtrees such as tableName, columnNameTypeOrConstraintList, and tableFileForma t constitute the body of createTableStatement; then, a custom syntax replacement rule can be used to perform a replacement operation on the tableName subtree of the AST tree to generate the following: Figure 4 In the new AST tree shown in (b), the tableName in the replaced new AST tree changes from the actual production table 't' to the temporary table 'tmp.uuid_0'. This ensures that the processing engine does not affect or change the online production data when executing the processing task. Finally, the AST tree can be deserialized to generate the rewritten Hiv e SQL statement, that is, the rewritten processing statement is: CREATE EXTERNAL TABLE IF NOT EXISTS tmp.uuid_0(iint,d dec) STORED AS orc.

[0114] In the above content, the original processing statements can be reconstructed and rewritten based on syntax parsing tools (for example, Antlr and Calcite), and the nodes of the abstract syntax tree can be replaced according to precise statement replacement rules to generate high-quality rewritten processing statements, thereby improving the accuracy of syntax replacement, ensuring the consistency and standardization of syntax rewriting and verification, and reducing new errors that may be introduced by manual rewriting.

[0115] Moreover, by rewriting the original processing tasks, it is possible to ensure that the subsequent processing engine executes the rewritten processing tasks rather than the original processing tasks, avoiding the impact on online production data caused by executing non-rewritten tasks in the production environment. For example, when replacing key elements such as table names and library names, reasonable rules are used to ensure data integrity and consistency, thereby ensuring the stability and accuracy of data during the migration process, avoiding repeated debugging and processing due to data problems, and further improving overall efficiency.

[0116] In addition, the above process can also assemble the rewritten processing statements according to the order of the original processing statements in the original processing task. Therefore, it can ensure that the order of processing statements in the rewritten processing task is consistent with the order of processing statements in the original processing task, and further ensure that except for some statements replaced based on statement replacement rules, other processing statements of the processing task have no changes, thereby ensuring the rationality of subsequent verification based on the execution results and improving the accuracy of data verification.

[0117] In one embodiment, for S130-S140, the source processing engine can be Hive. When Hive executes the rewrite processing task, it can be specifically implemented through the Hive execution engine server: HiveServer2; the target processing engine can be Spark (a big data computing engine). When Spark executes the rewrite processing task, it can be specifically implemented through the Spark execution engine server, wherein the Spark execution engine server can be connected through Kyuubi (a client), and this application does not impose any restrictions on this.

[0118] It is understandable that if the processing task is executed by building a test environment with the same services as the production environment (source engine and target engine), it will cause a waste of time and resources. For example, a large amount of hardware configuration, software installation and data migration resources will be consumed for compatibility testing; moreover, this method can only cover a limited number of test cases, and it is difficult to fully detect all possible compatibility issues. For example, for some complex processing statement queries, such as those involving multi-table associations, subqueries and function nesting, copying only part of the data may not be able to discover potential problems. These problems may be exposed in the actual production environment, bringing risks to the business. However, the present application can directly execute the rewritten original processing task, that is, the rewritten processing task, in the source processing engine and the target processing engine respectively, that is, directly use the production environment. Therefore, it can avoid the waste of a lot of time and resources and the emergence of potential problems such as insufficient compatibility, and achieve comprehensive detection of processing tasks and ensure compatibility during task processing.

[0119] In one embodiment, for S250, the following steps may be included: determining whether the source execution results and the corresponding target execution results corresponding to each of the multiple original processing tasks contain error information; determining a first specific processing task among the multiple original processing tasks, in which the corresponding source execution results and the corresponding target execution results do not contain error information; and determining the processing task to be migrated based on the first specific processing task.

[0120] Specifically, the above-mentioned determination of the processing task to be migrated based on the first specific processing task may include the following two situations:

[0121] In case 1, the first specific processing task is determined as the processing task to be migrated.

[0122] That is, in the first case, the original processing task corresponding to the execution result that does not contain error information can be determined as the processing task to be migrated that can be used for direct migration, which can improve the data verification efficiency.

[0123] Case 2: For the first specific processing task, determine the source task output data contained in the corresponding source execution result and the target task output data contained in the corresponding target execution result; compare the source task output data and the target task output data; and determine the processing tasks corresponding to the consistent source task output data and target task output data in the first specific processing task as the processing tasks to be migrated.

[0124] For example, the source task output data includes at least one of the following: table information of the source specific table, the total number of records of the source task output data, and the hash value (Hash value) of the data generated by the source task; the target task output data includes at least one of the following: table information of the target specific table, the total number of records of the target task output data, and the hash value of the data generated by the target task.

[0125] The table information may be the schema information of the table; the specific table may be the target object table of the INSERT in the task or a new table created using the CTAS syntax, but is not limited thereto.

[0126] That is to say, in case 2, the original processing task corresponding to the execution result that does not contain error information can be further verified. Specifically, the task output data corresponding to the two processing engines can be compared to determine the processing tasks to be migrated that can be used for direct migration, which can ensure the accuracy of data verification.

[0127] It is understandable that the successful execution of a processing task in a processing engine only indicates that the processing task can be migrated from the source processing engine to the target processing engine based on grammatical adaptation. However, in addition to grammatical adaptation, the correctness of the data output by the processing task after migration should also be guaranteed to truly migrate the processing task from the source processing engine to the target processing engine, thereby ensuring the accuracy of data migration to a greater extent. Therefore, through the second scenario, the correctness of the final data output of the successfully executed processing task can be verified, that is, the consistency of the data in the execution results corresponding to the source processing engine and the target processing engine can be verified to ensure the accuracy of data migration to a greater extent.

[0128] Compared with the method of manually writing data verification scripts, the above embodiment does not require writing data verification scripts, which can avoid the problem of obtaining incorrect verification results due to the error-prone writing of verification scripts; moreover, the above embodiment can realize the automation and timeliness requirements of verification execution results, avoid missing some compatibility and data consistency issues and bringing risks to the post-migration business, and can improve the reliability of migration.

[0129] In one embodiment, before determining the processing task to be migrated based on the first specific processing task, the above-mentioned steps also include: determining a retry processing task in which the corresponding source execution result and the corresponding target execution result contain error information among multiple original processing tasks; sending the retry processing task to the source processing engine and the target processing engine, so that the retry processing task is executed again by the source processing engine and the target processing engine respectively.

[0130] Specifically, the above-mentioned re-execution of the retry processing task by the source processing engine and the target processing engine respectively may include: executing the retry processing task a maximum of preset times by the source processing engine and the target processing engine respectively to avoid waste of resources and time caused by multiple retries and improve the retry success rate.

[0131] It is understandable that when the processing resources of the processing engine are insufficient, the execution results may also contain error information. Therefore, the retry processing task containing error information can be re-executed by retrying to avoid mistakenly identifying the processing task that can be executed correctly as containing error information.

[0132] In one embodiment, the server can also collect statistics on the tasks to be migrated, the output data of each source task, the output data of each target task, and the comparison results to obtain a task migration analysis report; and send the task migration analysis report to the client so that the client can display the task migration analysis report. The task migration analysis report can include various error information.

[0133] In addition, the server can persistently store the above execution results and task migration analysis reports.

[0134] Through this embodiment, a detailed compatibility report can be automatically generated without the need to manually record and analyze test results, thereby reducing user workload and avoiding omissions and misjudgments.

[0135] Furthermore, the server can perform statistical analysis from multiple angles: the verification object, that is, the execution results of the source processing engine and the target processing engine, can clearly identify which tables have been verified and which fields have been excluded; the total number of records, that is, the output data of each task, can clearly identify the total number of records of the output data of the processing task in the processing engine; the total hash value, that is, the hash value of the output data of the processing task in the processing engine. Thus, the execution results and verification results of all the above stages can be fully and comprehensively provided to the user for confirmation, making it easier to identify which specific stage has an error so that the user can analyze the problem.

[0136] In one embodiment, before S120, the following steps may also be included: for any target original processing task among multiple original processing tasks, converting the processing statement in the target original processing task into a task execution plan query statement to obtain a pre-execution processing task; sending the pre-execution processing task to the source processing engine and the target processing engine, so that the pre-execution processing task is executed by the source processing engine and the target processing engine respectively; receiving multiple source pre-execution results returned by the source processing engine and multiple target pre-execution results returned by the target processing engine; determining multiple second specific processing tasks in which the corresponding source pre-execution results and the corresponding target pre-execution results are successfully executed among the multiple original processing tasks.

[0137] Correspondingly, S120 includes: rewriting the plurality of second specific processing tasks to obtain a plurality of rewritten processing tasks.

[0138] Exemplarily, taking the case where the processing statements in the original processing task include SQL statements and Command statements, the server can modify the SQL statements and add the EXPLA IN keyword before all SQL statements to convert the SQL statements into execution plan query statements; if the source processing engine and the target processing engine are both compatible with Command statements, then the Command statements can be obtained without modifying the Command statements. The above-mentioned execution plan query statements can then be processed by the source processing engine and the target processing engine to output the execution plan phase of the SQL statement and the relationship between different phases. In addition, this process will not actually execute the original processing task, and it can be ensured that all operations performed at this stage will not affect the online data. Finally, if the execution results corresponding to the source processing engine and the target processing engine are both successfully executed, then the next stage of operation, i.e., S120, will be executed; otherwise, an error can be directly returned to the client.

[0139] Through this embodiment, before the task is rewritten, the original processing task can be initially verified, specifically by executing the task execution plan query statement of the processing task. During this process, the original processing task will not be actually executed, but only the pre-execution results (i.e., the execution plan stage of the processing task and the relationship between different stages) are compared to determine whether the source processing engine and the target processing engine can correctly execute the processing statement corresponding to the original processing task. Therefore, not only can the impact on the online data corresponding to the processing engine be avoided, but also a verification can be performed with a lower data processing volume, thereby improving data verification efficiency.

[0140] The following is an introduction to the technical solution of this application by means of a schematic diagram:

[0141] In one embodiment, in combination with the above embodiments, Figure 5As shown, the client (Client) can send a data verification request (Send Requst) to the server (Server), which includes the first original processing task and the statement reconstruction rule; then, the server can generate a task execution plan query statement (Pre-ExecuteSQL) corresponding to the original processing task, and send the task execution plan query statement to HiveServer2 and Kyuubi for execution respectively; after that, it can receive the execution results (End of Pre-Execute) returned by HiveServer2 and Kyuubi respectively. If both executions fail, an error message (Abort if Pre-Execute Failed) is directly returned to the client, indicating that the original processing task cannot be migrated. If the execution is successful, the subsequent steps are continued. Next, the server can generate a unique uuid (Gen tmp uu id) to independently identify each processing task. Then, based on the replacement rules, namely the statement reconstruction rules, the server can use the syntax parsing tool to parse the original processing statement, reconstruct the abstract syntax tree, and generate a rewritten processing statement (Rewrite SQL). The rewritten processing statement can then be sent to HiveServer2 and Kyuubi for execution (Execute SQL) and receive an execution completion (End of execute). For data operations (operations in the rewritten processing statement, such as write operations, etc.), a query SQL statement is generated for data validation (Gen Validation SQL For DML) and sent to HiveServer2 and Kyuubi for execution. After the query SQL statement is executed, the result is returned to the server (Return result set). The server can compare the above-mentioned results, namely the execution results of the rewritten statement by HiveServer2 and Kyuubi, and generate a result report (Compare result set and generate report). Finally, the result report is returned to the client.

[0142] It should be noted that all the above technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0143] Figure 6 A schematic diagram of a data verification device 600 provided in an embodiment of the present application is shown as follows: Figure 6As shown, the device 600 includes: a first determination module 601, a task rewriting module 602, a task sending module 603, a result receiving module 604, a second determination module 605, a content removal module 606, a report statistics module 607, a report sending module 608, an error determination module 609, an error retry module 610, a statement conversion module 611, a sending execution module 612, a data receiving module 613, and a task selection module 614.

[0144] In one embodiment, a first determination module 601 is used to determine multiple original processing tasks executed in a source processing engine; a task rewriting module 602 is used to rewrite the multiple original processing tasks respectively to obtain multiple rewritten processing tasks; a task sending module 603 is used to send the multiple rewritten processing tasks to the source processing engine and the target processing engine, so that the multiple rewritten processing tasks are executed by the source processing engine and the target processing engine respectively; a result receiving module 604 is used to receive multiple source execution results returned by the source processing engine and multiple target execution results returned by the target processing engine; a second determination module 605 is used to verify the multiple original processing tasks based on the multiple source execution results and the multiple target execution results, and determine the processing tasks to be migrated from the multiple original processing tasks for direct migration to the target processing engine for execution.

[0145] Exemplarily, the task rewriting module 602 is specifically used to: for any target original processing task among multiple original processing tasks, split the target original processing task into at least one original processing statement; for any target original processing statement among at least one original processing statement, parse the target original processing statement to obtain a first syntax structure tree of the target original processing statement; according to a preset statement reconstruction rule, adjust the specific nodes indicated by the statement reconstruction rule in the first syntax structure tree to obtain a second syntax structure tree; deserialize the second syntax structure tree to obtain a rewritten processing statement; assemble at least one rewritten processing statement according to the order of at least one original processing statement in the target original processing task to obtain a rewritten processing task corresponding to the target original processing task.

[0146] Exemplarily, the content removal module 606 is specifically configured to remove annotation content that explains the original processing statement in the target original processing task.

[0147] Exemplarily, the statement reconstruction rule includes at least one of the following: a mapping relationship between an original table and a target table used to replace the original table; an original field and a replacement rule for the original field; correspondingly, the task rewriting module 602 is specifically used to: traverse the first syntax structure tree, determine the first specific node containing the original table and / or the second specific node containing the original field in the first syntax structure tree; replace the original table in the first specific node with the corresponding target table according to the mapping relationship, and / or adjust the original field in the second specific node according to the replacement rule to obtain a second syntax structure tree.

[0148] Exemplarily, the statement reconstruction rules include rules submitted by the user based on the client.

[0149] Exemplarily, the second determination module 605 is specifically used to: determine whether the source execution results and the corresponding target execution results corresponding to multiple original processing tasks contain error information; determine the first specific processing task among the multiple original processing tasks, in which the corresponding source execution results and the corresponding target execution results do not contain error information; and determine the processing task to be migrated based on the first specific processing task.

[0150] Illustratively, the second determining module 605 is specifically configured to: determine the first specific processing task as the processing task to be migrated.

[0151] Exemplarily, the second determination module 605 is specifically used to: determine, for the first specific processing task, the source task output data contained in the corresponding source execution result and the target task output data contained in the corresponding target execution result; compare the source task output data and the target task output data; and determine the processing tasks corresponding to the consistent source task output data and target task output data in the first specific processing task as processing tasks to be migrated.

[0152] Exemplarily, the source task output data includes at least one of the following: table information of the source specific table, the total number of records of the source task output data, and the hash value of the data generated by the source task; the target task output data includes at least one of the following: table information of the target specific table, the total number of records of the target task output data, and the hash value of the data generated by the target task.

[0153] Exemplarily, the report statistics module 607 is used to: perform statistics on the tasks to be migrated, the output data of each source task, the output data of each target task, and each comparison result to obtain a task migration analysis report; the report sending module 608 is used to: send the task migration analysis report to the client so that the client can display the task migration analysis report.

[0154] Exemplarily, the error determination module 609 is used to: determine the retry processing task whose corresponding source execution results and corresponding target execution results contain error information among multiple original processing tasks; the error retry module 610 is used to: send the retry processing task to the source processing engine and the target processing engine, so that the retry processing task can be executed again by the source processing engine and the target processing engine respectively.

[0155] Exemplarily, the statement conversion module 611 is used to: for any target original processing task among multiple original processing tasks, convert the processing statement in the target original processing task into a task execution plan query statement to obtain a pre-execution processing task; the sending execution module 612 is used to: send the pre-execution processing task to the source processing engine and the target processing engine, so that the pre-execution processing task is executed by the source processing engine and the target processing engine respectively; the data receiving module 613 is used to: receive multiple source pre-execution results returned by the source processing engine and multiple target pre-execution results returned by the target processing engine; the task selection module 614 is used to: determine multiple second specific processing tasks in which the corresponding source pre-execution results and the corresponding target pre-execution results are successfully executed among multiple original processing tasks; the task rewriting module 602 is specifically used to rewrite multiple second specific processing tasks to obtain multiple rewritten processing tasks.

[0156] Exemplarily, the first determining module 601 is specifically configured to: receive a first original processing task submitted by a user based on a client; and collect and execute second original processing tasks in a source processing engine.

[0157] Exemplarily, the first determination module 601 is specifically used to: collect multiple third original processing tasks executed in each data interaction session in the source processing engine; send the multiple third original processing tasks to the message queue; for the multiple third original processing tasks in the message queue, obtain the third original processing task corresponding to any execution identifier under the same task identifier to obtain the second original processing task; wherein the task identifier is used to uniquely identify the third original processing task, and the execution identifier is used to uniquely identify the number of executions of the third original processing task.

[0158] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Figure 6 The device 600 shown can execute the above method embodiment, and the above and other operations and / or functions of each module in the device 600 are respectively for implementing the corresponding processes in the above method, which will not be repeated here for the sake of brevity.

[0159] The above describes the device 600 of the embodiment of the present application from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.

[0160] Figure 7 A schematic diagram of an electronic device 700 provided in an embodiment of the present application.

[0161] like Figure 7 As shown, the electronic device 700 may include:

[0162] The memory 710 and the processor 720 are configured to store computer programs and transmit the program code to the processor 720. In other words, the processor 720 can call and run the computer program from the memory 710 to implement the method in the embodiment of the present application.

[0163] For example, the processor 720 may be configured to execute the above method embodiments according to instructions in the computer program.

[0164] In some embodiments of the present application, the processor 720 may include but is not limited to:

[0165] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0166] In some embodiments of the present application, the memory 710 includes but is not limited to:

[0167] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SL DRAM), and direct RAM bus random access memory (DR RAM).

[0168] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 710 and executed by the processor 720 to implement the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0169] like Figure 7 As shown, the electronic device may further include:

[0170] The transceiver 730 may be connected to the processor 720 or the memory 710 .

[0171] The processor 720 may control the transceiver 730 to communicate with other devices. Specifically, the processor 720 may send information or data to other devices or receive information or data sent by other devices. The transceiver 730 may include a transmitter and a receiver. The transceiver 730 may further include one or more antennas.

[0172] It should be understood that the various components in the electronic device are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0173] The present application also provides a computer storage medium having a computer program stored thereon, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment. In other words, the present application also provides a computer program product containing instructions, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment.

[0174] When software is used to implement, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instruction is loaded and executed on a computer, the computer can be made to perform the corresponding flow in each method in the embodiment of the present application, generate the function that each method in the embodiment of the present application can realize in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instruction can be stored in a computer-readable storage medium, or transmitted from a computer-readable storage medium to another computer-readable storage medium. For example, the computer instruction can be transmitted from a website, computer, server, or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, computer, server, or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).

[0175] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0176] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the system, device or module can be electrical, mechanical or other forms.

[0177] Modules described as separate components may or may not be physically separate, and components displayed as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected based on actual needs to achieve the purpose of the present embodiment. For example, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module.

Claims

1. A data verification method, characterized in that: include: determining a plurality of raw processing tasks to be executed in a source processing engine; Rewriting the plurality of original processing tasks respectively to obtain a plurality of rewritten processing tasks; Sending the plurality of rewriting processing tasks to the source processing engine and the target processing engine, so that the plurality of rewriting processing tasks are executed by the source processing engine and the target processing engine respectively; receiving a plurality of source execution results returned by the source processing engine and a plurality of target execution results returned by the target processing engine; The plurality of original processing tasks are verified according to the plurality of source execution results and the plurality of target execution results, and processing tasks to be migrated for direct migration to the target processing engine for execution are determined among the plurality of original processing tasks.

2. The method according to claim 1, characterized in that The rewriting of the plurality of original processing tasks to obtain a plurality of rewritten processing tasks includes: For any target original processing task among the multiple original processing tasks, split the target original processing task into at least one original processing statement; For any target original processing statement in the at least one original processing statement, the target original processing statement is parsed to obtain a first syntax structure tree of the target original processing statement; According to a preset sentence reconstruction rule, adjusting a specific node indicated by the sentence reconstruction rule in the first grammar structure tree to obtain a second grammar structure tree; Deserializing the second syntax structure tree to obtain a rewriting processing statement; According to the order of at least one original processing statement in the target original processing task, at least one rewritten processing statement is assembled to obtain a rewritten processing task corresponding to the target original processing task.

3. The method according to claim 2, characterized in that Before splitting the target original processing task into at least one original processing statement, the method further includes: Remove the comment content in the target original processing task that explains the original processing statement.

4. The method according to claim 2, characterized in that The statement reconstruction rule includes at least one of the following: A mapping relationship between an original table and a target table used to replace the original table; Original fields and replacement rules for the original fields; Correspondingly, adjusting the specific node indicated by the sentence reconstruction rule in the first grammar structure tree according to the preset sentence reconstruction rule to obtain the second grammar structure tree includes: Traversing the first syntax structure tree to determine that the first syntax structure tree includes a first specific node of the original table and / or a second specific node of the original field; The original table in the first specific node is replaced with the corresponding target table according to the mapping relationship, and / or the original field in the second specific node is adjusted according to the replacement rule to obtain the second syntax structure tree.

5. The method according to claim 2, characterized in that The statement reconstruction rules include rules submitted by the user based on the client.

6. The method according to claim 1, characterized in that The verifying the plurality of original processing tasks according to the plurality of source execution results and the plurality of target execution results to determine the to-be-migrated processing tasks for directly migrating to the target processing engine for execution among the plurality of original processing tasks includes: Determine whether the source execution results and the target execution results corresponding to the multiple original processing tasks contain error information; Determine a first specific processing task among the multiple original processing tasks, the first specific processing task for which the corresponding source execution result and the corresponding target execution result do not include error information; The processing task to be migrated is determined according to the first specific processing task.

7. A data verification device, characterized in that: include: A first determining module, configured to determine a plurality of original processing tasks executed in a source processing engine; A task rewriting module, configured to rewrite the plurality of original processing tasks respectively to obtain a plurality of rewritten processing tasks; a task sending module, configured to send the plurality of rewriting processing tasks to the source processing engine and the target processing engine, so that the plurality of rewriting processing tasks are executed by the source processing engine and the target processing engine respectively; A result receiving module, configured to receive a plurality of source execution results returned by the source processing engine and a plurality of target execution results returned by the target processing engine; The second determining module is used to verify the multiple original processing tasks according to the multiple source execution results and the multiple target execution results, and determine the processing tasks to be migrated from the multiple original processing tasks for direct migration to the target processing engine for execution.

8. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to perform the method according to any one of claims 1 to 6 by executing the executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising instructions, characterized in that When the computer program product is run on an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 6.