Data consistency verification method, electronic equipment, storage medium and program product

By generating and comparing the query result signatures of new and old version database clusters, the inefficient data consistency verification problem in the upgrade of large versions of the database cluster is solved, and efficient and accurate data consistency checksum exception handling is achieved.

CN120371822APending Publication Date: 2025-07-25KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510377568.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When upgrading large versions of database clusters, comparing the field values of new and old versions of database clusters one by one in the prior art leads to inefficient data consistency verification.

Method used

By generating the signature of the query results of the old and new versions of the database cluster, the data consistency verification is performed using hash calculation and signature statistics to reduce the direct comparison of field values.

Benefits of technology

Improves the efficiency of data consistency verification, simplifies verification logic, reduces the possibility of misjudgment, and automatically recognizes and processes the causes of failure of verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371822A_ABST
    Figure CN120371822A_ABST
Patent Text Reader

Abstract

The invention provides a data consistency verification method, electronic equipment, a storage medium and a program product. The data consistency verification method comprises the steps that a first query result and a second query result are obtained, the first query result is a query result obtained after a new-version database cluster executes a target query instruction, and the second query result is a query result obtained after an old-version database cluster executes the target query instruction; the first query result and the second query result respectively comprise a plurality of queried records; a signature of each record in the multiple records is generated, a first signature result and a second signature result are obtained, the first signature result comprises the signature of each record in the multiple records of the first query result, the second signature result comprises the signature of each record in the multiple records of the second query result, and the signatures are used for representing the corresponding records; and performing data consistency verification on the first signature result and the second signature result to obtain a data consistency verification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data verification, etc., and particularly relates to a data consistency verification method, an electronic device, a storage medium, and a program product. Background Art

[0002] As the last link in the big data ecosystem, the OLAP (Online Analytical Processing) engine undertakes the important responsibilities of data storage and query. When performing a major version upgrade on the engine, in order to avoid data inconsistency between the new version database cluster and the old version database cluster, data consistency verification needs to be performed before the new version database cluster is put into production.

[0003] In the related art, each record stored in the database cluster includes field values of multiple fields. When performing data consistency verification, it is necessary to perform data comparison on the field values in the new version database cluster and the old version database cluster one by one.

[0004] However, since the number of records stored in the database cluster is large, that is, the number of field values is large, therefore, performing data comparison on the field values one by one has the problem of low efficiency. Summary of the Invention

[0005] The present disclosure provides a data consistency verification method, an electronic device, a storage medium, and a program product.

[0006] According to one aspect of the present disclosure, there is provided a data consistency verification method, including: obtaining a first query result and a second query result, where the first query result is the query result obtained after the new version database cluster executes a target query instruction, the second query result is the query result obtained after the old version database cluster executes the target query instruction, and the first query result and the second query result respectively include multiple records queried; respectively generating signatures for each record in the multiple records to obtain a first signature result and a second signature result, where the first signature result includes the signatures of each record in the multiple records of the first query result, the second signature result includes the signatures of each record in the multiple records of the second query result, and the signature is used to represent the corresponding record; and performing data consistency verification on the first signature result and the second signature result to obtain a data consistency verification result.

[0007] According to the data consistency verification method of at least one embodiment of the present disclosure, each of the records includes field values of at least one field; generating signatures for each of the multiple records respectively, including: for each of the multiple records, concatenating the field values of at least one field of the record to obtain a string corresponding to the record; performing a hash calculation on the string corresponding to the record to obtain a hash value corresponding to the record; and using the hash value corresponding to the record as the signature of the corresponding record in the multiple records.

[0008] According to the data consistency verification method of at least one embodiment of the present disclosure, performing data consistency verification on the first signature result and the second signature result to obtain a data consistency verification result, including: performing statistics on the signatures in the first signature result to obtain a first statistical result, and performing statistics on the signatures in the second signature result to obtain a second statistical result; and performing data consistency verification on the first statistical result and the second statistical result to obtain the data consistency verification result.

[0009] According to the data consistency verification method of at least one embodiment of the present disclosure, performing statistics on the signatures in the first signature result to obtain a first statistical result, and performing statistics on the signatures in the second signature result to obtain a second statistical result, including: for the first signature result and the second signature result, respectively determining the quantity of each signature in the signature result where the signature is located; generating key-value pairs with the signature as the key and the quantity of the signature in the signature result where the signature is located as the value; and using the key-value pairs corresponding to the signatures in the first signature result as the first statistical result, and using the key-value pairs corresponding to the signatures in the second signature result as the second statistical result.

[0010] According to the data consistency verification method of at least one embodiment of the present disclosure, performing data consistency verification on the first statistical result and the second statistical result to obtain the data consistency verification result, including: determining whether the first statistical result is the same as the second statistical result; in response to the first statistical result being the same as the second statistical result, determining that the data consistency verification result is verification passed; and in response to the first statistical result being different from the second statistical result, determining that the data consistency verification result is verification failed.

[0011] According to the data consistency verification method of at least one embodiment of the present disclosure, after performing data consistency verification on the first signature result and the second signature result to obtain a data consistency verification result, it further includes: in response to the data consistency verification result being verification failed, based on a first rewriting rule, rewriting the field values of at least one field in the first target object to obtain a rewritten first target object, where the first target object is one of the first query result and the second query result; comparing the rewritten first target object with a second target object, where the second target object is the other of the first query result and the second query result; and in the case where the rewritten first target object is consistent with the second target object, determining that the reason for the verification failure is the preset reason corresponding to the first rewriting rule.

[0012] According to the data consistency verification method of at least one embodiment of the present disclosure, after comparing the rewritten first target object with the second target object, it further includes: in the case where the rewritten first target object is inconsistent with the second target object, based on a second rewriting rule, respectively rewriting the field values of at least one field in the first query result and the field values of at least one field in the second query result to obtain a rewritten first query result and a rewritten second query result; comparing the rewritten first query result with the rewritten second query result; and in the case where the rewritten first query result is consistent with the rewritten second query result, determining that the reason for the verification failure is the preset reason corresponding to the second rewriting rule.

[0013] According to the data consistency verification method of at least one embodiment of the present disclosure, before obtaining the first query result and the second query result, it further includes: obtaining a plurality of preset query instructions; respectively calculating the signatures of each of the plurality of preset query instructions; screening the plurality of preset query instructions based on the signatures of the preset query instructions to obtain a plurality of preset query instructions with different signatures; and determining the target query instruction according to the plurality of preset query instructions with different signatures.

[0014] According to the data consistency verification method of at least one embodiment of the present disclosure, respectively calculating the signatures of each of the plurality of preset query instructions includes: for each of the plurality of preset query instructions, determining the string corresponding to the preset query instruction; performing a hash calculation on the string corresponding to the preset query instruction to obtain the hash value corresponding to the preset query instruction; and using the hash value corresponding to the preset query instruction as the signature of the corresponding preset query instruction among the plurality of preset query instructions.

[0015] The method for data consistency verification according to at least one embodiment of the present disclosure to determine the string corresponding to the preset query instruction includes: converting the preset query instruction into an abstract syntax tree to obtain the abstract syntax tree corresponding to the preset query instruction; replacing the parameter values in the abstract syntax tree corresponding to the preset query instruction with preset strings; and parsing the replaced abstract syntax tree corresponding to the preset query instruction into a string to obtain the string corresponding to the preset query instruction.

[0016] The method for data consistency verification according to at least one embodiment of the present disclosure to determine the target query instruction according to the plurality of preset query instructions with different signatures includes: deleting the aggregation operator in the first query instruction and retaining the aggregation object corresponding to the aggregation operator in the first query instruction to obtain the processed first query instruction, where the first query instruction is the preset query instruction with an aggregation operator among the plurality of preset query instructions with different signatures; and using the processed first query instruction and the second query instruction as the target query instruction, where the second query instruction is the preset query instruction without an aggregation operator among the plurality of preset query instructions with different signatures.

[0017] The method for data consistency verification according to at least one embodiment of the present disclosure to determine the target query instruction according to the plurality of preset query instructions with different signatures includes: generating a target sorting rule according to the query object in the third query instruction, where the third query instruction is the preset query instruction that limits the number of records in the query result and has no sorting rule among the plurality of preset query instructions with different signatures; adding the target sorting rule to the third query instruction to obtain the processed third query instruction, where the processed third query instruction is used to sort the multiple records corresponding to the query object based on the target sorting rule and select the top number of records from the sorting result as the query result; and using the processed third query instruction and the fourth query instruction as the target query instruction, where the fourth query instruction is the preset query instruction other than the third query instruction among the plurality of preset query instructions with different signatures.

[0018] According to another aspect of the present disclosure, there is provided an electronic device, including: a memory that stores execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the data consistency verification method according to any one of the embodiments of the present disclosure.

[0019] According to still another aspect of the present disclosure, there is provided a readable storage medium that stores execution instructions, and when the execution instructions are executed by a processor, they are used to implement the data consistency verification method according to any one of the embodiments of the present disclosure.

[0020] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the data consistency verification method according to any one of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. The drawings are included herein to provide a further understanding of the present disclosure and are part of this specification.

[0022] Figure 1 is a schematic flowchart of the data consistency verification method according to an embodiment of the present disclosure.

[0023] Figure 2 is a schematic diagram of the process of generating a signature of a record according to an embodiment of the present disclosure.

[0024] Figure 3 is a schematic diagram of the process of verifying data consistency according to an embodiment of the present disclosure.

[0025] Figure 4 is a schematic diagram of the process of counting signatures according to an embodiment of the present disclosure.

[0026] Figure 5 is a visualization example diagram of the signature counting process according to an embodiment of the present disclosure.

[0027] Figure 6 is a schematic diagram of the process of verifying data consistency according to another embodiment of the present disclosure.

[0028] Figure 7 is a schematic diagram of the process of determining the reason for failed verification according to an embodiment of the present disclosure.

[0029] Figure 8 is a schematic diagram of the process of determining the reason for failed verification according to another embodiment of the present disclosure.

[0030] Figure 9 is a schematic diagram of the process of determining a target query instruction according to an embodiment of the present disclosure.

[0031] Figure 10 is a schematic diagram of the process of generating a signature of a preset query instruction according to an embodiment of the present disclosure.

[0032] Figure 11 is a schematic diagram of the process of determining a string corresponding to a preset query instruction according to an embodiment of the present disclosure.

[0033] Figure 12It is a schematic diagram of the process for determining a target query instruction according to another embodiment of the present disclosure.

[0034] Figure 13 It is a visualization example diagram of a first query instruction and a processed first query instruction according to an embodiment of the present disclosure.

[0035] Figure 14 It is a schematic diagram of the process for determining a target query instruction according to still another embodiment of the present disclosure.

[0036] Figure 15 It is a visualization example diagram of a third query instruction and a processed third query instruction according to an embodiment of the present disclosure.

[0037] Figure 16 It is a schematic flowchart of a data consistency verification method according to another embodiment of the present disclosure.

[0038] Figure 17 It is a schematic block diagram of a data consistency verification device according to an embodiment of the present disclosure.

[0039] Figure 18 It is a schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Embodiments

[0040] The present disclosure will be further described in detail below with reference to the accompanying drawings and examples. It can be understood that the specific examples described herein are only for explaining the relevant content and are not intended to limit the present disclosure. Additionally, it should be noted that for the sake of description, only parts related to the present disclosure are shown in the drawings.

[0041] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.

[0042] In the related art, during the data consistency verification process before a new version database cluster is put into production, the field values of the same fields in the new version database cluster and the old version database cluster are compared one by one. Since the number of fields and field values is relatively large, the comparison process takes a large amount of time, thereby resulting in low data consistency verification efficiency.

[0043] Therefore, the present disclosure proposes a data consistency verification method.

[0044] The data consistency verification method of the present disclosure can be used by an electronic device to perform data consistency verification on a new version database cluster and an old version database cluster before the new version database cluster is put into production. In the present disclosure, the electronic device includes but is not limited to servers, mobile phones, tablet computers, laptop computers, personal computers, wearable devices, teller machines, etc.

[0045] For ease of description and to make the technical solutions of the specific embodiments of the present disclosure easier to understand, before describing the data consistency verification method implemented in the present disclosure, the technical terms involved in the specific embodiments of the present disclosure are explained as follows: A database cluster is a virtual single database logical image formed by using at least two or more database servers. It can provide transparent data services to clients like a single database system.

[0046] Data consistency verification refers to the process of comparing data between different data sources or data replicas to ensure that they are consistent in content.

[0047] A signature is data used to represent a data unit. The signature can be attached to the data unit it represents.

[0048] An abstract syntax tree is an abstract representation of the source code structure, which shows the syntax structure of a programming language in a tree-like form. Each node on the tree represents a structure in the source code.

[0049] SQL (Structured Query Language) is a special-purpose programming language, a database query and programming language used to access data and query, update, and manage relational database systems.

[0050] Hash calculation is a process of converting input data of any length into an output of a fixed length through a specific algorithm.

[0051] An aggregation function, also known as an aggregation operator, is a class of functions in a database query language used to perform calculations on a set of values and return a single value, such as summation, average, maximum / minimum value, etc.

[0052] An aggregation dimension is the basis for grouping data when performing an aggregation operation.

[0053] Figure 1 Fig. shows the overall flowchart of the data consistency verification method M100 according to an embodiment of the present disclosure. As Figure 1 shown, the method includes steps S110 to S140. Among them, the method can be executed by an electronic device such as a server, a mobile phone, or a computer.

[0054] Specifically, Figure 1 the method shown includes: S110. Obtain a first query result and a second query result, where the first query result is the query result obtained after a target query instruction is executed on a new version database cluster, and the second query result is the query result obtained after the target query instruction is executed on an old version database cluster. The first query result and the second query result each include multiple retrieved records; Exemplarily, an SQL statement can be used as the target query instruction. The target query instruction can include multiple SQL statements. One SQL statement can retrieve at least one record, and one record can include field values of at least one field.

[0055] The first query result and the second query result are obtained through the same target query instruction. For a new version database cluster, if an SQL statement that can be normally executed on an old version database cluster can also be normally executed on this new version database cluster, then it can be considered that this new version database cluster meets the SQL syntax compatibility.

[0056] S120. Generate signatures for each record among the multiple records respectively to obtain a first signature result and a second signature result, where the first signature result includes the signatures of each record among the multiple records in the first query result, and the second signature result includes the signatures of each record among the multiple records in the second query result. The signature is used to represent the corresponding record; A signature is data used to represent a record. The signature can be attached to the record it represents. Signatures corresponding to multiple records with different contents are different, and signatures corresponding to multiple records with the same content can be the same. Compared with the field values of numerous fields in one record, the signature is more concise. Therefore, using signatures instead of field values for data consistency verification can save more time and thus improve efficiency.

[0057] S130. Perform data consistency verification on the first signature result and the second signature result to obtain a data consistency verification result.

[0058] In an example, the first signature result includes multiple signatures, and the second signature result also includes multiple signatures. The signatures in the first signature result can be compared one by one with the signatures in the second signature result. If the signatures in the first signature result are exactly the same as the signatures in the second signature result, the data consistency verification result is verification passed; if there are differences between the signatures in the first signature result and the signatures in the second signature result, the data consistency verification result is verification failed. In addition, it is also possible to perform data consistency verification after respectively counting the signatures in the first signature result and the signatures in the second signature result, which is not limited herein.

[0059] In the data consistency verification method of the embodiments of the present disclosure, since the signature can represent the records in the query result, the data consistency verification result obtained by performing data consistency verification on the first signature result and the second signature result can represent the data consistency verification result obtained by performing data consistency verification on the first query result and the second query result. Furthermore, data consistency verification of the new version database cluster and the old version database cluster is realized through signatures. At the same time, since one signature can represent the field values of multiple fields in one record, data consistency verification is performed using the signature instead of the field values, which can save more time and thus improve efficiency.

[0060] Regarding step S120, in some embodiments of the present disclosure, each record includes field values of at least one field. Correspondingly, step S120 may include steps S121 to S123 as Figure 2 shown.

[0061] S121. For each of multiple records, concatenate the field values of at least one field of the record to obtain a string corresponding to the record.

[0062] Exemplarily, the concatenation order can be determined based on the arrangement of the records. For example, if the multiple records are arranged in rows, the concatenation order can be from left to right, that is, based on the order from left to right, the field values of all fields in this record are sequentially concatenated together to obtain the string corresponding to this record; if the multiple records are arranged in columns, the concatenation order can be from top to bottom, that is, based on the order from top to bottom, the field values of all fields in this record are sequentially concatenated together to obtain the string corresponding to this record. Among them, for empty field values, "null" can be used for concatenation. It can be directly concatenated or concatenated with commas in between, which is not limited here. The same concatenation order can be adopted for each of the multiple records to ensure uniform data format and improve the accuracy of the data consistency verification result.

[0063] S122. Perform a hash calculation on the string corresponding to the record to obtain a hash value corresponding to the record.

[0064] Hash calculation can convert different strings corresponding to different records into hash values of a unified length, which is convenient for data consistency verification. Exemplarily, any one of the MD5 (Message-Digest Algorithm), SHA (Secure Hash Algorithm), and HMAC (Hash-based Message Authentication Code) algorithms can be used for hash calculation.

[0065] S123. Use the hash value corresponding to the record as the signature of the corresponding record among multiple records.

[0066] For the data consistency verification method in the above embodiment, the hash calculation is performed on the string obtained by concatenating all field values in the record, so that the obtained hash value can accurately and completely represent the corresponding record. Moreover, using the hash value as the signature of the corresponding record can improve the accuracy of data consistency verification.

[0067] Regarding step S120, in other embodiments, it is also possible to perform feature extraction on the field values of at least one field in the record, and then use the feature extraction result as the signature of the record, thereby providing support for data consistency verification.

[0068] Regarding step S130, in some embodiments of the present disclosure, it may include steps S131 to S132 as Figure 3 shown.

[0069] S131. Perform statistics on the signatures in the first signature result to obtain a first statistical result, and perform statistics on the signatures in the second signature result to obtain a second statistical result.

[0070] S132. Perform data consistency verification on the first statistical result and the second statistical result to obtain a data consistency verification result.

[0071] For the data consistency verification method in the above embodiment, by performing statistics on the signature results and performing data consistency verification based on the statistical results, it is possible to better evaluate the consistency between the first signature result and the second signature result as a whole, that is, the consistency between the first query result and the second query result. At the same time, performing data consistency verification based on the statistical results can avoid the problem that it is impossible to accurately compare due to different data return orders, solve the problem of data order inconsistency caused by the change of the data sorting algorithm processing strategy in different versions of the cluster, and is particularly suitable for distributed query engines. In addition, performing statistics on the signatures instead of on the string obtained by concatenating the field values of multiple fields in the record, since the signature is more concise than the concatenated string, can improve the processing speed and simplify the data consistency verification logic.

[0072] Regarding step S131, in some embodiments of the present disclosure, it may include steps S1311 to S1313 as Figure 4 shown.

[0073] S1311. For the first signature result and the second signature result, respectively determine the number of each signature in the signature result where the signature is located.

[0074] Multiple identical signatures may exist in a signature result. By determining the number of each signature in the signature result where the signature is located, the data can be simplified and the efficiency of data consistency verification can be improved. Specifically, for the first signature result, determine the number of each signature in the first signature result; for the second signature result, determine the number of each signature in the second signature result.

[0075] S1312. Generate key-value pairs with the signature as the key and the number of the signature in the signature result where the signature is located as the value.

[0076] Specifically, for each signature in the first signature result, generate key-value pairs with the signature as the key and the number of the signature in the first signature result as the value; for each signature in the second signature result, generate key-value pairs with the signature as the key and the number of the signature in the second signature result as the value.

[0077] S1313. Take the key-value pairs corresponding to the signatures in the first signature result as the first statistical result, and take the key-value pairs corresponding to the signatures in the second signature result as the second statistical result.

[0078] Exemplarily, the first statistical result and the second statistical result can be in the form of a map respectively. Map is a data structure used to store a set of key-value pairs.

[0079] Please combine Figure 5 , in an example, the first query result includes 3 records, namely "Zhang San 22", "Zhang San 22" and "Li Si 15". Concatenate the field values in each record in the first query result respectively to obtain 3 strings: "Zhang San, 22", "Zhang San, 22" and "Li Si, 15". Then generate the signatures corresponding to each string. Among them, "Zhang San, 22" corresponds to the first signature, "Li Si, 15" corresponds to the second signature, and the first signature is different from the second signature. After statistics, it can be determined that the number of the first signature is 2 and the number of the second signature is 1. Store the key-value pairs with the signature as the key and the signature number as the value in the form of a map to obtain the first statistical result. The second query result includes 3 records, namely "Zhang San 22", "Zhang San 22" and "Li Si 15.0". Concatenate the field values in each record in the second query result respectively to obtain 3 strings: "Zhang San, 22", "Zhang San, 22" and "Li Si, 15.0". Then generate the signatures corresponding to each string. Among them, "Zhang San, 22" corresponds to the first signature, "Li Si, 15.0" corresponds to the third signature, and the third signature is different from the first signature and the second signature. After statistics, it can be determined that the number of the first signature is 2 and the number of the third signature is 1. Store the key-value pairs with the signature as the key and the signature number as the value in the form of a map to obtain the second statistical result. Finally, by comparing the first statistical result and the second statistical result, the data consistency verification result can be obtained.

[0080] The data consistency verification method of the above embodiment stores the signature and its quantity in the form of key-value pairs, which is convenient for management and query, and improves the efficiency of data processing. Moreover, by generating key-value pairs, the distribution of each signature in different query results becomes more intuitive, which helps the electronic device to complete data consistency verification faster.

[0081] Regarding step S132, in some embodiments of the present disclosure, it may include steps S1321 to S1323 as Figure 6 shown.

[0082] S1321. Determine whether the first statistical result is the same as the second statistical result.

[0083] Exemplarily, when both the first statistical result and the second statistical result are in the form of a map, it is possible to first determine whether the number of key-value pairs of the two maps is the same. If they are not the same, it is determined that the first statistical result is different from the second statistical result. If they are the same, it is further determined whether the keys in one map exist with the same keys and the same values in the other map. If so, it is determined that the first statistical result is the same as the second statistical result; otherwise, it is determined that the first statistical result is different from the second statistical result.

[0084] S1322. In response to the first statistical result being the same as the second statistical result, determine that the data consistency verification result is verification passed.

[0085] The first statistical result is the same as the second statistical result, that is, the first signature result is the same as the second signature result, which is also the same as the first query result and the second query result. Furthermore, the data consistency verification result is verification passed.

[0086] S1323. In response to the first statistical result being different from the second statistical result, determine that the data consistency verification result is verification failed.

[0087] The data consistency verification method of the above embodiment can quickly determine the result of data consistency verification by directly comparing the first statistical result and the second statistical result, reducing the possibility of misjudgment. Since the statistical result is the result of signature statistics, and the signature represents an entire record, it is possible to efficiently perform data consistency verification without knowing the SQL query fields, avoiding the dependence on SQL query fields. It can be understood that in the related art, since the field values of each field in the record are compared one by one, it is necessary to pre-parse the SQL statement to determine the query fields in the SQL statement. For complex queries and aggregation operations, this method is inefficient and error-prone.

[0088] As a possible implementation, after step S130, it may further include asFigure 7 Steps S140 to S160 as shown.

[0089] S140. In response to the data consistency check result being that the check fails, based on the first rewriting rule, rewrite the field values of at least one field in the first target object to obtain the rewritten first target object, where the first target object is one of the first query result and the second query result.

[0090] Exemplarily, the first rewriting rule can be to update the field value with a target number of decimal places in the first query result or the second query result to the sum value of this field value and a preset decimal (such as 0.000001). The preset decimal is a decimal with a target number of digits. It is possible to rewrite all field values with a target number of decimal places in the first target object, or only rewrite the field values with a target number of decimal places that are different from the field values of the same field in the second target object. This is not limited here.

[0091] It can be understood that the processing strategies for decimals exceeding the target number of digits in the new and old version database clusters may be different. For example, the processing strategy of the new version database cluster is rounding, and the processing strategy of the old version database cluster is truncation. When the target number of digits is 6, for "623.0156796", the field value stored in the new version database cluster is "623.015680", and the field value stored in the old version database cluster is "623.015679", which will cause the data consistency check to fail. Therefore, based on the first rewriting rule, rewriting the field values of at least one field in the first target object can eliminate the problem of data consistency check failure caused by different processing strategies for decimals exceeding the target number of digits. The first rewriting rule can be preset based on the problem (i.e., the preset reason) that may cause the data consistency check to fail.

[0092] S150. Compare the rewritten first target object with the second target object, where the second target object is the other one of the first query result and the second query result.

[0093] Exemplarily, the rewritten first target object and the second target object can be compared with reference to steps S120 and S130, thereby improving the data comparison efficiency.

[0094] S160. When the rewritten first target object is consistent with the second target object in data, determine that the original reason for the check failure is the preset reason corresponding to the first rewriting rule.

[0095] After the first target object and the second target object are rewritten and the data is consistent, it indicates that the first rewriting rule can eliminate the problem that causes the data consistency check to fail. In other words, the reason for the failure in step S130 is the preset reason corresponding to the first rewriting rule (for example, different processing strategies for decimals exceeding the target number of digits).

[0096] For the data consistency check method in the above embodiment, when the data consistency check fails, the first target object is rewritten based on the first rewriting rule, and further data comparison is performed, which helps to automatically and accurately determine the reason for the failure in the check, greatly reducing the cost of manually analyzing a large number of abnormal SQLs, and realizing efficient classification of abnormal SQLs and problem location.

[0097] As another possible embodiment, after step S150, steps S170 to S190 as shown in Figure 8 may also be included.

[0098] S170. In the case where the data of the first target object after rewriting is inconsistent with the data of the second target object, based on the second rewriting rule, the field values of at least one field in the first query result and the field values of at least one field in the second query result are rewritten respectively to obtain the rewritten first query result and the rewritten second query result.

[0099] Exemplarily, the second rewriting rule may be to increase the number of decimal places of the field values in the first query result and the second query result with the number of decimal places less than the second target number of digits (for example, 6 digits) to the second target number of digits by adding 0 at the end of the decimal. It is possible to rewrite all the field values in the first target object with the number of decimal places less than the second target number of digits, or only rewrite the field values in the first target object that are different from the field values of the same field in the second target object and have the number of decimal places less than the second target number of digits, which is not limited herein.

[0100] It can be understood that the retention strategies for invalid 0s at the end of decimals in the new and old version database clusters may be different. For example, the processing strategy of the new version database cluster is to retain two invalid 0s (such as 100.23400), and the processing strategy of the old version database cluster is to retain three invalid 0s (such as 100.234000), which will cause the data consistency check to fail. Therefore, based on the second rewriting rule, the field values of at least one field in the first query result and the field values of at least one field in the second query result are rewritten respectively, which can eliminate the problem that causes the data consistency check to fail due to different retention strategies for invalid 0s at the end of decimals. The second rewriting rule can be preset based on the problems (i.e., preset reasons) that may cause the data consistency check to fail.

[0101] S180. Compare the data of the first rewritten query result with the second rewritten query result.

[0102] Exemplarily, steps S120 and S130 can be referred to for comparing the data of the first rewritten query result with the second rewritten query result, thereby improving the data comparison efficiency.

[0103] S190. When the data of the first rewritten query result is consistent with that of the second rewritten query result, determine that the preset reason corresponding to the second rewriting rule is the reason for the failed verification.

[0104] When the data of the first rewritten query result is inconsistent with that of the second rewritten query result, relevant personnel can be prompted to manually determine the reason for the failed data consistency verification.

[0105] In the data consistency verification method of the above embodiments, when the data of the first target object is inconsistent with that of the second target object, the first query result and the second query result are rewritten based on the second rewriting rule, and further data comparison is performed, which helps to automatically and accurately determine the reason for the failed verification, greatly reducing the cost of manually analyzing a large number of abnormal SQLs, and realizing efficient classification of abnormal SQLs and problem location.

[0106] In some embodiments of the present disclosure, before step S110, steps S210 to S240 as shown in Figure 9 may also be included.

[0107] S210. Obtain multiple preset query instructions. The preset query instruction can be an SQL statement.

[0108] S220. Calculate the signature of each preset query instruction among the multiple preset query instructions respectively.

[0109] The signature of the preset query instruction can represent the preset query instruction. The signatures of multiple preset query instructions may be the same or different.

[0110] S230. Based on the signatures of the preset query instructions, screen the multiple preset query instructions to obtain multiple preset query instructions with different signatures.

[0111] S240. Determine the target query instruction according to the multiple preset query instructions with different signatures.

[0112] The multiple preset query instructions with different signatures can be directly used as the target query instruction, or can be used as the target query instruction after certain processing of the multiple preset query instructions with different signatures, which is not limited herein.

[0113] The data consistency verification method of the above embodiment filters multiple preset query instructions based on signatures, thereby filtering out the preset query instructions with duplicate signatures, reducing the number of target query instructions, and thus being able to improve the acquisition speed of the first query result and the second query result, and avoiding wasting resources due to repeatedly executing a large number of identical target query instructions. At the same time, determining the target query instructions based on multiple preset query instructions with different signatures ensures the diversity and representativeness of the target query instructions, increases the coverage of data consistency verification, and improves the accuracy of data consistency verification.

[0114] Regarding step S220, in some embodiments of the present disclosure, it may include steps S221 to S223 as Figure 10 shown.

[0115] S221. For each preset query instruction among the multiple preset query instructions, determine the string corresponding to the preset query instruction.

[0116] The content of a preset query instruction can be directly used as the string corresponding to the preset query instruction, or the content of a preset query instruction can be processed and then used as the string corresponding to the preset query instruction, which is not limited herein. Exemplarily, the string corresponding to the preset query instruction can at least characterize the structure of the preset query instruction.

[0117] S222. Perform a hash calculation on the string corresponding to the preset query instruction to obtain the hash value corresponding to the preset query instruction.

[0118] The hash calculation can convert different strings corresponding to different preset query instructions into hash values of a unified length, thus facilitating screening. Exemplarily, any one of the MD5 (Message-Digest Algorithm), SHA (Secure Hash Algorithm), and HMAC (Hash-based Message Authentication Code) algorithms can be used for the hash calculation.

[0119] S223. Use the hash value corresponding to the preset query instruction as the signature of the corresponding preset query instruction among the multiple preset query instructions.

[0120] The data consistency verification method of the above embodiment performs a hash calculation on the string corresponding to the preset query instruction, and the obtained hash value can accurately represent the corresponding preset query instruction, thereby improving the accuracy of screening the preset query instructions.

[0121] Regarding step S221, in some embodiments of the present disclosure, it may include asFigure 11 Steps S2211 to S2213 shown above.

[0122] S2211. Convert the preset query instruction into an abstract syntax tree to obtain the abstract syntax tree corresponding to the preset query instruction.

[0123] The preset query instruction can be converted into an abstract syntax tree by performing lexical analysis and syntax analysis on the preset query instruction.

[0124] S2212. Replace the parameter values in the abstract syntax tree corresponding to the preset query instruction with preset strings.

[0125] Multiple preset query instructions may have the same structure, only the parameter values are different. For example, one preset query instruction queries records with an age greater than 20, and another preset query instruction queries records with an age greater than 40. If all such multiple preset query instructions are executed once in the new and old version database clusters, it will cause unnecessary resource waste. Therefore, only one of the preset query instructions needs to be selected for execution.

[0126] By replacing the parameter values in the abstract syntax tree corresponding to the preset query instruction with preset strings, the abstract syntax trees of preset query instructions with the same structure but different parameter values can be converted into the same abstract syntax tree, thus facilitating the subsequent screening of preset query instructions.

[0127] The preset string may include letters and / or constants.

[0128] S2213. Parse the replaced abstract syntax tree corresponding to the preset query instruction into a string to obtain the string corresponding to the preset query instruction.

[0129] The structure of the replaced abstract syntax tree can be traversed and translated through the StarRocks component in related technologies, and then the replaced abstract syntax tree corresponding to the preset query instruction can be parsed into a string.

[0130] In the data consistency verification method of the above embodiment, by replacing the parameter values in the abstract syntax tree corresponding to the preset query instruction with preset strings, the strings corresponding to multiple preset query instructions with the same structure but different parameter values are unified, eliminating the problem that the strings corresponding to preset query instructions with the same structure are different due to different parameter values. Furthermore, the signatures determined by multiple preset query instructions with the same structure but different parameter values based on the same string are the same, thus facilitating the screening of preset query instructions with the same structure but different parameter values and being beneficial to improving the acquisition speed of the first query result and the second query result.

[0131] Regarding step S240, as a possible implementation, it may include as Figure 12Steps S241 to S242 shown above.

[0132] S241. Delete the aggregation operator in the first query instruction, and retain the aggregation object corresponding to the aggregation operator in the first query instruction, to obtain the processed first query instruction, where the first query instruction is a preset query instruction with an aggregation operator among multiple preset query instructions with different signatures.

[0133] Aggregation operators include, but are not limited to, COUNT(), SUM(), AVG(), MAX(), and MIN(), etc. Among them, COUNT() is used to calculate the number of rows in a table or the number of rows that meet specific conditions, SUM() is used to calculate the sum, AVG() is used to calculate the average, MAX() is used to find the maximum value, and MIN() is used to find the minimum value. The aggregation object is the object on which the set operator acts, usually located within the parentheses of the aggregation operator.

[0134] In some embodiments, the first query instruction further includes an aggregation dimension and a keyword for specifying the aggregation dimension. Correspondingly, in step S241, in addition to deleting the aggregation operator in the first query instruction and retaining the aggregation object corresponding to the aggregation operator in the first query instruction, the aggregation dimension and the keyword for specifying the aggregation dimension in the first query instruction are also deleted, so as to obtain the processed first query instruction. In this way, it is ensured that the processed first query instruction conforms to the syntax regulations.

[0135] Please refer to Figure 13 , in an example, the first query instruction is SELECT age,SUM(score) FROM Student GROUP by age, where the included aggregation operator is SUM(), the aggregation object is score, the aggregation dimension is age, and GROUP by is the keyword for specifying the aggregation dimension. Then, delete the aggregation operator, the aggregation dimension, the keyword for specifying the aggregation dimension, and retain the aggregation object corresponding to the aggregation operator in the first query instruction in the first query instruction, to obtain the processed first query instruction as SELECT age,score FROM Student, where both age and score are query objects in the processed first query instruction.

[0136] S242. Use the processed first query instruction and the second query instruction as the target query instruction, where the second query instruction is a preset query instruction without an aggregation operator among multiple preset query instructions with different signatures.

[0137] In the data consistency verification method of the above embodiment, since the aggregation operator in the first query instruction is deleted and the aggregation object is retained, there is no longer an aggregation object in the target query instruction, simplifying the target query instruction and facilitating the improvement of the acquisition speed of the first query result and the second query result. At the same time, since the aggregation object is still retained as the query object in the target query instruction, the first query result and the second query result obtained can include the field values corresponding to the aggregation object (i.e., the detailed data before aggregation), thus facilitating the subsequent positioning of the reason for the failure of data consistency verification based on the first query result and the second query result, solving the problem that it is difficult to locate the reason for the failure of data consistency verification in the case of aggregation calculation for complex target query instructions, and improving the accuracy of locating the reason for the failure of data consistency verification.

[0138] Regarding step S240, as another possible embodiment, it may include steps S241' to S243' as Figure 14 shown.

[0139] S241'. Generate a target sorting rule according to the query object in the third query instruction, where the third query instruction is a preset query instruction among multiple preset query instructions with different signatures that limits the number of records in the query result and has no sorting rule.

[0140] Exemplarily, when the number of query objects in the third query instruction is one and the field value corresponding to the query object is of numerical type, the target sorting rule may be sorting according to the numerical size; when the number of query objects in the third query instruction is one and the field value corresponding to the query object is of text type, the target sorting rule may be sorting according to alphabetical order / lexicographical order; when the number of query objects in the third query instruction is multiple, the target sorting rule may be sorting according to the priority order of multiple query objects and sorting level by level according to the field value type corresponding to each query object.

[0141] S242'. Add the target sorting rule to the third query instruction to obtain the processed third query instruction, where the processed third query instruction is used to sort multiple records corresponding to the query object based on the target sorting rule and select the top number of records from the sorting result as the query result.

[0142] It can be understood that if the number of records in the query result is restricted in the preset query instruction, then when the number of records that meet the conditions retrieved is greater than the restricted number, the randomness of the distributed engine will cause different rows to be returned by the new and old version database clusters. For example, if the restricted number is 10 and the number of records that meet the conditions retrieved is 100 (with IDs from 1 to 100), then the first query result returned by the new version database cluster may be 10 records with IDs from 1 to 10, while the second query result returned by the old version database cluster may be 10 records with IDs from 21 to 30. In this way, even if the data in the new and old version database clusters is consistent, due to the different selection of query results, the data consistency verification result may be wrongly determined as failed verification. Since the processed third query instruction can be used to sort multiple records corresponding to the query object based on the target sorting rule and select the top number of records from the sorting result as the query result, therefore, when the data in the new and old version database clusters is consistent and the number of records that meet the conditions retrieved is greater than the restricted number, the first query result and the second query result can be kept consistent, thereby avoiding misjudgment and improving the accuracy of the data consistency verification result.

[0143] Please combine Figure 15 , in an example, the third query instruction is SELECT age,phone_number FROM Student LIMIT 10, where the query object is age,phone_number. Then, the target sorting rule generated according to the query object is ORDER by age,phone_number, that is, sort by age first, and if the ages are the same, sort by phone_number. Finally, add the target sorting rule to the third query instruction, and the obtained processed third query instruction is SELECT age,phone_number FROM Student ORDER by age,phone_number LIMIT 10.

[0144] It should be noted that the specific values mentioned above are only for detailed illustration of the implementation of the present disclosure as examples and should not be construed as a limitation of the present disclosure. In other examples or embodiments or implementations, other values can be selected according to the present disclosure, and no specific limitation is made here.

[0145] S243’: Use the processed third query instruction and the fourth query instruction as the target query instructions, where the fourth query instruction is the preset query instruction other than the third query instruction among multiple preset query instructions with different signatures.

[0146] In the data consistency verification method of the above embodiment, a target sorting rule is added to the third query instruction, so that the processed third query instruction is used to sort multiple records corresponding to the query object based on the target sorting rule and select the top several records from the sorting result as the query result. Furthermore, when the data in the new and old version database clusters is consistent and the number of records that meet the conditions is greater than the limit number, the first query result and the second query result can be kept consistent, thereby avoiding misjudgment and improving the accuracy of the data consistency verification result.

[0147] Please combine Figure 16 , in an example, the data consistency verification method may include the following steps S201 to step S218. The content related to steps S201 to S218 can be referred to the description of the above embodiment. For the sake of brevity, it will not be repeated here.

[0148] In step S201, obtain multiple preset query instructions.

[0149] In step S202, calculate the signature of each preset query instruction in the multiple preset query instructions respectively.

[0150] In step S203, based on the signatures of the preset query instructions, filter the multiple preset query instructions to obtain multiple preset query instructions with different signatures.

[0151] In step S204, determine the target query instruction according to the multiple preset query instructions with different signatures.

[0152] In step S205, based on the target query instruction, obtain the first query result and the second query result.

[0153] In step S206, generate the signature of each record in the multiple records respectively.

[0154] In step S207, perform data consistency verification on the first signature result and the second signature result.

[0155] In step S208, determine whether the data consistency verification passes. If so, enter step S209; otherwise, enter step S210.

[0156] In step S209, when the data consistency verification passes, prompt the data consistency verification result.

[0157] In step S210, when the data consistency verification fails, rewrite the field value of at least one field in the first target object based on the first rewriting rule to obtain the rewritten first target object.

[0158] In step S211, compare the rewritten first target object with the second target object for data.

[0159] In step S212, determine whether the rewritten first target object is consistent with the second target object in data. If so, proceed to step S213; otherwise, proceed to step S214.

[0160] In step S213, when the rewritten first target object is consistent with the second target object in data, determine that the preset reason for the failed verification is the preset reason corresponding to the first rewriting rule.

[0161] In step S214, when the rewritten first target object is inconsistent with the second target object in data, based on the second rewriting rule, rewrite the field values of at least one field in the first query result and the field values of at least one field in the second query result respectively to obtain the rewritten first query result and the rewritten second query result.

[0162] In step S215, compare the rewritten first query result with the rewritten second query result for data.

[0163] In step S216, determine whether the rewritten first query result is consistent with the rewritten second query result in data. If so, proceed to step S217; otherwise, proceed to step S218.

[0164] In step S217, when the rewritten first query result is consistent with the rewritten second query result in data, determine that the preset reason for the failed verification is the preset reason corresponding to the second rewriting rule.

[0165] In step S218, when the rewritten first query result is inconsistent with the rewritten second query result in data, prompt for manual intervention.

[0166] Based on any of the above embodiments, the present disclosure further provides a data consistency verification device. Figure 17 It is a structural schematic block diagram of a data consistency verification device according to an embodiment of the present disclosure. As Figure 17 shown, the data consistency verification device includes: A first acquisition module 110, configured to acquire a first query result and a second query result, where the first query result is the query result obtained after the new version database cluster executes a target query instruction, the second query result is the query result obtained after the old version database cluster executes the target query instruction, and the first query result and the second query result respectively include multiple queried records; A generation module 120 is configured to generate signatures for each record in multiple records respectively, obtaining a first signature result and a second signature result, where the first signature result includes signatures for each record in multiple records of a first query result, and the second signature result includes signatures for each record in multiple records of a second query result, and the signature is used to represent the corresponding record. A verification module 130 is configured to perform data consistency verification on the first signature result and the second signature result, obtaining a data consistency verification result.

[0167] The above data consistency verification device may be in the form of computer software, and each module of the above data consistency verification device may be implemented by computer software modules.

[0168] In some embodiments of the present disclosure, each record includes field values of at least one field; correspondingly, the generation module 120 is configured to: for each record in multiple records, splice the field values of at least one field of the record to obtain a string corresponding to the record; perform a hash calculation on the string corresponding to the record to obtain a hash value corresponding to the record; and use the hash value corresponding to the record as the signature of the corresponding record in multiple records.

[0169] In some embodiments of the present disclosure, the verification module 130 is configured to: perform statistics on the signatures in the first signature result to obtain a first statistical result, and perform statistics on the signatures in the second signature result to obtain a second statistical result; and perform data consistency verification on the first statistical result and the second statistical result to obtain a data consistency verification result.

[0170] In some embodiments of the present disclosure, the verification module 130 is configured to: for the first signature result and the second signature result, respectively determine the quantity of each signature in the signature result where the signature is located; generate key-value pairs with the signature as the key and the quantity of the signature in the signature result where the signature is located as the value; and use the key-value pairs corresponding to the signatures in the first signature result as the first statistical result, and use the key-value pairs corresponding to the signatures in the second signature result as the second statistical result.

[0171] In some embodiments of the present disclosure, the verification module 130 is configured to: determine whether the first statistical result is the same as the second statistical result; in response to the first statistical result being the same as the second statistical result, determine that the data consistency verification result is verification passed; and in response to the first statistical result being different from the second statistical result, determine that the data consistency verification result is verification failed.

[0172] In some embodiments of the present disclosure, the data consistency verification device further includes: a first rewriting module, configured to, in response to the data consistency verification result being verification failed, rewrite the field values of at least one field in the first target object based on a first rewriting rule to obtain a rewritten first target object, where the first target object is one of a first query result and a second query result; a first comparison module, configured to perform data comparison between the rewritten first target object and a second target object, where the second target object is the other of the first query result and the second query result; and a first determination module, configured to, when the rewritten first target object is consistent with the second target object in data, determine that the reason for the verification failure is a preset reason corresponding to the first rewriting rule.

[0173] In some embodiments of the present disclosure, the data consistency verification device further includes: a second rewriting module, configured to, when the rewritten first target object is inconsistent with the second target object in data, rewrite the field values of at least one field in the first query result and the field values of at least one field in the second query result respectively based on a second rewriting rule to obtain a rewritten first query result and a rewritten second query result; a second comparison module, configured to perform data comparison between the rewritten first query result and the rewritten second query result; and a second determination module, configured to, when the rewritten first query result is consistent with the rewritten second query result in data, determine that the reason for the verification failure is a preset reason corresponding to the second rewriting rule.

[0174] In some embodiments of the present disclosure, the data consistency verification device further includes: a second acquisition module, configured to acquire a plurality of preset query instructions; a calculation module, configured to calculate the signature of each of the plurality of preset query instructions respectively; a screening module, configured to screen the plurality of preset query instructions based on the signatures of the preset query instructions to obtain a plurality of preset query instructions with mutually different signatures; and a third determination module, configured to determine a target query instruction according to the plurality of preset query instructions with mutually different signatures.

[0175] In some embodiments of the present disclosure, the calculation module is configured to: for each of the plurality of preset query instructions, determine the string corresponding to the preset query instruction; perform hash calculation on the string corresponding to the preset query instruction to obtain the hash value corresponding to the preset query instruction; and use the hash value corresponding to the preset query instruction as the signature of the corresponding preset query instruction in the plurality of preset query instructions.

[0176] In some embodiments of the present disclosure, the calculation module is configured to: convert a preset query instruction into an abstract syntax tree to obtain the abstract syntax tree corresponding to the preset query instruction; replace the parameter values in the abstract syntax tree corresponding to the preset query instruction with preset strings; and parse the replaced abstract syntax tree corresponding to the preset query instruction into a string to obtain the string corresponding to the preset query instruction.

[0177] In some embodiments of the present disclosure, the third determination module is configured to: delete the aggregation operator in the first query instruction and retain the aggregation object corresponding to the aggregation operator in the first query instruction to obtain the processed first query instruction, where the first query instruction is a preset query instruction with an aggregation operator among multiple preset query instructions with different signatures; and use the processed first query instruction and the second query instruction as the target query instruction, where the second query instruction is a preset query instruction without an aggregation operator among multiple preset query instructions with different signatures.

[0178] In some embodiments of the present disclosure, the third determination module is configured to: generate a target sorting rule according to the query object in the third query instruction, where the third query instruction is a preset query instruction that limits the number of records in the query result and has no sorting rule among multiple preset query instructions with different signatures; add the target sorting rule to the third query instruction to obtain the processed third query instruction, where the processed third query instruction is used to sort multiple records corresponding to the query object based on the target sorting rule and select the top records from the sorting result as the query result; and use the processed third query instruction and the fourth query instruction as the target query instruction, where the fourth query instruction is a preset query instruction other than the third query instruction among multiple preset query instructions with different signatures.

[0179] The implementation processes of the functions and roles of each module in the above device are specifically described in detail in the implementation processes of the corresponding steps in the above method, and will not be repeated here.

[0180] The execution subject of the data consistency verification method in the specific embodiments of the present disclosure may be an electronic device such as a server, a mobile phone, or a computer.

[0181] Therefore, based on any of the above embodiments, the present disclosure further provides an electronic device, which can execute the data consistency verification method of any of the above embodiments described in the present disclosure.

[0182] Figure 18 It is a structural schematic diagram of an electronic device 1000 according to an embodiment of the present disclosure.

[0183] The hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application of the hardware and the overall design constraints. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0184] The bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Component (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one connection line is shown in this figure, but it does not mean that there is only one bus or one type of bus.

[0185] The present disclosure also provides a readable storage medium. A computer program is stored in the readable storage medium and is used to implement the above - mentioned method when executed by a processor. The "readable storage medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read - only memory (ROM), an erasable programmable read - only memory (EPROM or flash memory), an optical fiber device, and a portable read - only memory (CDROM), etc.

[0186] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part.

[0187] Computer programs or instructions can be stored in a readable storage medium or transmitted from one readable storage medium to another. For example, the computer programs or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any available medium that can be accessed or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.

[0188] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, system, or computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0189] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of multiple flows and / or blocks.

[0190] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of multiple flows and / or blocks.

[0191] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or boxes Figure 1 one process or a plurality of processes and / or boxes Figure 1 steps for implementing the functions specified in one box or a plurality of boxes.

[0192] In the description of this specification, the descriptions with reference to the terms "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc. mean that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described may be combined in any one or more embodiments / ways or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.

[0193] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0194] Those skilled in the art should understand that the above embodiments are only for clearly explaining the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or variations can be made on the basis of the above disclosure, and these changes or variations are still within the scope of the present disclosure.

Claims

1. A data consistency verification method, characterized in that, Including: Obtain a first query result and a second query result, where the first query result is the query result obtained after a target query instruction is executed by a new version database cluster, and the second query result is the query result obtained after the target query instruction is executed by an old version database cluster. The first query result and the second query result each include multiple records retrieved; Generate signatures for each of the multiple records respectively to obtain a first signature result and a second signature result, where the first signature result includes the signatures of each of the multiple records in the first query result, and the second signature result includes the signatures of each of the multiple records in the second query result. The signature is used to represent the corresponding record; and Perform data consistency verification on the first signature result and the second signature result to obtain a data consistency verification result.

2. The data consistency verification method according to claim 1, wherein Each of the records includes field values of at least one field; Generating signatures for each of the multiple records respectively includes: For each of the multiple records, splice the field values of at least one field of the record to obtain a string corresponding to the record; Perform a hash calculation on the string corresponding to the record to obtain a hash value corresponding to the record; and Use the hash value corresponding to the record as the signature of the corresponding record in the multiple records.

3. The data consistency verification method according to claim 1, wherein Performing data consistency verification on the first signature result and the second signature result to obtain a data consistency verification result includes: Perform statistics on the signatures in the first signature result to obtain a first statistical result, and perform statistics on the signatures in the second signature result to obtain a second statistical result; and Perform data consistency verification on the first statistical result and the second statistical result to obtain the data consistency verification result.

4. The data consistency verification method according to claim 1, wherein After performing data consistency verification on the first signature result and the second signature result to obtain a data consistency verification result, it further includes: In response to the data consistency verification result being verification failed, based on a first rewriting rule, rewrite the field values of at least one field in a first target object to obtain a rewritten first target object, where the first target object is one of the first query result and the second query result; Perform data comparison between the rewritten first target object and a second target object, where the second target object is the other of the first query result and the second query result; and When the rewritten first target object is consistent with the second target object in data, determine that the reason for the verification failure is the preset reason corresponding to the first rewriting rule.

5. The data consistency verification method according to any one of claims 1 to 4, characterized in that Before obtaining the first query result and the second query result, it further includes: Obtain multiple preset query instructions; Calculate the signatures of each of the multiple preset query instructions respectively; Based on the signatures of the preset query instructions, filter the multiple preset query instructions to obtain multiple preset query instructions with distinct signatures; and Determine the target query instruction according to the multiple preset query instructions with distinct signatures.

6. The data consistency verification method according to claim 5, wherein Determining the target query instruction according to a plurality of preset query instructions with different signatures includes: Deleting the aggregation operator in the first query instruction and retaining the aggregation object corresponding to the aggregation operator in the first query instruction to obtain a processed first query instruction, where the first query instruction is a preset query instruction with an aggregation operator among the plurality of preset query instructions with different signatures; and Using the processed first query instruction and the second query instruction as the target query instruction, where the second query instruction is a preset query instruction without an aggregation operator among the plurality of preset query instructions with different signatures.

7. The data consistency verification method according to claim 5, wherein Determining the target query instruction according to a plurality of preset query instructions with different signatures includes: Generating a target sorting rule according to the query object in the third query instruction, where the third query instruction is a preset query instruction that limits the number of records in the query result and has no sorting rule among the plurality of preset query instructions with different signatures; Adding the target sorting rule to the third query instruction to obtain a processed third query instruction, where the processed third query instruction is used to sort the multiple records corresponding to the query object based on the target sorting rule and select the top number of records from the sorting result as the query result; and Using the processed third query instruction and the fourth query instruction as the target query instruction, where the fourth query instruction is a preset query instruction other than the third query instruction among the plurality of preset query instructions with different signatures.

8. An electronic device, characterized in that, Including: A memory that stores execution instructions; And A processor that executes the execution instructions stored in the memory, enabling the processor to execute the data consistency verification method according to any one of claims 1 to 7.

9. A readable storage medium, characterized in that, Execution instructions are stored in the readable storage medium, and when the execution instructions are executed by a processor, they are used to implement the data consistency verification method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the data consistency verification method according to any one of claims 1 to 7.